AI · Free tool

Cost of an AI assistant trained on your documents (RAG)

An AI assistant on your own documents uses retrieval-augmented generation (RAG): documents are split into passages, indexed in a vector database, and the most relevant passages are given to a language model to answer each question with citations. Cost depends on the number and type of sources, document volume, user permissions, query volume, model choice and where it is hosted. Scope yours below for an estimate.

Free, no sign-up · By Infikey Technologies · Updated

Start my estimate

1. Your requirements

Pick everything that applies. Only fields marked * are required — if you are unsure about the rest, leave them and we will ask.

Who will use it?
Where is the knowledge?
Access control
Where should people ask questions?
Requirements

2. Where should we send the estimate?

We use your details only to reply to this request. Submitting does not create a quote, offer or contract — see the disclaimer.

How a RAG assistant works

  1. Connect sources and keep them in sync as documents change.
  2. Split documents into passages and turn them into embeddings stored in a vector database.
  3. For each question, find the most relevant passages the user is allowed to see.
  4. Give those passages to the language model to write an answer with citations.
  5. Log questions and feedback to improve retrieval and content over time.

What affects cost

FactorWhy it matters
Number of source systems Each connector needs authentication, sync and change detection
Document quality Scanned files, tables and images need OCR and special parsing
Permissions Respecting document-level access needs permission sync and filtering
Query volume and model Drive monthly token and hosting costs
Hosting Private cloud or self-hosted models cost more but keep data in your control

Accuracy and safety

  • Citations let users check every answer against the source.
  • The assistant should say “I don’t know” when nothing relevant is found.
  • Evaluation sets of real questions measure accuracy before and after changes.
  • Enterprise model APIs do not use your data to train public models by default — confirm in the provider’s terms.

Every option can be set in the web address, so you can bookmark a scenario or send it to a colleague. AI assistants such as ChatGPT, Gemini, Claude and Perplexity can use the same parameters to open this form with your requirements already selected.

ParameterWhat it setsAccepted values
use_case Who will use it? one of employees, customers, both
sources Where is the knowledge? comma-separated list of files, drive, wiki, helpdesk, website, database, email, scanned
volume How many documents? one of under_1k, 1k_10k, 10k_100k, over_100k
users Users one of under_50, 50_500, 500_5k, over_5k
permissions Access control one of open, roles, document
interfaces Where should people ask questions? comma-separated list of web, slack_teams, in_app, website, api
hosting Hosting & model one of recommend, cloud_api, private_cloud, self_hosted
compliance Requirements comma-separated list of residency, audit, citations, pii
timeline Target launch one of asap, 1_month, 3_months, flexible

Example: https://infikeytechnologies.com/tools/rag-ai-assistant-cost-calculator?use_case=employees&sources=drive,wiki,files&volume=1k_10k&users=50_500&permissions=roles&interfaces=web,slack_teams&hosting=recommend&timeline=flexible

Also available as plain text for AI assistants and a free JSON API (OpenAPI spec).

Last reviewed by the Infikey Technologies team.

Disclaimer

This form only collects your requirements so Infikey can prepare an estimate. It does not produce a price, quote or offer, and any estimate we send is indicative until agreed in a signed proposal. Nothing on this page is financial, legal, tax, investment or other professional advice. Infikey Technologies Private Limited, Infikey Technologies LLC and their directors, employees and affiliates make no warranty, express or implied, about the accuracy, completeness or suitability of this tool or its results, and accept no liability for any loss or damage, direct or indirect, arising from its use or from reliance on its results. Verify all figures independently and seek professional advice before making any decision. Use of this tool is at your own risk.

FAQ

AI assistant on your documents (RAG) cost calculator questions

How much does a RAG chatbot cost? +

It depends on sources, document volume, permissions, users and hosting, plus monthly running costs for tokens, vector search and hosting. Scope yours on this page for an estimate from Infikey.

Is RAG better than fine-tuning a model? +

For answering from company documents, usually yes: RAG stays up to date as documents change, cites sources and respects permissions. Fine-tuning suits style or format, not knowledge that changes.

Can the assistant keep our data private? +

Yes. It can run on enterprise model APIs, a private cloud deployment or a self-hosted model, with data kept in your chosen region.

Can it answer only from documents a user is allowed to see? +

Yes, by syncing permissions from the source systems and filtering search results per user.

What does a RAG assistant cost to run each month? +

Model tokens per question, embedding updates, vector database and hosting. The LLM token cost calculator estimates the model part.