Services Case Studies Insights About Start a project →

The real cost of RAG in production.

AI Published July 10, 2026 9 min read

Why the RAG bill surprises people

Almost every team we meet budgets for a retrieval-augmented generation system as if it were a single line item: the price of the generation call. That is the visible number, the one on the model provider's pricing page, and it is usually the smallest part of what the system actually costs to run. The bill that arrives is bigger, and the reason it surprises people is that a production RAG system is not one API call. It is a pipeline of them, plus infrastructure that runs whether or not anyone is asking questions, plus the ongoing human cost of keeping the answers trustworthy.

This post breaks a production RAG system into the cost centres we actually see on invoices, walks through an illustrative worked example, and then covers the levers that move the total most. The numbers here are illustrative and deliberately rounded; your real figures depend on your corpus, your traffic, and your quality bar. The structure, though, is consistent across almost every deployment we have costed.

The cost centres nobody quotes you

There are six places money goes in a production RAG system, and generation is only one of them.

  • Embedding your corpus. Every document has to be chunked and embedded before it can be retrieved, and re-embedded whenever the content changes or you switch embedding models. For a large, frequently-updated corpus this is a recurring cost, not a one-off. A few million chunks re-embedded on a schedule adds up.
  • The vector store. A managed vector database bills you continuously for stored vectors and for the compute that serves queries, whether traffic is one query an hour or a thousand. This is fixed infrastructure. At scale it is frequently the single largest line on the bill, and it is the one teams forget entirely at planning time.
  • Retrieval at query time. Each user question is itself embedded before it can be matched, so you pay a small embedding cost on every request in addition to the vector search.
  • Re-ranking. Naive nearest-neighbour retrieval is rarely good enough, so most serious systems retrieve a wide set of candidates and re-rank them with a cross-encoder or a model call. That is an extra inference on every query, and it scales directly with traffic.
  • Generation. The visible cost. But note that RAG makes it larger than a bare chat call, because you are stuffing retrieved context into the prompt. Input tokens are the hidden multiplier here: retrieve eight generous chunks and you can spend far more on input than on the answer the model writes.
  • Evaluation, observability, and human review. The cost of knowing the system still works. Eval runs are themselves model calls, often many per change. Tracing and logging infrastructure has its own bill. And in regulated or high-stakes settings, a human reviewing a sample of answers is a real, recurring line item that dwarfs the compute.

An illustrative worked example

Consider a system serving 100,000 questions a month against a corpus of a few hundred thousand documents. The generation call per query might be a cent or two. Add the query-time embedding, the vector search, and a re-ranking pass, and the marginal cost per query can quietly double or triple before the model writes a single word. Multiply by 100,000 and the per-query costs are real, but they are still not the whole story.

Now layer on the fixed costs. The vector store runs around the clock and bills whether traffic is high or low. Re-embedding the corpus on a weekly cadence is a scheduled expense. The eval suite runs on every prompt or model change, and each run is a batch of model calls. Add tracing and, if your domain demands it, a human spot-checking a percentage of answers. It is common for these standing costs to rival or exceed the entire per-query bill. The lesson is blunt: a RAG system at low traffic is dominated by fixed infrastructure and quality overhead, not by usage. You are paying to keep the lights on, not to answer questions.

The levers that actually move the number

Once you can see the cost centres, the optimisation targets are obvious, and most of them cost nothing but engineering judgment.

  • Retrieve less, not more. The input-token multiplier is the easiest big win. Better chunking and better retrieval let you pass three excellent chunks instead of ten mediocre ones. This cuts generation input cost and usually improves answer quality at the same time, because the model is not drowning in near-duplicates.
  • Route models by difficulty. Most queries do not need your most expensive model. A cheaper, smaller model handles the routine questions, and you escalate only the hard ones. Model routing is frequently the difference between a system that ships and one that is too expensive to justify.
  • Cache aggressively. Real user traffic is repetitive. Caching embeddings for repeated queries, and caching whole answers for common questions, removes work you were paying to do twice. In production this alone can take a meaningful bite out of the per-query cost.
  • Right-size the vector store. Because it is a fixed cost, the vector database rewards attention. Smaller embedding dimensions, quantisation, and pruning stale content all reduce a bill that you pay every hour of every day. For low-traffic systems, a Postgres extension may beat a dedicated managed vector service outright.
  • Make evals cheap enough to run often. A smaller, well-chosen eval set that you can afford to run on every change is worth more than an exhaustive one you run twice a year, because it catches regressions before they reach users, and a wrong answer in production is the most expensive outcome of all.

When the cost tells you to stop

Sometimes the honest output of costing a RAG system is that RAG is the wrong tool. If your corpus is small and stable, you may not need retrieval at all; the content might fit in a prompt, or the problem might be better served by classification than generation. If your queries are narrow and repetitive, a fine-tuned smaller model can be cheaper to run than a retrieval pipeline, even after the cost of training. And if your data is a mess, no retrieval architecture will save you, and the money is better spent fixing the source. The point of costing the system properly is not only to make it cheaper. It is to find out, before you have built and staffed it, whether it should exist in the form you imagined.

That is exactly the question our fixed-price feasibility work is built to answer: what will this actually cost in production, and is it the right approach at all, written up in a way you can take to a budget conversation.

Keep reading

Want a straight answer on what your RAG system will cost?

Start a conversation →
KT Solutions Assistant

Before we start, please share a few details so we can follow up with you.

Please enter your name and a valid email address.

End this conversation? Your chat will be emailed to us.