Almost every team we meet budgets for a retrieval-augmented generation system as if it were a single line item: the price of the generation call. That is the visible number, the one on the model provider's pricing page, and it is usually the smallest part of what the system actually costs to run. The bill that arrives is bigger, and the reason it surprises people is that a production RAG system is not one API call. It is a pipeline of them, plus infrastructure that runs whether or not anyone is asking questions, plus the ongoing human cost of keeping the answers trustworthy.
This post breaks a production RAG system into the cost centres we actually see on invoices, walks through an illustrative worked example, and then covers the levers that move the total most. The numbers here are illustrative and deliberately rounded; your real figures depend on your corpus, your traffic, and your quality bar. The structure, though, is consistent across almost every deployment we have costed.
There are six places money goes in a production RAG system, and generation is only one of them.
Consider a system serving 100,000 questions a month against a corpus of a few hundred thousand documents. The generation call per query might be a cent or two. Add the query-time embedding, the vector search, and a re-ranking pass, and the marginal cost per query can quietly double or triple before the model writes a single word. Multiply by 100,000 and the per-query costs are real, but they are still not the whole story.
Now layer on the fixed costs. The vector store runs around the clock and bills whether traffic is high or low. Re-embedding the corpus on a weekly cadence is a scheduled expense. The eval suite runs on every prompt or model change, and each run is a batch of model calls. Add tracing and, if your domain demands it, a human spot-checking a percentage of answers. It is common for these standing costs to rival or exceed the entire per-query bill. The lesson is blunt: a RAG system at low traffic is dominated by fixed infrastructure and quality overhead, not by usage. You are paying to keep the lights on, not to answer questions.
Once you can see the cost centres, the optimisation targets are obvious, and most of them cost nothing but engineering judgment.
Sometimes the honest output of costing a RAG system is that RAG is the wrong tool. If your corpus is small and stable, you may not need retrieval at all; the content might fit in a prompt, or the problem might be better served by classification than generation. If your queries are narrow and repetitive, a fine-tuned smaller model can be cheaper to run than a retrieval pipeline, even after the cost of training. And if your data is a mess, no retrieval architecture will save you, and the money is better spent fixing the source. The point of costing the system properly is not only to make it cheaper. It is to find out, before you have built and staffed it, whether it should exist in the form you imagined.
That is exactly the question our fixed-price feasibility work is built to answer: what will this actually cost in production, and is it the right approach at all, written up in a way you can take to a budget conversation.
Before we start, please share a few details so we can follow up with you.
End this conversation? Your chat will be emailed to us.