Three failure modes we see repeatedly in production retrieval systems, and the architectural shifts that fix them. With real numbers from a 100k-document deployment.
Retrieval-augmented generation became the default architecture for grounding a language model in a company's own documents somewhere around 2023, and for a lot of use cases it is still the right call. Chunk the documents, embed them, store the vectors, retrieve the nearest neighbours for a query, stuff them into a prompt, and let the model write the answer. It works well in a demo, and it works well for the first few thousand documents in production too. The trouble starts later, once the corpus grows, once documents start contradicting each other, and once real users start typing real questions instead of the tidy ones from the test set.
We have now rebuilt or repaired enough of these systems that the failure modes have stopped surprising us. They are structural, not accidental, and they show up in almost the same order every time. This post covers the three we see most often, what we changed to fix them, and roughly what that was worth, drawn from a recent engagement where the client's internal knowledge base had grown past 100,000 documents and the existing RAG pipeline had quietly become more of a liability than a feature.
Most RAG pipelines are built around naive fixed-size chunking, usually somewhere between 300 and 800 tokens, sometimes with a small overlap. This is a reasonable place to start and a bad place to stay. Fixed-size chunks slice through section boundaries, split tables in half, and separate a clause from the heading that gives it meaning. The embedding model then has to represent a fragment that, on its own, barely means anything.
The second half of this problem is what researchers call "lost in the middle." Once you retrieve, say, the top eight chunks and concatenate them into a context window, the model pays disproportionate attention to whatever is at the start and end of that context, and it quietly deprioritises the middle. So even when the correct chunk is technically retrieved, it can still lose the argument to a chunk sitting in a better position in the prompt.
Together, these two effects create what we call a retrieval ceiling: past a certain corpus size, adding more documents does not improve answer quality, it degrades it, because the retriever is returning more plausible-looking noise for every genuinely relevant chunk. In the deployment we mentioned above, accuracy on a held-out set of 200 real support questions had dropped from 81 percent at 10,000 documents to 58 percent at 100,000 documents, using the exact same pipeline. Nothing was broken in the conventional sense. The architecture simply did not scale with corpus size.
The second failure mode is less about retrieval mechanics and more about what happens when your document corpus is not a clean, single source of truth, which describes almost every real company. Policy documents get superseded but the old version stays in the drive. A pricing page from eighteen months ago sits next to this quarter's actual pricing. Two teams write two different runbooks for the same incident type, each convinced theirs is current.
A standard RAG pipeline has no concept of any of this. It retrieves by semantic similarity, not by validity, so a stale document that happens to phrase things clearly will often out-rank the current one, which might use different terminology. The model then confidently synthesises an answer from both, sometimes blending the old policy and the new one into something that never existed and that nobody would sign off on. This is, in our experience, the failure mode that does the most reputational damage, because the output reads as fluent and certain even when it is wrong.
In the same 100k-document deployment, we found that roughly one in six retrieved answers touching policy or pricing questions contained at least a partial contradiction with the source of truth once we audited them by hand. The pipeline had no signal that would have told anyone this was happening. It just kept answering.
The third failure mode shows up at the query end rather than the document end. Real users do not phrase questions the way documents are written. A support agent typing "why did the refund fail" is trying to match against a document titled "Payment Reconciliation and Chargeback Handling Procedure," and the embedding distance between those two phrasings is larger than most teams assume, especially for domain-specific terminology, internal product names, or abbreviations that only exist inside the company.
Dense embeddings are good at capturing broad semantic similarity, but they are noticeably weaker at exact term matching, which is precisely what you need when a user query contains a product SKU, an error code, or an internal acronym. Pure vector search will happily return five documents that are thematically related and miss the one that contains the literal string the user needed.
None of the fixes here are exotic. What matters is applying them together rather than picking one and hoping.
After these changes, the same held-out set of 200 questions scored 84 percent accuracy at 100,000 documents, above the original 10,000-document baseline, and the contradiction rate on policy and pricing questions dropped from roughly one in six to under one in twenty. Latency did increase, from an average of 900 milliseconds to about 1.4 seconds per query, because reranking is not free, and that was a trade-off the client was glad to make once they saw the accuracy numbers next to a support ticket volume that had been quietly rising because of bad answers.
Before we touch a line of retrieval code on a new engagement, we ask a client four questions, because the answers usually tell us which of the three failure modes we are dealing with before we have read a single log.
RAG has not stopped being useful. It has stopped being a solved problem you can bolt on once and ignore, which is a different thing, and a more interesting one to build for. If you are weighing whether your own retrieval pipeline needs this kind of rework, that is exactly the kind of question our AI consultancy work is built around.
Before we start, please share a few details so we can follow up with you.
End this conversation? Your chat will be emailed to us.