Services Case Studies Insights About Start a project →

When RAG stops being the answer.

AI Published June 24, 2026 8 min read

Three failure modes we see repeatedly in production retrieval systems, and the architectural shifts that fix them. With real numbers from a 100k-document deployment.

The default that stopped scaling

Retrieval-augmented generation became the default architecture for grounding a language model in a company's own documents somewhere around 2023, and for a lot of use cases it is still the right call. Chunk the documents, embed them, store the vectors, retrieve the nearest neighbours for a query, stuff them into a prompt, and let the model write the answer. It works well in a demo, and it works well for the first few thousand documents in production too. The trouble starts later, once the corpus grows, once documents start contradicting each other, and once real users start typing real questions instead of the tidy ones from the test set.

We have now rebuilt or repaired enough of these systems that the failure modes have stopped surprising us. They are structural, not accidental, and they show up in almost the same order every time. This post covers the three we see most often, what we changed to fix them, and roughly what that was worth, drawn from a recent engagement where the client's internal knowledge base had grown past 100,000 documents and the existing RAG pipeline had quietly become more of a liability than a feature.

Failure mode one: the retrieval ceiling

Most RAG pipelines are built around naive fixed-size chunking, usually somewhere between 300 and 800 tokens, sometimes with a small overlap. This is a reasonable place to start and a bad place to stay. Fixed-size chunks slice through section boundaries, split tables in half, and separate a clause from the heading that gives it meaning. The embedding model then has to represent a fragment that, on its own, barely means anything.

The second half of this problem is what researchers call "lost in the middle." Once you retrieve, say, the top eight chunks and concatenate them into a context window, the model pays disproportionate attention to whatever is at the start and end of that context, and it quietly deprioritises the middle. So even when the correct chunk is technically retrieved, it can still lose the argument to a chunk sitting in a better position in the prompt.

Together, these two effects create what we call a retrieval ceiling: past a certain corpus size, adding more documents does not improve answer quality, it degrades it, because the retriever is returning more plausible-looking noise for every genuinely relevant chunk. In the deployment we mentioned above, accuracy on a held-out set of 200 real support questions had dropped from 81 percent at 10,000 documents to 58 percent at 100,000 documents, using the exact same pipeline. Nothing was broken in the conventional sense. The architecture simply did not scale with corpus size.

Failure mode two: contradicting and stale sources

The second failure mode is less about retrieval mechanics and more about what happens when your document corpus is not a clean, single source of truth, which describes almost every real company. Policy documents get superseded but the old version stays in the drive. A pricing page from eighteen months ago sits next to this quarter's actual pricing. Two teams write two different runbooks for the same incident type, each convinced theirs is current.

A standard RAG pipeline has no concept of any of this. It retrieves by semantic similarity, not by validity, so a stale document that happens to phrase things clearly will often out-rank the current one, which might use different terminology. The model then confidently synthesises an answer from both, sometimes blending the old policy and the new one into something that never existed and that nobody would sign off on. This is, in our experience, the failure mode that does the most reputational damage, because the output reads as fluent and certain even when it is wrong.

In the same 100k-document deployment, we found that roughly one in six retrieved answers touching policy or pricing questions contained at least a partial contradiction with the source of truth once we audited them by hand. The pipeline had no signal that would have told anyone this was happening. It just kept answering.

Failure mode three: the vocabulary gap

The third failure mode shows up at the query end rather than the document end. Real users do not phrase questions the way documents are written. A support agent typing "why did the refund fail" is trying to match against a document titled "Payment Reconciliation and Chargeback Handling Procedure," and the embedding distance between those two phrasings is larger than most teams assume, especially for domain-specific terminology, internal product names, or abbreviations that only exist inside the company.

Dense embeddings are good at capturing broad semantic similarity, but they are noticeably weaker at exact term matching, which is precisely what you need when a user query contains a product SKU, an error code, or an internal acronym. Pure vector search will happily return five documents that are thematically related and miss the one that contains the literal string the user needed.

What we changed, and what it did

None of the fixes here are exotic. What matters is applying them together rather than picking one and hoping.

  • Structure-aware chunking We rebuilt the ingestion pipeline to chunk along document structure, sections, tables, and headings, rather than fixed token counts, and we attach the parent heading to every chunk as metadata so a fragment never loses its context.
  • Hybrid search with reranking We pair dense vector retrieval with BM25 keyword search, then pass the combined candidate set through a cross-encoder reranker before anything reaches the prompt. This directly closes the vocabulary gap, because BM25 still catches the literal error code or SKU that the embedding model glosses over.
  • Recency and validity metadata Every document now carries an effective date and a superseded-by pointer where one exists, and the retriever filters or down-weights anything marked stale before reranking, not after generation.
  • Query rewriting A cheap, fast model rewrites the incoming query into two or three retrieval-oriented variants before search runs, which measurably helps with the short, informal queries real users actually type.
  • Grounded citation and refusal The generation step is required to cite the specific chunk it drew from for any factual claim, and it is explicitly instructed to say it does not know rather than blend conflicting sources. This alone caught a large share of the contradiction cases we found in the audit.

After these changes, the same held-out set of 200 questions scored 84 percent accuracy at 100,000 documents, above the original 10,000-document baseline, and the contradiction rate on policy and pricing questions dropped from roughly one in six to under one in twenty. Latency did increase, from an average of 900 milliseconds to about 1.4 seconds per query, because reranking is not free, and that was a trade-off the client was glad to make once they saw the accuracy numbers next to a support ticket volume that had been quietly rising because of bad answers.

The checklist we now run first

Before we touch a line of retrieval code on a new engagement, we ask a client four questions, because the answers usually tell us which of the three failure modes we are dealing with before we have read a single log.

  • How big is the corpus, and how fast is it growing Retrieval ceilings show up predictably around the 30,000 to 50,000 document mark for most naive pipelines, sometimes earlier if documents are long.
  • Is there more than one version of the truth in the corpus If two teams maintain overlapping documentation, assume contradiction risk exists whether or not anyone has noticed it yet.
  • Who is asking the questions, and how do they phrase them Internal technical staff and external customers use different vocabulary, and a pipeline tuned for one will underperform for the other.
  • What does the system do when it is not confident If the honest answer is "it always answers," that is the single highest-leverage fix available, often before any retrieval architecture work at all.

RAG has not stopped being useful. It has stopped being a solved problem you can bolt on once and ignore, which is a different thing, and a more interesting one to build for. If you are weighing whether your own retrieval pipeline needs this kind of rework, that is exactly the kind of question our AI consultancy work is built around.

Keep reading

Debugging a RAG system that used to work?

Start a conversation →
KT Solutions Assistant

Before we start, please share a few details so we can follow up with you.

Please enter your name and a valid email address.

End this conversation? Your chat will be emailed to us.