Your Context Window Did Not Kill RAG. It Just Moved the Failure Mode.
Bigger context windows were supposed to make RAG unnecessary. The research says models read long inputs unevenly, length alone can cost accuracy, and routing between the two can keep quality at lower cost.
- Subject
- RAG (retrieval-augmented generation) against long context windows: what the research says, and when to route between them
- Published
- 19 AUG 2026
- Reading time
- 9 min
- Film
- Watch · 4:36
In short
- The claim
- Huge model inputs were meant to make RAG unnecessary. Models still read long inputs unevenly.
- What the research found
- Accuracy sags mid-input, falls on harder tests well before the rated length, and often drops with length alone.
- Where long input wins
- In one comparison, with cost no object, it scored higher than RAG on average. Trying RAG first kept quality comparable at lower cost.
- What to do
- A small collection that fits: try putting it all in. Large or mixed: retrieve. In between: route.
In this post · 7 sections
- —In short
- 01Accuracy sags in the middle of a long input
- 02A model’s rated length is more than the length it handles well
- 03Length costs accuracy even when retrieval is perfect
- 04Long context can outperform retrieval
- 05Routing between the two keeps most of the quality
- 06What this looks like inside a compliance boundary
- 07A decision rule
Retrieval-augmented generation (RAG), fetching the few passages a question needs before a model answers, was supposed to die once context windows grew. A context window is everything a model reads at once, measured in tokens, words or pieces of words.
Meta rated Llama 4 Scout at 10 million tokens in April 2025, up from 128,000 for Llama 3.1. The pitch that follows is simple: skip the search step, put every document in the prompt, and let the model sort it out. The research below shows where that costs accuracy, and where it doesn’t.
Meta, “The Llama 4 herd” (April 2025) and “Introducing Llama 3.1” (July 2024). ai.meta.com/blog/llama-4-multimodal-intelligence · ai.meta.com/blog/meta-llama-3-1
Accuracy sags in the middle of a long input
In 2023, a Stanford-led team placed the one passage that answers a question at controlled positions inside a long input. Accuracy was highest with the passage at the very start or the very end, and dropped significantly when it sat in the middle. The curve held even for models built for long inputs. With 20 or 30 documents in the input, GPT-3.5-Turbo’s worst case was below its 56.1% score with no documents at all. The paper named the effect: Lost in the Middle.
“Lost in the Middle: How Language Models Use Long Contexts,” arXiv:2307.03172 (2023; TACL vol. 12, 2024). arxiv.org/abs/2307.03172
The second cliff comes from the other direction, from pouring in too much retrieved text. A 2026 study of chunking strategies found answer quality falling beyond about 2,500 tokens of retrieved context, using SPLADE retrieval and a Ministral-8B generator on the Natural Questions set. It called this a “context cliff”, and notes the exact point depends on the model. A 2024 study saw the same shape as retrieved chunks were added: quality rose, peaked, then fell. A third, in 2026, found that many newer ways of cutting documents into chunks claim gains on narrow cases, with little evidence that they hold elsewhere.
“A Systematic Analysis of Chunking Strategies for Reliable Question Answering,” arXiv:2601.14123 (2026). arxiv.org/abs/2601.14123 · “In Defense of RAG in the Era of Long-Context Language Models,” arXiv:2409.01666 (2024). arxiv.org/abs/2409.01666 · “Chunking Methods on Retrieval-Augmented Generation — Effectiveness Evaluation Against Computational Cost and Limitations,” arXiv:2606.00881 (2026). arxiv.org/abs/2606.00881
A model’s rated length is more than the length it handles well
The usual check hides one fact in a long text and asks for it back. Models pass it almost perfectly, and it flatters them. The RULER benchmark (2024) added harder tasks and tested 17 models, all rated at 32,000 tokens or more. Only half held up at 32,000. Its pass mark is a small model’s score on short inputs, and GPT-4, rated at 128,000 tokens, cleared it only up to 64,000.
NoLiMa (2025) removed a shortcut: its questions share few words with the hidden fact, so the model has to make the connection itself. Of 13 models rated at 128,000 tokens or more, 11 fell below half their short-input score at 32,000. GPT-4o went from 99.3% to 69.7%. Chroma, a vector-database company, tested 18 models in July 2025 and reported the same pattern: the longer the input, the less reliable the answer.
“RULER: What’s the Real Context Size of Your Long-Context Language Models?” arXiv:2404.06654 (2024). arxiv.org/abs/2404.06654 · “NoLiMa: Long-Context Evaluation Beyond Literal Matching,” arXiv:2502.05167 (2025). arxiv.org/abs/2502.05167 · Chroma, “Context Rot,” vendor technical report (July 2025). trychroma.com/research/context-rot
Length costs accuracy even when retrieval is perfect
A model that misses a passage might simply fail to find it. A 2025 study ruled that out. It gave five models the evidence they needed, on maths, question answering and code, and only made the input longer: up to 30,000 added tokens, inside every model’s rated length. In most combinations of model and task, accuracy fell, though a few held steady. Llama 3.1 8B’s score on a summing task went from 96% to 11% with 30,000 tokens of unrelated text between the evidence and the question.
Accuracy also fell when the added text was blank space. On two open models, it fell when the added tokens were masked out of the model’s attention entirely. It fell when the evidence sat right before the question, at the end of a long blank input. Length on its own can cost accuracy, so part of what retrieval buys is a shorter input.
“Context Length Alone Hurts LLM Performance Despite Perfect Retrieval,” arXiv:2510.05381 (2025). arxiv.org/abs/2510.05381
Long context can outperform retrieval
A 2024 comparison across public benchmarks found that, given enough resources, long context beat RAG on average. RAG cost far less, and for 63% of questions the two gave exactly the same prediction, right or wrong. A re-evaluation later that year found long context generally ahead of chunk-based retrieval on Wikipedia-style questions, with retrieval of summaries roughly level. Retrieval did better on dialogue and general queries.
Anthropic’s own guidance, published in September 2024, is that a knowledge base under 200,000 tokens, about 500 pages, can simply go into the prompt. And the two combine: a 2023 study found retrieval improved models whatever their window size, with the best results from using both.
“Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach,” arXiv:2407.16833 (2024). arxiv.org/abs/2407.16833 · “Long Context vs. RAG for LLMs: An Evaluation and Revisits,” arXiv:2501.01880 (2024). arxiv.org/abs/2501.01880 · Anthropic, “Introducing Contextual Retrieval” (September 2024). anthropic.com/news/contextual-retrieval · “Retrieval meets Long Context Large Language Models,” arXiv:2310.03025 (2023). arxiv.org/abs/2310.03025
Routing between the two keeps most of the quality
The same authors then tested a retrieval-first router. Each question goes to retrieval first, and the model is told to reply “unanswerable” if the passages can’t answer it. Only those questions go to long context with the whole document. The authors report quality comparable to long context alone, at 65% lower cost on Gemini-1.5-Pro and 39% lower on GPT-4o, counted in tokens processed.
What this looks like inside a compliance boundary
One of our own production systems is built the retrieval way. It runs meaning-based and keyword search in parallel across more than 24,000 documents and 1,100-plus vector collections, entirely inside a compliance boundary. Nothing leaves the boundary but a cited answer.
The design came from the question an auditor asks: which passages was the model given, and why those? Any system should log its inputs and check its citations. With retrieval, that log is a short list of passages a person can read. With one long input, it is the whole pile.
A decision rule
If the collection is small and fits the model’s window, try putting it all in; Anthropic’s guidance puts that line at about 500 pages. For large or mixed collections and high-stakes precision, retrieval keeps the input short and the evidence easy to check. Long inputs fail quietly: no error, just a wrong answer in the same confident tone as a right one. In between, route: retrieve first, and escalate to long context only when the passages can’t answer. Run the cheap test on your own model and data first.
Updated 4 October 2026: added research on rated versus tested length, on length alone, and on routing; corrected the second drawing’s caption to include input length alongside placement.
Agnizar builds AI into your core systems, then hands it over or keeps it running. Every job starts small: one bounded piece of work, one named result, one clear decision. Book an AI Architecture Review; a senior engineer replies within one business day.