Your Context Window Did Not Kill RAG. It Just Moved the Failure Mode.

Bigger context windows were supposed to make RAG unnecessary. The research says models read long inputs unevenly, length alone can cost accuracy, and routing between the two can keep quality at lower cost.

Subject
RAG (retrieval-augmented generation) against long context windows: what the research says, and when to route between them
Published
19 AUG 2026
Reading time
9 min

In short

The claim
Huge model inputs were meant to make RAG unnecessary. Models still read long inputs unevenly.
What the research found
Accuracy sags mid-input, falls on harder tests well before the rated length, and often drops with length alone.
Where long input wins
In one comparison, with cost no object, it scored higher than RAG on average. Trying RAG first kept quality comparable at lower cost.
What to do
A small collection that fits: try putting it all in. Large or mixed: retrieve. In between: route.
In this post · 7 sections
  1. —In short
  2. 01Accuracy sags in the middle of a long input
  3. 02A model’s rated length is more than the length it handles well
  4. 03Length costs accuracy even when retrieval is perfect
  5. 04Long context can outperform retrieval
  6. 05Routing between the two keeps most of the quality
  7. 06What this looks like inside a compliance boundary
  8. 07A decision rule
The same argument in 4:36. Our graphics, AI narration.

Retrieval-augmented generation (RAG), fetching the few passages a question needs before a model answers, was supposed to die once context windows grew. A context window is everything a model reads at once, measured in tokens, words or pieces of words.

Meta rated Llama 4 Scout at 10 million tokens in April 2025, up from 128,000 for Llama 3.1. The pitch that follows is simple: skip the search step, put every document in the prompt, and let the model sort it out. The research below shows where that costs accuracy, and where it doesn’t.

Meta, “The Llama 4 herd” (April 2025) and “Introducing Llama 3.1” (July 2024). ai.meta.com/blog/llama-4-multimodal-intelligence · ai.meta.com/blog/meta-llama-3-1

Two ways to give a model your documents. One: put every document into one long input; the model reads all of it and answers. Two, retrieval-augmented generation (RAG): a question goes to a search step that picks a few passages; the model reads only those, and the passages it was given are logged with the answer.TWO WAYS INHOW A MODEL GETS YOUR DOCUMENTSLONG CONTEXT, OR RETRIEVE FIRSTDWG Nº 36SCHEMATIC1 · PUT EVERYTHING INlong context2 · RETRIEVE FIRST (RAG)retrieval-augmented generationDOCUMENTSall of themLONG INPUTevery pageMODELreads it allANSWERfrom the pileQUESTIONone queryRETRIEVALa few passagesMODELreads thoseCITED ANSWERpassages loggedLONG CONTEXTsimplest to buildinput and bill grow with the pilethe log is the whole pileRAGa search step to build and tuneshort input, chosen by youthe log is a few passages
Two ways in: pour every document into one long input, or retrieve the few passages a question needs first. Both can log what the model was given; with retrieval, that log is short enough to read.

Accuracy sags in the middle of a long input

In 2023, a Stanford-led team placed the one passage that answers a question at controlled positions inside a long input. Accuracy was highest with the passage at the very start or the very end, and dropped significantly when it sat in the middle. The curve held even for models built for long inputs. With 20 or 30 documents in the input, GPT-3.5-Turbo’s worst case was below its 56.1% score with no documents at all. The paper named the effect: Lost in the Middle.

“Lost in the Middle: How Language Models Use Long Contexts,” arXiv:2307.03172 (2023; TACL vol. 12, 2024). arxiv.org/abs/2307.03172

Two separate failure modes: answer accuracy dipping in the middle of a context window, and quality falling once too much retrieved material is poured in1 · WHERE THE ANSWER SITSposition: lost in the middle, 2023ONE LONG CONTEXT WINDOWSTARTMIDDLEENDX — WHERE IN THE WINDOW IT LANDSACCURACYTHE SAGFACT✓ FOUND✕ MISSED✓ FOUND2 · HOW MUCH YOU POUR INdilution: a different cliffQUALITYX — HOW MANY CHUNKS YOU POUR INTHE CLIFFmore in, worse outCHUNKS RETRIEVED — THE LAST FEW ARE NOISERETRIEVAL: PUT IT WHERE IT IS READRETRIEVAL: STOP BEFORE IT DILUTESTWO DIFFERENT FAILURESNOT ONE CLIFF THAT MOVEDRETRIEVAL CHOOSES BOTHSCHEMATIC OF PUBLISHED FINDINGS — SHAPES ARE QUALITATIVE, NOT MEASURED VALUES
Same window, two outcomes: in these experiments accuracy was higher with the evidence near the start or end of the input, and too much retrieved text pulled quality down again. Retrieval answers both by choosing what goes in, where it sits, and how much of it there is.

The second cliff comes from the other direction, from pouring in too much retrieved text. A 2026 study of chunking strategies found answer quality falling beyond about 2,500 tokens of retrieved context, using SPLADE retrieval and a Ministral-8B generator on the Natural Questions set. It called this a “context cliff”, and notes the exact point depends on the model. A 2024 study saw the same shape as retrieved chunks were added: quality rose, peaked, then fell. A third, in 2026, found that many newer ways of cutting documents into chunks claim gains on narrow cases, with little evidence that they hold elsewhere.

“A Systematic Analysis of Chunking Strategies for Reliable Question Answering,” arXiv:2601.14123 (2026). arxiv.org/abs/2601.14123 · “In Defense of RAG in the Era of Long-Context Language Models,” arXiv:2409.01666 (2024). arxiv.org/abs/2409.01666 · “Chunking Methods on Retrieval-Augmented Generation — Effectiveness Evaluation Against Computational Cost and Limitations,” arXiv:2606.00881 (2026). arxiv.org/abs/2606.00881

A model’s rated length is more than the length it handles well

The usual check hides one fact in a long text and asks for it back. Models pass it almost perfectly, and it flatters them. The RULER benchmark (2024) added harder tasks and tested 17 models, all rated at 32,000 tokens or more. Only half held up at 32,000. Its pass mark is a small model’s score on short inputs, and GPT-4, rated at 128,000 tokens, cleared it only up to 64,000.

NoLiMa (2025) removed a shortcut: its questions share few words with the hidden fact, so the model has to make the connection itself. Of 13 models rated at 128,000 tokens or more, 11 fell below half their short-input score at 32,000. GPT-4o went from 99.3% to 69.7%. Chroma, a vector-database company, tested 18 models in July 2025 and reported the same pattern: the longer the input, the less reliable the answer.

“RULER: What’s the Real Context Size of Your Long-Context Language Models?” arXiv:2404.06654 (2024). arxiv.org/abs/2404.06654 · “NoLiMa: Long-Context Evaluation Beyond Literal Matching,” arXiv:2502.05167 (2025). arxiv.org/abs/2502.05167 · Chroma, “Context Rot,” vendor technical report (July 2025). trychroma.com/research/context-rot

The rated length and the tested one. In the RULER benchmark (2024), GPT-4 was rated at 128,000 tokens and passed the benchmark's bar up to 64,000; of 17 models all rated at 32,000 tokens or more, only half held up at 32,000. In the NoLiMa benchmark (2025), GPT-4o scored 99.3% with under 1,000 tokens of input and 69.7% at 32,000; 11 of 13 models fell below half their short-input score at that length.RATED LIMITWHAT A MODEL ACCEPTS, WHAT IT USESTWO PUBLISHED BENCHMARKSDWG Nº 372024 · 2025RULER, 2024 · GPT-4length in tokens, one scaleRATEDPASSES TO128K64KThe bar: a small model’s score on short inputs. 17 models, all rated at32K or more: only half held up at 32K.NOLIMA, 2025 · GPT-4osame task, accuracy, one scale to 100%UNDER 1KAT 32K99.3%69.7%13 models rated at 128K or more: at 32K, 11 fell below halftheir short-input score. K = thousand tokens.
Rated is not tested: GPT-4, rated at 128K tokens, passed RULER’s bar only up to 64K; GPT-4o fell from 99.3% to 69.7% in NoLiMa once the input reached 32K tokens.

Length costs accuracy even when retrieval is perfect

A model that misses a passage might simply fail to find it. A 2025 study ruled that out. It gave five models the evidence they needed, on maths, question answering and code, and only made the input longer: up to 30,000 added tokens, inside every model’s rated length. In most combinations of model and task, accuracy fell, though a few held steady. Llama 3.1 8B’s score on a summing task went from 96% to 11% with 30,000 tokens of unrelated text between the evidence and the question.

Accuracy also fell when the added text was blank space. On two open models, it fell when the added tokens were masked out of the model’s attention entirely. It fell when the evidence sat right before the question, at the end of a long blank input. Length on its own can cost accuracy, so part of what retrieval buys is a shorter input.

“Context Length Alone Hurts LLM Performance Despite Perfect Retrieval,” arXiv:2510.05381 (2025). arxiv.org/abs/2510.05381

Length alone can cost accuracy. Five inputs that each contain the evidence and the question: a short one; three with extra material between the evidence and the question, as unrelated text, as blank space, and as tokens masked out of the model's attention (tested on two open models); and one with blank space first and the evidence moved next to the question. In most model and task combinations, accuracy fell as the input grew. Llama 3.1 8B's score on a summing task went from 96% to 11% with 30,000 tokens of unrelated text.LENGTH ONLYSAME EVIDENCE, LONGER INPUTRETRIEVAL WAS PERFECTDWG Nº 38SCHEMATICINPUTWHAT THE MODEL IS GIVENACCURACYSHORT INPUTevidence and question onlyevidencequestionBASELINE+ UNRELATED TEXTbetween the twoevidenceunrelated textquestion▼ FALLS+ BLANK SPACEbetween the twoevidenceblank spacequestion▼ FALLS+ MASKED OUTtwo open models onlyevidencemasked outquestion▼ FALLSEVIDENCE MOVEDnext to the questionblank spaceevidencequestion▼ FALLSLLAMA 3.1 8B, A SUMMING TASK: 96% → 11%with 30,000 tokens of unrelated text. 5 models; maths, question answering,code. Most combinations fell, a few held, all inside the rated length.
Length alone: given the evidence they needed, most models lost accuracy as the input grew, even when the added text was blank or masked out, or the evidence sat right next to the question. One model’s summing task fell from 96% to 11%.

Long context can outperform retrieval

A 2024 comparison across public benchmarks found that, given enough resources, long context beat RAG on average. RAG cost far less, and for 63% of questions the two gave exactly the same prediction, right or wrong. A re-evaluation later that year found long context generally ahead of chunk-based retrieval on Wikipedia-style questions, with retrieval of summaries roughly level. Retrieval did better on dialogue and general queries.

Anthropic’s own guidance, published in September 2024, is that a knowledge base under 200,000 tokens, about 500 pages, can simply go into the prompt. And the two combine: a 2023 study found retrieval improved models whatever their window size, with the best results from using both.

“Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach,” arXiv:2407.16833 (2024). arxiv.org/abs/2407.16833 · “Long Context vs. RAG for LLMs: An Evaluation and Revisits,” arXiv:2501.01880 (2024). arxiv.org/abs/2501.01880 · Anthropic, “Introducing Contextual Retrieval” (September 2024). anthropic.com/news/contextual-retrieval · “Retrieval meets Long Context Large Language Models,” arXiv:2310.03025 (2023). arxiv.org/abs/2310.03025

Routing between the two keeps most of the quality

The same authors then tested a retrieval-first router. Each question goes to retrieval first, and the model is told to reply “unanswerable” if the passages can’t answer it. Only those questions go to long context with the whole document. The authors report quality comparable to long context alone, at 65% lower cost on Gemini-1.5-Pro and 39% lower on GPT-4o, counted in tokens processed.

A cheap test before you trust either one: take real questions with known answers. Move the supporting passage between the start, middle and end of the input you’d normally send, keeping everything else fixed. Separately, pad the input with unrelated text. Run each version a few times and compare how often the answer is right. A repeatable drop is your own cliff.
Routing between retrieval and long context, as in the Self-Route study (2024). A question goes to retrieval first, and the model is asked whether the retrieved passages answer it. If yes, it answers from them, cheaply. If no, the whole document goes to a long-context model, which answers. The authors report quality comparable to long context alone, at 65% lower cost on Gemini-1.5-Pro and 39% lower on GPT-4o, counted in tokens processed; for 63% of questions the two approaches gave exactly the same prediction, right or wrong.ROUTINGRETRIEVE FIRST, ESCALATE IF NEEDEDSELF-ROUTE, A 2024 STUDYDWG Nº 39SCHEMATICyesnoQUESTIONone queryTHE ROUTERRETRIEVE FIRSTa few passages, thenask the model: canthese answer it?ANSWERfrom the passages, cheapLONG CONTEXTwhole documentANSWERWHAT ROUTING SAVEDquality comparable to long context alonereported cost: 65% lower on Gemini-1.5-Pro,39% lower on GPT-4o, in tokens processedsame prediction on 63% of questions, right or wrong
Route between them: retrieve first, and send a question to long context only when the model says the passages can’t answer it. The Self-Route authors report quality close to long context alone, at 65% lower cost on one model and 39% on another.

What this looks like inside a compliance boundary

One of our own production systems is built the retrieval way. It runs meaning-based and keyword search in parallel across more than 24,000 documents and 1,100-plus vector collections, entirely inside a compliance boundary. Nothing leaves the boundary but a cited answer.

The design came from the question an auditor asks: which passages was the model given, and why those? Any system should log its inputs and check its citations. With retrieval, that log is a short list of passages a person can read. With one long input, it is the whole pile.

A decision rule

If the collection is small and fits the model’s window, try putting it all in; Anthropic’s guidance puts that line at about 500 pages. For large or mixed collections and high-stakes precision, retrieval keeps the input short and the evidence easy to check. Long inputs fail quietly: no error, just a wrong answer in the same confident tone as a right one. In between, route: retrieve first, and escalate to long context only when the passages can’t answer. Run the cheap test on your own model and data first.

Updated 4 October 2026: added research on rated versus tested length, on length alone, and on routing; corrected the second drawing’s caption to include input length alongside placement.

AGNIZAR
Production AI · System architecture · Fractional CTO

Agnizar builds AI into your core systems, then hands it over or keeps it running. Every job starts small: one bounded piece of work, one named result, one clear decision. Book an AI Architecture Review; a senior engineer replies within one business day.

Book an AI Architecture Review