Why AI Misreads Long Documents, and What Actually Fixes It
Published October 1, 2026
A pattern repeats across models: accuracy peaks when the relevant information sits at the start or end of a document, and collapses in the middle.

Every time an AI model answers a long document question wrongly, one explanation gets repeated most often: the context window is not big enough. It is a reassuring answer, because it implies a technical fix is coming. The reality is that after context windows grew many times over, the errors did not disappear. They moved.
In 2023, a team from Stanford and Samaya AI tested language models on two tasks that require finding information inside long documents. They found a pattern consistent across every model they tested, including ones built specifically for long context: accuracy peaked when the relevant information sat at the beginning or the end of the input, and dropped sharply when it sat in the middle. Not by a few points, either. It dropped.
They called it "lost in the middle". The name is accurate, because what gets lost is not the data but the model's attention to it. The information is inside the context, the model can technically reach it, and the result is as if it were never there.
The shape of the curve is a U
The most useful finding in that paper is not a quote but a curve. If you move a correct document from the start to the middle to the end and measure accuracy at each position, you do not see a gradual decline. You see a letter U.
The numbers above represent the shape of the curve, not a single paper's exact measurement. What matters is the pattern: high at both ends, low in the middle, and roughly symmetric.
Those two ends have names: primacy bias for the early one and recency bias for the late one. Neither is an accident. A network trained on next-token prediction gets a clearer gradient when the target is nearby. The start of the context has patterns the model has already internalised before the question appears. The end of the context has the most recent tokens and sits closest to the question. The middle has neither advantage.
Compare with what happens in the middle. In one scenario, when the relevant information was placed in the middle of the input, GPT-3.5-Turbo scored lower than when it answered with no document at all. That is not a slightly wrong answer. That is worse than not reading the document.
Why this matters in practice, and what to add
Three direct implications for anyone using AI on work documents. All three reduce to one table, and the difference between them is not right versus wrong but which symptom you are actually seeing:
| What you observe | What is really happening | What to change |
|---|---|---|
| The model says the information is not in the document | The context window is cutting it off | Enlarge the window, or split the document into parts |
| The model is right when the key sits at the start or end, and wrong or making things up when it sits in the middle | Positional bias, the lost-in-the-middle effect | Reorder, and put the decisive parts at the edges |
| Answers shift around for no clear reason on long documents | Both problems above at once | Split the document, then reorder each piece |
Third, and this is the one most often skipped: explicitly deleting the middle beats adding window. If you already know a fact sits on page 200 of a 400-page document, splitting that document into two 200-page halves and handing over both means the fact lands at the edge of one of them rather than buried in the middle. Removing context you know is irrelevant is not waste. It is the cheapest accuracy improvement available.
The first is implications for anyone using AI on work documents.
First, document order changes answers. If you upload a contract, a report, or a technical spec, the parts that determine the outcome and that get asked about most are better placed near the start or near the end, not buried mid-document with appendices behind them. This is not theory. This is measured behaviour.
Second, "a bigger context window" is not an automatic fix. In 2025, the LongRoPE2 paper tried to extend LLaMA3-8B to 128,000 tokens while keeping more than 98.5% of its short-context performance, using only 10 billion training tokens. Those numbers are impressive. But the Lost in the Middle paper does not pretend that a longer window automatically fixes the U-shaped bias. Those are two different problems, and solving one does not subsidise the other.
The cause sits inside attention itself
A 2024 paper, "Found in the Middle" in Findings of ACL, traced the underlying cause: the positional bias is intrinsic to the model. Tokens at the beginning and end of the input receive more attention by default, regardless of whether their content is relevant.
The analogy is reading a long table under lights that only come on at each end. You can see the whole table, but anything placed in the middle shadows is harder to read. Making the table longer does not fix the lighting.
That paper proposed a neat fix: calibrate the positional bias so the model can attend according to relevance rather than position. It flattens the attention distribution. Tested on retrieval and retrieval-augmented generation, it beat existing methods by up to 10 percentage points.
Third, something tutorials rarely mention: splitting the document yourself often beats enlarging the window. If a 400-page document is broken into five 80-page chunks, with the correct order preserved in the prompt, each chunk enters the model intact, and the information you care about lands at the beginning or the end of each chunk rather than in the middle of the original. A large window makes all of it fit, but the content still sits in segments inside that window. Removing the middle explicitly gives you a benefit no larger window can provide, at no extra compute cost.
One detail worth noting: the same U-shaped bias appears in small models like GPT-2, not only in large, expensive ones. It is a property of the attention architecture itself, not an artifact of data scale. Which means it is not a temporary problem that better scale will solve.
Quadratic cost, and how FlashAttention removed it
There is a second problem that gets confused with this one. Attention computes every pair of words in a document, so the cost rises quadratically with length. At 1,000 words that is about a million pairwise computations. At 100,000 words, ten billion. At a million-token window, a quadrillion.
Several approaches tried to solve this by trading away quality: linear attention, sparse attention, Performer. The problem is that many of them reduced the computational work on paper without delivering real wall-clock speedup, because the cost just shifted elsewhere.
FlashAttention, from Stanford in 2022, attacked the problem from a different angle. Instead of simplifying attention, it recognised that the real bottleneck is not the number of operations but the volume of data moving between two levels of GPU memory: the large, slow main memory and the small on-chip cache. The algorithm uses tiling so attention is computed inside the cache and the large attention matrix never gets written to slow memory.
The results, from the paper's own table:
| Metric | Standard attention | FlashAttention |
|---|---|---|
| HBM reads/writes (GB) | 35.3 | 4.4 |
| Runtime (ms) | 35.1 | 11.7 |
| Memory complexity vs length | quadratic | linear |
The word "exact" in the paper's title matters. FlashAttention gives up nothing in attention precision. It produces exactly the same numbers as standard attention, just with far less memory movement. When FlashAttention-2 arrived in 2023 it roughly halved runtime again and reached 73% of theoretical maximum throughput on A100.
Two problems, two fixes, one rule
Back to practice. If your document is long and the model keeps getting it wrong, there are two different causes with two different fixes.
If the information is still being cut off. Symptoms: the model says the information is not in the document. Cause: the context window. Fix: enlarge the window, or split the document into smaller pieces.
If the information is there and being ignored. Symptoms: the model answers correctly when the information is at the start or end, and guesses or fabricates when it is in the middle. Cause: positional bias. Fix: reorder — put the decisive parts at the edges rather than trusting the model to find something buried mid-document.
Telling those apart matters, because the wrong fix makes things worse. Enlarging a 32,000-token window will not fix a 5,000-token document whose key clause is in paragraph eight. Restructuring that document, or chunking it, will.
The practical rule that falls out of all of this is one sentence: inside the context window, position is information. When you put a document into a model, you are not only conditioning it with the contents. You are also conditioning it with your ordering.
Eight years after the Transformer paper made attention the default, the problems that surfaced from standardising on it turn out not to be flaws in the core idea but things the architecture never handled. Positional bias and quadratic cost both follow directly from the decision to design attention so it could run in parallel. Both have also been solved by modifying that architecture rather than replacing it. And each solution changed how people should structure their input, which has not made it into anyone's documentation yet.
Every time the context window gets enlarged again, the same question returns: does a bigger window make the model better at the middle of a document, or does it only hide the problem for longer?
One measurable thing did change: the cost of the workaround. When retrieval-augmented pipelines first appeared, the accepted approach was to split documents, retrieve the top few passages, and send only those. Each extra retrieval round trip adds latency, cost, and a failure mode where the right passage never gets retrieved at all. As context windows grew, sending the whole document became possible, and that removed the retrieval step from a whole class of problems entirely. The tradeoff moved from "find the right passage" to "place it at the edge", which is a problem with an answer.
That is the part worth taking away. Both problems here are real and neither is exotic, but both have moved from unsolved to handled in a handful of years. What took nine years for the Transformer to be challenged at all took the field far longer to notice it had already solved things nobody had formally asked for.
Related tools
Free browser tools that apply to this topic.
Share this article
Share to
Related articles

October 1, 2026
The 2017 Transformer Architecture That Changed How Machines Learn
Eight P100 GPUs and 3.5 days were enough to beat the BLEU record. The paper did it by deleting the recurrence the field once considered mandatory.

September 30, 2026
When Did AI First Appear? Not 1956, but 13 Years Earlier
The name was coined at Dartmouth in 1956, but the idea was written down in 1943. AI's history shows a pattern that keeps repeating.

September 30, 2026
20 Countries Cap Social Media for Kids, Indonesia Moves First
Australia from December 2025, Indonesia from March 2026. At least 20 countries have similar rules, but only some are actually in force.



