← Bridging the gap
AI

The 10-Million-Token Lie

Aug 2026 · 4 min

On the belief that a bigger window will save you. “Why are you building all this? Windows are about to hit ten million tokens. Drop the whole company in and let the model sort it out.”

I hear a version of this every time a lab announces a longer context, which is to say constantly. It comes from smart people, and it deserves the honest part said first: scale has humiliated objections like mine before. Translation was supposed to need linguists. Code was supposed to need symbolic scaffolding. Both fell to capacity anyway, and long context specifically has already buried real infrastructure: the chunking strategies, the sliding windows, the seven-step retrieval choreography that existed only because windows were small. If you have bet against scale in this field, you have mostly lost. The people making this argument know that history, and about that history they are right.

They are still wrong about this, because this is not that kind of problem.

Your window holds your March pricing decision. It also holds the July reversal. Both are well written, both are confident, both are equally retrievable, and nothing in the tokens marks which one your company acts on today. Ask what the model should do there. Not what it can process. What it should believe.

Translation yielded to scale because translation is an inference problem: the answer is a function of the input. Given the sentence, the translation is computable from it. “Which of these decisions is in force” is not a function of the two documents. Two careful humans can read both memos in full and still disagree, because the answer was never in the memos. It was created in a meeting, by people with the standing to create it, and if nobody recorded the verdict, it does not exist in any corpus at any size. Reading is not deciding. You cannot infer your way to a fact that was never written down.

An answer computed from its input, beside one that arrives from outsideOn the left, a circle joined by an arrow to a second circle that holds a solid dot: the answer is contained in what was given. On the right, two identical circles, and a marked line descending from beyond the top edge of the frame to settle a dot into one of them. Nothing in the pair itself chooses between them.INFERENCESTANDING
Fig. 01A translation is computable from its input. Which decision is in force is not — the verdict enters from outside the corpus, or it does not exist at all.

Notice that this is not a claim about capability, which is why capability cannot answer it. Everything the long-context optimists predict will come true. Windows will get bigger. Recall will get sharper. I will not even lean on the degradation research, because degradation is an engineering problem and engineering problems get solved. Grant perfect recall, at any length, forever.

Perfect recall makes it worse.

Today the model misses things in long context, and the missing is visible, so your team stays skeptical and a human checks the output. Perfect recall removes the visible failure and keeps the invisible one: the model quotes both memos flawlessly and still guesses which is alive, except now it guesses with citations. Capability converts the failures you can see into failures you cannot. Your trust rises while the gap stays exactly where it was.

Visible misses in one row, a uniform row beneath itAn upper row of dots in which three are drawn as open marked rings — the failures you can see. A lower row of the same length in which every dot is solid and identical. One of them is wrong, and nothing in the row says which.TODAYPERFECT RECALL
Fig. 02While the misses show, somebody checks. Perfect recall makes every answer look alike — including the one quoting a decision that died in July.

And the gap is structural. A company that learns produces contradictions as exhaust, because deciding means reversing earlier decisions; the corpus fills with March-July pairs on pricing, on architecture, on hiring, none of them marked. A claim that was true when written stays word-for-word identical on the day reality makes it false. The window has no operation for any of this. Belief revision is a write: retire the old claim, name its successor. A context window can only append. It is a log, and no log, at any length, has ever issued a verdict.

An append-only stack beside a stack with a retirementTwo identical stacks of six bars. In the left stack every bar carries the same weight and nothing relates them. In the right stack one bar is drawn faintly and a marked line carries from it down to a later bar, naming that bar as its successor.APPENDWRITE
Fig. 03Appending is the only move a window has. Revision is a write — the old claim retired, its successor named — and no amount of length adds that operation.

There is also a quieter cost that never makes the launch post. When context was scarce, somebody had to choose what went into the window, and the choosing was a crude, accidental form of curation. When everything fits, nobody chooses anything. The pile grows fastest at the exact moment the last pressure to sort it disappears. Storage went through this a generation ago, and filing never came back.

None of this argues against long context. Buy the bigger window; I do, gladly, every upgrade. It is a superb workbench: the whole codebase on the table, the full transcript in reach, everything one task needs in one place, and my company runs on that workbench daily. But a workbench is a place to work, not a place where truth gets decided, and adding table space does not change what kind of place it is.

So here is the wager, in a form you can hold me to. The day a model, handed the March memo and the July memo and nothing else, reliably tells you which is in force — reversals included, scoped exceptions included, with reasoning a human can audit and overrule — I retire this essay in public and say so in this series. Until then, every quarter will bring a bigger window, the chorus will sing that the knowledge problem is about to dissolve, and somewhere an agent with flawless recall will quote a dead decision to a customer. I will take that bet every quarter.

More from Bridging the gap, soon. Back to the index

Say hello.

Always up for a digital coffee — building products, what AI is doing to design, and how teams keep what they know.