The 10-Million-Token Lie
On the belief that a bigger window will save you. “Why are you building all this? Windows are about to hit ten million tokens. Drop the whole company in and let the model sort it out.”
I hear a version of this every time a lab announces a longer context, which is to say constantly. It comes from smart people, and it deserves the honest part said first: scale has humiliated objections like mine before. Translation was supposed to need linguists. Code was supposed to need symbolic scaffolding. Both fell to capacity anyway, and long context specifically has already buried real infrastructure: the chunking strategies, the sliding windows, the seven-step retrieval choreography that existed only because windows were small. If you have bet against scale in this field, you have mostly lost. The people making this argument know that history, and about that history they are right.
They are still wrong about this, because this is not that kind of problem.
Your window holds your March pricing decision. It also holds the July reversal. Both are well written, both are confident, both are equally retrievable, and nothing in the tokens marks which one your company acts on today. Ask what the model should do there. Not what it can process. What it should believe.
Translation yielded to scale because translation is an inference problem: the answer is a function of the input. Given the sentence, the translation is computable from it. “Which of these decisions is in force” is not a function of the two documents. Two careful humans can read both memos in full and still disagree, because the answer was never in the memos. It was created in a meeting, by people with the standing to create it, and if nobody recorded the verdict, it does not exist in any corpus at any size. Reading is not deciding. You cannot infer your way to a fact that was never written down.
Notice that this is not a claim about capability, which is why capability cannot answer it. Everything the long-context optimists predict will come true. Windows will get bigger. Recall will get sharper. I will not even lean on the degradation research, because degradation is an engineering problem and engineering problems get solved. Grant perfect recall, at any length, forever.
Perfect recall makes it worse.
Today the model misses things in long context, and the missing is visible, so your team stays skeptical and a human checks the output. Perfect recall removes the visible failure and keeps the invisible one: the model quotes both memos flawlessly and still guesses which is alive, except now it guesses with citations. Capability converts the failures you can see into failures you cannot. Your trust rises while the gap stays exactly where it was.
And the gap is structural. A company that learns produces contradictions as exhaust, because deciding means reversing earlier decisions; the corpus fills with March-July pairs on pricing, on architecture, on hiring, none of them marked. A claim that was true when written stays word-for-word identical on the day reality makes it false. The window has no operation for any of this. Belief revision is a write: retire the old claim, name its successor. A context window can only append. It is a log, and no log, at any length, has ever issued a verdict.
There is also a quieter cost that never makes the launch post. When context was scarce, somebody had to choose what went into the window, and the choosing was a crude, accidental form of curation. When everything fits, nobody chooses anything. The pile grows fastest at the exact moment the last pressure to sort it disappears. Storage went through this a generation ago, and filing never came back.
None of this argues against long context. Buy the bigger window; I do, gladly, every upgrade. It is a superb workbench: the whole codebase on the table, the full transcript in reach, everything one task needs in one place, and my company runs on that workbench daily. But a workbench is a place to work, not a place where truth gets decided, and adding table space does not change what kind of place it is.
So here is the wager, in a form you can hold me to. The day a model, handed the March memo and the July memo and nothing else, reliably tells you which is in force — reversals included, scoped exceptions included, with reasoning a human can audit and overrule — I retire this essay in public and say so in this series. Until then, every quarter will bring a bigger window, the chorus will sing that the knowledge problem is about to dissolve, and somewhere an agent with flawless recall will quote a dead decision to a customer. I will take that bet every quarter.