Million-Token Context Windows: A Ceiling, Not a Budget
Uneven attention, rate limits below the advertised window, and linear cost on every call. Useful — and frequently misapplied.
Jump to 7 sections
Quick answer: A million-token context window is real, but comes with three caveats vendors rarely state together: attention quality degrades unevenly across the window, rate limits often cap usable context below what's advertised, and filling it on every call gets expensive fast. Treat the number as a ceiling, not a working budget.
Context windows grew roughly a thousandfold in four years, and the marketing has consistently run ahead of the measurement. Almost every lab now advertises a very large window. Almost none publishes what retrieval quality looks like across it.
This review examines the million-token context window as a shipped feature: what it enables, where it degrades, what it costs, and how to decide between a long context and a retrieval pipeline. Written for teams considering an architecture that depends on it.
What a million-token window actually promises
That the API will accept that many tokens in a request without erroring. That is the whole guarantee, and it is worth stating plainly because three other things get inferred from it that are not promised.
It does not promise uniform attention across the window. It does not promise that your account tier can send that many tokens per minute. And it certainly does not promise the request is affordable.
Anthropic's published pricing page states context "up to 1M, varies by model," which is a more careful formulation than most coverage of long context manages. The window is a per-model property, not a family-wide one, and the model you are actually calling may offer considerably less.
What it genuinely enables is real. Whole-repository code reasoning, long document analysis without chunking, and multi-hour agent transcripts that retain their own history all become possible rather than awkward. Those are meaningful capability unlocks, and they are why this feature matters despite the caveats.
Where long-context quality degrades
Unevenly, and in ways a single aggregate score hides. The widely-replicated finding across long-context research is that retrieval accuracy varies by position — information at the very start and very end of a long context is recovered more reliably than material buried in the middle.
The commonly cited test for this, a needle-in-a-haystack retrieval probe, is also the least representative of real use. Finding one distinctive planted sentence in a long document is far easier than synthesizing across twelve related passages scattered through it. A model can score near-perfectly on the first while struggling with the second.
The evaluation literature on long-context behavior, much of it published openly on arXiv's computation and language section, has repeatedly found this gap between retrieval probes and multi-hop reasoning across long inputs. Model Drop's position is that any long-context claim unaccompanied by a multi-hop evaluation should be treated as unmeasured.
The practical consequence: test recall at your own context length, with your own documents, using questions that require combining information from several places. A vendor's needle test tells you very little about whether your use case works.
What it costs to fill
Linearly, on every call, and that linearity is what surprises people. Context is not stored between requests unless you explicitly cache it — each call re-sends and re-pays for the entire window.
| Context sent | At $10/1M input | At $2/1M | 1,000 calls at $10/1M |
|---|---|---|---|
| 10,000 tokens | $0.10 | $0.02 | $100 |
| 100,000 tokens | $1.00 | $0.20 | $1,000 |
| 500,000 tokens | $5.00 | $1.00 | $5,000 |
| 1,000,000 tokens | $10.00 | $2.00 | $10,000 |
A thousand requests each filling a million-token window costs $10,000 in input tokens alone at flagship pricing. That is not a pathological example — it is a modest daily volume for a document-processing product.
Prompt caching is what makes this viable. A stable prefix cached across calls is billed at a steep discount on subsequent reads, which changes long-context economics from prohibitive to reasonable for workloads that query the same corpus repeatedly. The caching mechanics and the tier structure they sit inside are covered in our breakdown of LLM API pricing across the 2026 tiers.
The rate limit problem nobody mentions
Your account's tokens-per-minute limit is a separate number from the context window, published in a different place, and frequently lower than the window itself.
When that happens the advertised capability is theoretical. A million-token context paired with a lower per-minute throughput allowance means a single maximum-size request consumes your entire minute, and concurrency collapses to one. The model supports it; your account does not.
This is exactly the class of omission our guide to reading launch announcements flags: the headline number and the operational constraint are published separately, and only one of them appears in the announcement. Check both before designing around a long context.
When should you actually use a million-token context window?
Use a long context when your corpus is small enough to fit inside the window, stable enough to cache across calls, and queried often enough to amortize the cost. Reasoning over one large codebase repeatedly, or a fixed document set that doesn't change, are the cases where this actually pays off. Everything else usually does better with retrieval.
There is a latency dimension too, and it usually gets discovered late. Time to first token scales with input length, so a request carrying several hundred thousand tokens of context takes materially longer to begin responding than a short one. For a batch job that is irrelevant. For anything a person is waiting on, it can rule the approach out entirely regardless of cost. Standardized inference measurement of the kind MLCommons publishes through its MLPerf suites is more informative here than vendor throughput claims, because the methodology is fixed in advance.
Model choice within a family matters as well. Context ceilings are per-model properties, and the cheaper tiers that make long context affordable frequently offer smaller windows than the flagship — which means the configuration you can afford and the configuration you want are not always the same one. The ladder is laid out in our roundup of the model families that matter in late 2026.
Use retrieval instead when the corpus is large, changes frequently, or when most queries need only a small slice of it. Sending a million tokens to answer a question that three paragraphs would have answered is the most expensive architecture available.
At Model Drop our working rule is that long context is a caching strategy in disguise. If the content is stable enough to cache, the window is economical. If it changes every call, the bill scales with the window and retrieval almost always wins.
How to actually test recall before you commit to it
Build a test set from your own documents before trusting a vendor's benchmark. Take twenty to thirty real questions your product actually needs answered, and make sure at least half require combining information from two or more places in the document rather than pulling one isolated fact.
Run that set at three context lengths: a short version, roughly half your target length, and your full target length. If accuracy holds flat across all three, the window is doing real work. If it drops meaningfully at the longer lengths, you've found the actual ceiling for your use case, not the vendor's advertised one.
Vary where the answer sits in the document across your test questions, not just the document length. Put some answers near the start, some in the middle third, and some near the end. A model that scores well on start-and-end placement but poorly in the middle is exhibiting the exact position bias the long-context literature describes, and averaging the score across placements hides it.
We've found this exercise takes under a day and consistently changes the architecture decision. Teams that skip it and trust the headline window number tend to discover the gap in production, on a real support ticket, rather than in a cheap test run where being wrong costs nothing but an afternoon.
The bottom line
A million-token context window is a real ceiling with three practical constraints: uneven attention across position, rate limits that often cap usable throughput below the window, and linear cost on every call. It earns its place for stable, repeatedly-queried corpora and loses badly to retrieval everywhere else.
Your next step: measure recall on your own documents at your actual context length using multi-hop questions, then check your account's tokens-per-minute limit against the window you are planning to use.
By Ashley Quon, Contributing Writer at Model Drop. Reviewed September 2026. Pricing figures are illustrative at published list rates; long-context behavior claims reflect published research rather than Model Drop's own reproduction.
Model Drop covers AI launches — new models, platforms, features, and tools — for the people who have to decide what to actually ship on.