A million-token window is cheap now. The expensive part is assuming it buys quality.
The finding first: a very large context window no longer commands a consistent premium. In the current catalogue, million-token-class models range from effectively free tiers and a few cents per million input tokens up to $30 per million, and that spread does not line up neatly with measured quality. You can buy a lot of room for very little money. What you cannot safely assume is that paying more for room buys a better model.
That matters because “1M context” is often treated as a product category of its own, as if the window size answers the purchasing question. It does not. In practice, the catalogue now has enough large-window models that context capacity, input price and measured quality have come apart.
Start with the supply side. Almanac’s catalogue currently lists 396 models on OpenRouter, with a largest advertised context of 2.0M tokens. That top end is no longer hypothetical. The biggest windows in the current list look like this:
- 2.0M tokens: SpaceXAI’s Grok 4.20 and Grok 4.20 Multi-Agent at $1.25/M input tokens.
- 1.311M tokens: Z.ai GLM 5.3 Flash at $0.075/M; DeepSeek V4 Flash Latest at $0.03/M; DeepSeek V4 Flash 0731 at $0.065/M; Meta Llama 4 Scout at $0.11/M; Z.ai GLM 5.3 at $1.40/M.
- ~1.05M tokens: OpenAI GPT-5.6 Luna at $0.20/M; Xiaomi MiMo-V2.5 at $0.14/M; Google Gemini 3.7 Flash at $0.75/M; OpenAI GPT-5.5 at $5/M; OpenAI GPT-5.5 Pro and GPT-5.4 Pro at $30/M.
That is the first practical answer to the question. A “very large” context window now costs, in raw input terms, anywhere from $0.03 to $30 per million tokens among models in roughly the same million-token class. The ratio between the cheap end and the expensive end is not a rounding error; it is three orders of magnitude.
The second answer is more useful: the market is not charging one clear price for context itself. It is charging a mixture of things — model family, vendor positioning, benchmark strength, likely serving cost, and in some cases simple willingness to price high. The context window is attached to those choices, but it does not explain the bill on its own.
The clearest evidence is inside the large-window cohort itself. Among models at or above roughly 1M tokens, there are already obvious low-cost outliers. DeepSeek V4 Flash Latest advertises 1.311M tokens at $0.03/M. GLM 5.3 Flash offers the same 1.311M class at $0.075/M. Xiaomi’s MiMo-V2.5 sits around 1.05M at $0.14/M. OpenAI’s GPT-5.6 Luna is also around 1.05M, but at $0.20/M. None of those prices look like a hard “million-token tax.”
Then look at the high side. OpenAI’s GPT-5.5 Pro and GPT-5.4 Pro are also around 1.05M context, but listed at $30/M input tokens. That is not an incremental premium for context storage. It is premium pricing for the model as a whole.
Put differently: if your workload needs a huge prompt, the catalogue suggests you should treat context length as a filter first, and only then compare the survivors on price and quality. If you do it the other way around — starting from the prestige model and assuming the window explains the price — you will overpay without learning much.
The obvious follow-up is whether those expensive large-window models at least buy measured quality. The answer is: sometimes, but not in proportion to the context bill, and not reliably enough to use price as a proxy.
On Almanac’s measured quality-versus-price frontier, some of the best value models are cheap rather than large or prestigious. Z.ai’s GLM 5.3 Flash scores 57.5 on the quality index at $0.075/M input tokens. Xiaomi’s MiMo-V2.5 scores 38.0 at $0.14/M. OpenAI’s GPT-5.6 Luna reaches 52.3 at $0.20/M. Meanwhile Meta’s Llama 4 Scout, despite a 1.311M-token window, scores just 10.3 at $0.11/M.
That last comparison is the point. A very large window and a low price do not guarantee strong measured quality. But the inverse is also true: a very high price does not buy quality in a neat linear way, and a huge window certainly does not.
There is some relationship between price and quality in the broad catalogue — truly capable models do not all live at the floor — but for large-context buyers it is weak enough to be operationally dangerous. You can find strong measured quality at low prices, mediocre measured quality at low prices, and premium prices attached to million-token windows that tell you more about vendor strategy than about the marginal value of another 800,000 tokens.
So what does a very large context window actually cost in practice?
For admission, surprisingly little. The catalogue now contains at least 15 models with 500K+ context at $1/M or less. If the job is “accept a very long prompt without immediately blowing the budget,” that capability has become cheap.
For use, the arithmetic is still real. A million-token prompt on a model priced at $0.03/M costs about 3 cents in input. On $0.20/M, it costs 20 cents. On $1.25/M, it costs $1.25. On $30/M, it costs $30 before the model produces an output token. Large context is affordable now, but only on some models; the same prompt can still be cheap enough to ignore or expensive enough to dominate unit economics depending on where you send it.
The practical conclusion is unglamorous. Buyers should stop treating context size as a luxury signal. In this catalogue it is increasingly a commodity feature with a very wide price distribution. Screen for the minimum window your workload actually needs, discard anything smaller, and then compare the remainder on measured quality and token price separately. The data does not support paying a premium merely because a model has a million-token badge.
A large context window now often costs less than expected. Assuming it means a better model is what gets expensive.
Both projects are on GitHub and PyPI. Install them.