Claude's 1M-token context window: when it is worth the cost, and when retrieval wins
A 1M-token prompt costs about $5 of input per request on Opus, resent every turn. When long context beats retrieval, and how caching changes the math.
A million-token context window makes a tempting shortcut: skip the retrieval pipeline and put the whole document set in the prompt. Sometimes that is exactly right. But the window is priced per token like everything else, and the part people miss is that a chat API is stateless — every turn resends the entire context. A prompt that is cheap to send once becomes expensive to send ten times.
The math for one turn
Claude Opus 4.8 has a 1M-token context window and bills $5 per million input tokens and $25 per million output tokens at Anthropic's current list price. So a full million tokens of input is $5.00 per request, before the model writes a word. Run a realistic turn — a 900K-token document plus a short question, about a thousand tokens of answer — through the same cost function the kit meters every request with:
// usage: { inputTokens, outputTokens, cacheReadTokens, cacheCreationTokens }
costUsd("claude-opus-4-8", { inputTokens: 900_000, outputTokens: 1_000, cacheReadTokens: 0, cacheCreationTokens: 0 });
// (900_000 * 5 + 1_000 * 25) / 1_000_000 = $4.525 per turn
costUsd("claude-opus-4-8", { inputTokens: 20_000, outputTokens: 1_000, cacheReadTokens: 0, cacheCreationTokens: 0 });
// the same question with 20K tokens of retrieved context = $0.125 per turn
Over a ten-turn conversation about that document, that is roughly $45 in long context versus $1.25 with retrieval. The output cost is identical; the entire gap is input tokens you resend on every turn.
Caching changes the math, not the conclusion
Prompt caching is what makes long context viable at all. Put the document in a stable prefix, mark it with cache_control, and after the first request the cached tokens bill at about a tenth of the input rate, while the turn that writes the cache costs about 1.25x. For the same ten turns:
- First turn (cache write): about $5.65.
- Each later turn (cache read): about $0.48.
- Ten turns total: about $9.93 — a 4.5x saving over no caching, still about 8x the retrieval version.
The cache has a lifetime (five minutes by default, refreshed on each hit), so a conversation with long pauses pays the write again. And the prefix has to be byte-identical, which is why the kit keeps system prompts deterministic — the full break-even analysis is in the prompt-caching cost math.
When long context actually wins
Paying for the full window is worth it when the task genuinely needs the whole thing at once: cross-referencing a long contract against itself, reasoning over a codebase where the relevant pieces are not knowable in advance, or summarizing a document end to end. Retrieval is a bet that you can find the relevant 2% before the model sees anything; when that bet is bad, retrieval returns confident answers built on the wrong excerpts. Long context is also faster to ship — no chunking, embeddings, or index to maintain — which matters for a prototype or a low-volume internal tool.
Retrieval wins when the questions are narrow and the corpus is large or shared: support bots over a help center, Q&A over a knowledge base, anything at high request volume. There the per-turn input bill dominates, and the per-user cost of a long-context design can erase the margin on a plan.
The routing trap: not every model has the window
Context windows differ by model. Opus 4.8 and Sonnet 5 take a million tokens; Haiku 4.5 takes 200K. If your plans route free users to Haiku — as the plan-gating pattern does — a long-context feature that works fine on Pro fails for free users the moment a request crosses 200K. Gate the feature by plan, or by the routed model's window, rather than discovering it in an error log.
Things that bite
- Count before you send. The SDK's token-counting endpoint tells you how big a request is before you pay for it; use it to enforce a ceiling instead of letting a user upload push a single turn to $5.
- Check for a long-context price tier. The models above bill at a flat rate in the current price table, but some models charge more beyond a prompt-size threshold (Haiku 5.5 does above 100K tokens). Read your model's pricing before you design around a huge window.
- Bigger prompts are slower. More input means a longer wait for the first token; a full window is a bad fit for anything that has to feel instant.
- Meter it. Long-context turns are exactly the requests that blow up a per-user cost average. Record cache reads and writes per request so you can see the effect of caching, not just hope for it.
Shipwright — the Next.js and Claude starter kit this blog documents — records input, output, cache-read, and cache-write tokens on every request and prices them per model, so the cost of a long-context design shows up in your numbers on day one. Watch per-request cost in the live demo.