Up front: by September 2026, GPT-5.5, Claude Opus 4.8 / Sonnet 4.6, and Gemini 3.7 Flash all advertise a window of about one million tokens. “1M tokens” is not “a million Chinese characters,” and filling the window is not the same as using it. The headline limits are close. Effective retrieval, max output, and price steps are not. If you work with JSON, the right move is trim, validate, then decide how much to stuff — not to paste the entire dump.
This piece is written as of 5 September 2026. Figures follow public docs: OpenAI’s gpt-5.5 API at about 1,050,000 tokens total and 128K output; Anthropic’s Opus 4.8 / Sonnet 4.6 at a default 1M and 128K output with no long-context surcharge; Google’s gemini-3.7-flash at 1,048,576 input / 65,536 output. Product surfaces (ChatGPT, Codex, claude.ai) are often narrower. Do not treat a marketing page as the API limit.
What 1M tokens means
A token is the smallest chunk a model counts and bills — not a character, not a word. In English, about 1 token ≈ 0.75 words, or roughly 4 characters. In Chinese, one token often covers 1–2 characters, depending on the tokenizer and punctuation. JSON is more expensive: braces, quotes, repeated keys, and indentation all consume budget.
1M tokens = a context window of about one million tokens. A rough conversion: ~750,000 English words, or ~500,000–800,000 Chinese characters, or one thick book to a few medium ones, or a mid-size repo plus docs. It is not “a million-character novel” and not “a million lines of code.” The same 1M holds very different amounts of signal in prose, tables, and minified JSON.
The context window is the shared budget for a single request: system prompt + history + tool definitions + tool results + this turn’s output. GPT-5-class models also count reasoning tokens. A 900K-token input plus high reasoning effort will hit context_length_exceeded before the model “finishes reading.” The budget is gone.
| Phrase | What it actually means | Common misread |
|---|---|---|
| 1M tokens | A ~1,000,000-token cap on input + output (and reasoning, on some models) | The model can read a million characters or a million lines |
| Context window | Shared budget for one request | Account-wide memory, or automatic recall of every past project |
| Max output | How many tokens this turn may generate | Equal to the window (Gemini’s ~64K output is far below 1M) |
| Effective context | Length where multi-needle retrieval and long-doc reasoning stay stable | Equal to the number on the pricing page |
In 2023 the mainstream was 4K–32K. In 2024 Gemini 1.5 made 1M a headline. By 2025, 200K was a mid-tier default. As of September 2026, flagship APIs have aligned the ad number at 1M. The race is no longer “who is longer,” but who still finds needles past 200K, and whose cache is cheap enough to use every day.
What a long context window is for
A long window solves seeing many materials in one inference. It does not replace a database, and it is not free infinite memory. The jobs below usually need chunking or RAG at 128K. At 1M you can sometimes turn “retrieve, then ask” into “load, then ask” — if what you load was filtered first.
- Whole-repo / multi-file coding: keep related modules, tests, and interface defs in one turn so the model is less likely to miss a file. Agent coding still bottlenecks on tool loops and test feedback, not on the window ad — we covered that in Gemini 3.8 Flash and AI coding.
- Long-document review: contracts, filings, specs, several PDFs side by side. Good for “find the clash between §47 and Appendix B.” Bad for “paste a year of mail and ask for a summary.”
- Long multimodal input: Gemini counts video, audio, images, and text in the same window. An hour of video burns tokens fast. 1M is a quota, not an invitation to upload uncompressed.
- Multi-step agents: tool JSON, error stacks, and the last patch can stay in the thread for a while. A bigger window is not a reason to re-inject raw stacks — see the Agent JSON data flow.
- Large JSON alignment: OpenAPI / JSON Schema, sample payloads, and production error logs in one turn. This is the 1M use most readers of this site should care about, and the easiest way to blow the bill.
- Skipping one RAG layer: when the corpus is stable, bounded, and you need verbatim cites every time, stuffing can be simpler than a vector store. When the corpus changes daily or queries repeat, retrieval is still cheaper and fresher.
The converse: chat, single-label classification, short JSON extraction. A 1M window is waste. Latency, prefills, and per-token billing all punish you. Short jobs get a short model or a short budget.
GPT vs Claude vs Gemini
The table below is what you can put in a doc on 5 September 2026. Prices are official per-million-token list rates; product-surface caps are separate. Gemini 3.8 Flash is still not GA as of this article — keep production on 3.7.
| Vendor / model | Advertised window | Max output | Input / output per 1M | Window caveats |
|---|---|---|---|---|
| OpenAI GPT-5.5 API | ~1,050,000 | 128,000 | $5 / $30 | Reasoning counts against the total; Codex product surface is 400K |
| OpenAI GPT-5.5 Pro | ~1,050,000 | 128,000 | $30 / $180 | Same window; you pay for output and reasoning |
| OpenAI GPT-5.4 | ~1M | 128,000 | $2.50 / $15 | Same 1M class, lower unit price |
| Anthropic Claude Opus 4.8 | 1M (default, no beta header) | 128,000 | $5 / $25 | No long-context surcharge; prompt cache ~90% off |
| Anthropic Claude Sonnet 4.6 | 1M | 128,000 | $3 / $15 | Context awareness (it tracks remaining budget) |
| Anthropic Claude Haiku 4.5 | 200K | 64,000 | $1 / $5 | Fast and cheap — not a 1M tier |
| Google Gemini 3.7 Flash | 1,048,576 | 65,536 | Intro $0.75 / $3.75 through 2026-12-31 | Standard $1.50 / $7.50 from 2027-01-01 |
| Google Gemini 3.1 Pro | 1,048,576 | 65,536 | About $2 / $4 (often a step above 200K) | Do not write third-party 2M / 10M claims into a contract |
How to read it: all three flagships can ingest about 1M. The differences are output headroom, whether reasoning eats the window, whether tokens past 200K cost more, and whether the cache is worth using.
- You need a long structured result in one shot (big JSON, a long patch, a multi-file diff): GPT-5.5 and Claude 4.6/4.8 at 128K output are about 2× Gemini’s ~64K. Gemini is better at “read a lot, write one table.”
- The same system prompt + tool schema every call: Claude prompt cache, Gemini Context Caching, and OpenAI cached input all beat paying full input each time. Agent
toolsJSON belongs in cache — see Tool Calling and JSON Schema. - Tight budget, lots of material, short output: Gemini 3.7 Flash’s intro price is still the cheapest 1M tier of the three. Past 200K, measure multi-needle on your own data; do not trust the window number alone.
- Product ≠ API: GPT-5.5 on ChatGPT / Codex may be 400K; some cloud marketplaces still serve Claude at 200K. Open the model card for the surface you ship on before you write an SLA.
Advertised window vs usable window
By 2026 the consensus is hard to ignore: advertising 1M and using 1M are different jobs. On multi-needle retrieval (MRCR v2-style: find several scattered facts in a huge text), accuracy usually starts to fall after 128K, splits between 200K and 512K, and near 1M only a few models still behave as if they read the whole file.
Public evals and third-party compilations (setups are not identical — treat them as a trend, not an acceptance test) look roughly like this:
- GPT-5.5: a step up from prior GPT-5.x on multi-needle reasoning in the 512K–1M band, and one of the APIs most often named when people actually fill the window. Single-needle “find this sentence” is fine for all three inside 128K.
- Claude Opus 4.6: multi-needle at 1M has been reported around 76%, with a flatter curve. 4.7 / 4.8 lean harder on calibration — refuse or mark uncertainty instead of inventing a location. Long-document compliance review usually wants the latter.
- Gemini 3.x: still strong at retrieval and long-doc understanding inside 128K; multi-needle falls off more steeply past 200K. Good for “read a thick stack, emit a short JSON.” Bad for “guarantee the 8th needle in 900K tokens of logs.”
In production, use three bands — not the price sheet:
| Band | Treat it as | Do this |
|---|---|---|
| ≤128K | Almost fully usable | Default workspace on all three |
| 128K–256K | Measure on samples | Run multi-needle on your JSON / repo; ignore Arena scores |
| 256K–500K | Usable; do not assume every fact is found | Pre-extract key fields with JSONPath; cache repeated prefixes |
| 500K–1M | Fits; “fits” ≠ “understood” | Only stable, high-value material; lock answers with Schema |
“Lost in the middle” did not vanish when windows hit 1M. Put hard constraints at the start of the system prompt and at the end of the user message; park appendices in the middle. Models remember “output JSON that matches the schema” more often than an enum buried at token 400,000.
JSON: stuffing is not validation
A 1M window is the first time “the whole OpenAPI, three days of error logs, and a draft schema” can fit in one request. It does not make illegal JSON legal, and it does not tighten a loose schema. The window is visibility; the contract is shape. We already used that rule in From prompt to Structured Output and OpenAI vs Gemini Structured Output: block illegal tokens at decode time with Schema, then validate again after the call.
Stuffing an 800K-token production dump usually fails for reasons other than “the window was too small”: repeated keys waste budget, middle fields get dropped, array lengths are miscounted, output clips at the 64K/128K cap, and you pay full input. The order is trim → validate → then feed.
{
"name": "longContextPack",
"description": "Pack sent into a 1M window: only validated, trimmed JSON",
"parameters": {
"type": "object",
"additionalProperties": false,
"properties": {
"task": { "type": "string", "enum": ["align_fields", "diff_versions", "extract_errors"] },
"schemaId": { "type": "string", "description": "Stable schema id — do not paste a giant schema inline" },
"focusPaths": {
"type": "array",
"minItems": 1,
"maxItems": 32,
"items": { "type": "string", "description": "JSONPath, e.g. $.paths./v2/orders.post" }
},
"payload": { "type": "object", "description": "Already syntax-checked and size-trimmed — not a raw log string" },
"tokenBudget": { "type": "integer", "minimum": 1000, "maximum": 1000000 }
},
"required": ["task", "schemaId", "focusPaths", "payload", "tokenBudget"]
}
}
Drop that pack into the JSON toolbox for a local syntax check, pull focusPaths with JSONPath, and Diff two payload versions. The model should see a trimmed object, not a 20MB .json file. Swap GPT / Claude / Gemini by changing the model string. The pack’s field names should not change.
What to do now
- Count tokens before you stuff: estimate system, tools, history, and attachments separately. Reserve reasoning on GPT-5.5; count video / PDF pages on Gemini. If you are over budget, cut — do not gamble on “it should still fit.”
- Cache repeated prefixes: system + tools are almost the same every call, so they belong in prompt cache / Context Caching. The first 1M-window invoice is an input bill; after cache hits, the long window becomes daily-drivable.
- Constrain output with Structured Output, not “please return JSON”: OpenAI’s
response_format.json_schema, Gemini’sresponseMimeType+responseJsonSchema, Claude tools or prefills. A longer window will still invent a field and drop a required one if the prompt is loose. - Past the band you measured, retrieve — do not cram: if your own multi-needle failed at 256K, do not treat 700K of logs as full-text understanding. JSONPath on error codes and timestamps is cheaper than letting the model swim.
- Split answers at the output cap: ~64K on Gemini, ~128K on GPT/Claude. A 200K JSON report is shards plus a schema per shard, not one heroic completion.
- Run local fixtures before you hit the API: the same payload through JSON validation and Diff in the browser. Nothing is uploaded. It is the same habit as checking a REST contract before launch.
FAQ
How many words is 1M tokens?
There is no fixed conversion. English is about 750,000 words; Chinese is about 500,000–800,000 characters, depending on punctuation and mixed Latin. JSON, code, and tables consume tokens faster. Do not plan capacity as “a million words.” Count real samples with the vendor tokenizer.
Whose 1M window is best — GPT, Claude, or Gemini?
It depends on the job. For full-window multi-needle reasoning, public evals often put GPT-5.5 first. If you would rather have a refusal than a fabricated citation, Claude 4.7/4.8 is steadier. If you need to read a lot and emit a short JSON table on a budget, Gemini 3.7 Flash’s intro price is the cheapest of the three. The largest advertised number does not win.
Do I still need RAG if I have a 1M window?
Yes. 1M is for stable, bounded material you must quote verbatim. When the corpus changes daily, queries repeat, or your measured effective window is far below 1M, retrieval is still cheaper and easier to update. Long context and RAG complement each other; they do not replace each other.
Can I use the full 1M on ChatGPT or Claude.ai?
Not necessarily. The gpt-5.5 API is about 1.05M; the Codex product surface is often 400K. Some cloud marketplaces still serve Claude at 200K. Trust the model card for the product you are actually using, not the API column copied into a chat-app expectation.
Is Gemini already at 2M or 10M?
As of September 2026, public docs for gemini-3.7-flash and gemini-3.1-pro list 1,048,576 input tokens. 2M / 10M usually comes from earlier experimental tiers, aggregator pages, or not-yet-GA variants. For contracts and quota tables, use Google’s current model card, not a blog headline.
If I stuff a whole JSON file into a 1M window, do I still need a schema?
Yes. The window only widens visibility; it does not constrain output shape. Extraction, alignment, and agent tool arguments still need JSON Schema, then a syntax and structure check after the call. You can swap GPT for Gemini. Field names and required lists should not move.
Takeaways
1M tokens is the 2026 flagship-API ad that GPT-5.5, Claude 4.6/4.8, and Gemini 3.7 can all print on a limits page. It is useful when one inference must see a repo, a long document, or a trimmed JSON pack. It is not free memory, and filling it is not understanding. Usable windows often sit between 128K and 500K; treat a full 1M as capacity, not comprehension. Output caps, whether reasoning consumes the window, and whether tokens past 200K cost more are the real selection fields.
The job is not chasing the next “2M” rumor. It is packing what you send so it survives a schema check: JSONPath to trim paths, Diff to compare versions, validate, then call the API. This site already covers Structured Output and Tool Calling; this article only adds the context-window reading.