Up front: Structured Output is not the prompt “please return JSON.” It is the API blocking illegal tokens at decode time with a JSON Schema, so the final answer can be parsed by a program. GPT, Gemini, and Claude made it a first-class feature not because the phrase looks good on a slide, but because agents, extraction, and form-fill have to wire the model into a pipeline. Prose fails JSON.parse. It fails the downstream schema even harder.
This article is dated 8 September 2026. All three can now constrain the final JSON to the user or the next service: OpenAI via response_format.json_schema (strict), Gemini via responseMimeType + responseJsonSchema, Claude via GA output_config.format (the older beta output_format still works during the transition). How to fill the fields, and how the subsets differ, we already unpacked in August. This piece answers two questions: what it is, and why all three had to ship it. For the how-to, see From prompt to Structured Output. For OpenAI vs Gemini, see the Structured Output API comparison.
What Structured Output is
Structured Output means: you supply a JSON Schema, and the model’s final reply must be JSON that matches it. The guarantee happens while each token is generated, not after the model “tries to look like JSON.” The names differ: OpenAI says Structured Outputs, Google says Structured Output, Anthropic writes structured outputs / JSON outputs. The extra s is branding. The job is the same.
Think compiler and type checker. A prompt is a comment — the model might listen. A schema is the type system — a wrong field name, a missing required, a string where a number belongs never get emitted. What your program receives is an object, not prose wrapped in a ```json fence.
| Phrase | What it actually means | Common misread |
|---|---|---|
| Structured Output | Constrained decoding of the final reply against a JSON Schema | The model got smarter, or “it can write JSON” |
| JSON Schema | The contract for fields, types, required, enums | A longer prompt |
| Constrained decoding | Illegal tokens are filtered as they are generated | Regex cleanup after the fact |
| strict / hard constraint | The API guarantees shape on a stricter schema subset | Facts are true and numbers are not invented |
A schema all three can read is usually flat: root object, explicit properties / required, additionalProperties: false. Under OpenAI strict, “optional” is often nullable rather than omitted from required. The subsets are not identical; take the intersection first.
{
"type": "object",
"additionalProperties": false,
"properties": {
"task": { "type": "string", "enum": ["extract", "classify", "summarize"] },
"ok": { "type": "boolean" },
"fields": {
"type": "object",
"additionalProperties": false,
"properties": {
"orderId": { "type": "string" },
"total": { "type": "number" },
"note": { "type": ["string", "null"] }
},
"required": ["orderId", "total", "note"]
}
},
"required": ["task", "ok", "fields"]
}
It is not JSON Mode, and not Tool Calling
Three names get flattened into one. They are not the same layer:
| Capability | What it guarantees | What it does not |
|---|---|---|
| Prompt: “return JSON” | A higher probability | Syntax, field names, required lists |
| JSON Mode | The text is parseable JSON | Shape, types, enums |
| Structured Output | The final reply matches the schema | Semantic truth, or that a tool ran |
| Tool Calling | Tool arguments match a schema and the host executes them | The shape of the user-facing answer |
JSON Mode only guarantees matching braces and a successful JSON.parse. The model can still invent order_id when you asked for orderId, or emit the amount as a string. In production, “it parses” is not “it can be inserted.”
Tool Calling / Function Calling constrains the hand that reaches for a tool, not the last sentence to the user. Inventory lookups, file writes, MCP tools/call belong on the tool schema. Extracting mail, classifying a ticket, emitting JSON for a downstream API belong on Structured Output. A full agent often enables both — arguments on tools, the final reply on an output schema. For the layering see What MCP is and the Agent JSON data flow.
Why all three started supporting it
In 2023 you could still gamble on a prompt. In 2026 an agent embeds the model in a loop: the output hits a database, the next tool, or another vendor’s model. The three labs did not coordinate a press cycle. They hit the same product pressure and the same contract — JSON Schema.
- The downstream consumer is a program, not a reader. Chat can be prose. A pipeline needs objects. One missing comma, one renamed field, and the overnight retry queue fills. Vendors would rather cut illegal paths in the decoder than watch every customer write a repairer.
- Agents made stable shape a requirement. In a multi-step loop, last turn’s JSON is this turn’s input. One drift and everything after is wrong. Tool Calling answers “how to reach out.” Structured Output answers “how to hand the conclusion back.” Both need a schema — see whether JSON Schema is becoming the Agent contract.
- Prompts proved they were not enough. “JSON only, no markdown” looks fine on a bench, then drops fields, adds fences, and paraphrases enums once the context is long, tools re-inject, or languages mix. Constrained decoding turns “sometimes” into an API 400 or a retryable schema error.
- JSON Schema was already the lowest common denominator. OpenAPI, MCP
inputSchema, Pydantic / Zod exports — all of it. A private IDL on the model side would force the Host to translate twice. Hooking the final reply to the same schema is what makes vendor swaps cheap. - The race became “can this ship to production,” not “can it chat.” Once one vendor shipped a hard constraint, gateways, agent frameworks, and procurement lists wrote it in as required. The other two either follow or they do not plug into the same graph. In September 2026, a flagship API without Structured Output is hard to sell to anyone who inserts rows.
That is why the dates cluster: OpenAI made Structured Outputs GA in August 2024; Gemini folded MIME + schema into generation config; Claude was still on a beta header in late 2025 and now ships output_config.format as a stable field. The names never matched. The pressure did.
How GPT, Gemini, and Claude turn it on
Align the concept. Do not paste fields across vendors. The table is what you can put in a doc on 8 September 2026 — not a full SDK tutorial.
| Vendor | Entry point | Where the schema hangs | What to watch in 2026 |
|---|---|---|---|
| OpenAI (GPT-5.5 and kin) | response_format on Chat Completions; text.format on the Responses API | type: json_schema + strict: true | Under strict, every object wants additionalProperties: false and properties usually all sit in required; optional becomes nullable |
| Google (Gemini 3.7 Flash and kin) | MIME + schema on the generation config | responseMimeType: application/json + responseJsonSchema (SDK often response_schema) | No switch named strict; older responseSchema used OpenAPI uppercase types; the newer channel uses lowercase JSON Schema |
| Anthropic (Claude 4.6 / 4.8 and kin) | output_config.format on the Messages API | type: json_schema + schema | GA — no structured-outputs-2025-11-13 header required; old output_format still works during transition. Tool-side strict: true is Tool Calling, not the final reply |
The wrappers differ. The schema body should be the same file. Swapping models changes the envelope, not orderId and required. A Claude sketch (spec fields; swap in your business schema):
{
"model": "claude-sonnet-4-6",
"max_tokens": 1024,
"messages": [
{"role": "user", "content": "Extract orderId and total from the order text"}
],
"output_config": {
"format": {
"type": "json_schema",
"schema": {
"type": "object",
"additionalProperties": false,
"properties": {
"orderId": { "type": "string" },
"total": { "type": "number" }
},
"required": ["orderId", "total"]
}
}
}
}
OpenAI puts the same schema in response_format.json_schema and turns on strict. Gemini puts it in responseJsonSchema and declares the JSON MIME type. Full Python side-by-side is still in OpenAI vs Gemini. Product surfaces (ChatGPT / claude.ai / the Gemini web app) do not always expose the same hard constraint. Write the SLA against the API you actually call.
What constrained decoding actually blocks
Without Structured Output the model samples the full vocabulary and hopes the prompt makes it look like JSON. With Structured Output the decoder keeps a legal prefix from the schema: the next token can only be something still valid — a ", orderId, true, or }. Illegal paths get probability zero.
It blocks shape: trailing commas, markdown fences, missing required fields, type drift, extra keys when additionalProperties is false. It does not block invention: total is a number and the number can be made up; a legal enum value can still be the wrong one. Production still runs the same schema through a validator; on failure you retry, degrade, or go to a human. Constrained decoding cuts parse incidents, not hallucinations.
A bigger window does not change that. 1M tokens only widen what is visible; they do not constrain output shape. If you stuff a dump, you still need a schema — see 1M token context windows.
What to do now
- Write the schema before you pick a model. Field names, required lists, and enums are the product contract. GPT / Gemini / Claude are swappable backends. Keep the contract in the repo, not in the prompt.
- Extraction, classification, form-fill → Structured Output. Side effects → Tool Calling. Do not pretend Structured Output already hit the inventory API. Cross-process reuse is when you add MCP.
- Take the schema intersection across vendors: flat objects,
additionalProperties: false, shallow$ref, no rootanyOf. OpenAI strict turns “optional” into nullable. Do not maintain three drifting field tables. - The API passing is not the last check. Save the schema and two or three good / bad fixtures as JSON; validate and Diff on this site. Nothing is uploaded. That is the second gate after constrained decoding.
- Feed failures back as structure: if parse or the second validator fails, return an object (which field, expected type). Do not pour a raw stack into the next turn.
FAQ
Is Structured Output just “make the model return JSON”?
No. A prompt or JSON Mode can emit JSON text. Structured Output filters tokens against a JSON Schema at decode time. Field names, types, and required lists are enforced by the API, not by the model behaving.
Why did GPT, Gemini, and Claude all ship this — isn’t one vendor enough?
Customers want multi-model failover and price shopping. Gateways and agent frameworks already wire “schema in, JSON out.” A vendor without a hard constraint does not plug into that pipeline. Competitive pressure and engineering need are the same fact.
Does Claude still need a fake tool to pretend it has Structured Output?
Not as the main path. In 2026 the Messages API ships JSON Schema output via output_config.format. Tool-level strict still only covers tool arguments. The old beta header and output_format remain in a transition window; new code should use output_config.
If Structured Output is on, do I still validate?
Yes. It guarantees shape and types, not true values or business rules. Run the same schema again in the app; on failure, retry or escalate. In the browser, check fixtures with the JSON toolbox first.
How do I choose among this, MCP, and Tool Calling?
Final reply for a program: Structured Output. External action: Tool Calling. Tools in another process, reused across Hosts: MCP. You can stack the three. Do not let one layer impersonate another.
Can one JSON Schema be sent to all three as-is?
The body can be shared; the request wrapper cannot. A flat object, no extra properties, optionals as nullables, wins most often. OpenAI’s strict subset is the tightest — pass that first, then hand the same file to Gemini / Claude, instead of three drifting schemas.
Takeaways
Structured Output is the 2026 flagship-API socket: the final reply is decoded against a JSON Schema, so programs stop gambling on brackets in a prompt. GPT, Gemini, and Claude all shipped it because agents and extraction wrote “stable shape” into acceptance tests, and JSON Schema was the contract all three already spoke. It is not JSON Mode. It does not replace Tool Calling or MCP.
Swap the model, swap only the wrapper fields. Keep field names and required in the repo, and validate samples locally against the same schema before you go live. How to configure each API, and how that splits from the tool layer, this site already covers. This article only makes “what” and “why” unmistakable.