How AI Generates JSON That Matches JSON Schema: From Prompt to Structured Output

From prompt-only JSON and JSON Mode to vendor Structured Outputs — how JSON Schema constrains model output, OpenAI / Gemini / Anthropic comparison, validation pipeline, and how it differs from Tool Calling.

Earlier posts in this series covered why agents need JSON (Tool Calling to MCP data flow), why JSON Schema became infrastructure (Schema, Function Calling, and MCP evolution), and Gemini-specific Structured Output config (Gemini API guide).

This article zooms out: regardless of which model API you use, how do you go from “please reply in JSON” to “output must match this JSON Schema”? Vendors converged on this in 2024–2026 under names like Structured Outputs / JSON Schema mode — same idea, slightly different field names, Schema subsets, and boundaries with Tool Calling.

Four levels, each stricter than the last

Teams typically use four approaches to get JSON from a model — reliability differs by an order of magnitude:

LevelApproachWhat you controlTypical failure
L0Prompt only: “output JSON”Soft constraint```json fences, prose, single quotes, trailing commas
L1Prompt + few-shot JSON examplesShape by example, no hard ruleField name drift, missing fields, mixed types
L2JSON Mode (response_format: json_object, etc.)Output must be valid JSONParses, but price may be a string
L3Structured Output + JSON SchemaFields, types, enum, requiredSemantic hallucination, truncation, ignored keywords

For production extraction, classification, or form filling, aim for L3. L0–L1 suit exploration; L2 when shape varies and you only need JSON.parse. L3 is the contract programs can consume directly.

What JSON Schema controls — and what it does not

JSON Schema describes document structure: fields, types, required keys, enums, ranges, array item shape. Vendor Structured Outputs compile that Schema into generation — not just paste it in the prompt.

Schema can enforce: syntax shape (object / array / string / integer), required, enum, minimum / maximum, additionalProperties: false, nested objects and arrays.

Schema cannot enforce business correctness. Example: “total_cents must equal sum of line items” — assert that in code after Schema validation. Schema also does not fact-check: a well-typed fabricated invoice number is still hallucination.

Tool inputSchema uses the same language; Structured Output constrains the final reply, Tool Calling constrains tool arguments. See data flow guide.

Constrained decoding: why Schema beats Prompt

Prompts only raise the odds of compliance. Structured Output uses constrained decoding: at each token, the decoder suppresses tokens that would break JSON syntax or violate the Schema.

You usually get parseable, shape-correct JSON without regex-stripping markdown fences. Implementations differ (FSM, grammar, logit masks), but the developer contract is the same: pass Schema to the API, not only the prompt.

Constrained decoding guarantees structure, not semantics. Always re-validate with the same Schema and add business rules in production.

OpenAI, Gemini, Anthropic comparison

Same concept, different field names. Example: extract one invoice object.

VendorJSON ModeStructured Output / SchemaNotes
OpenAIresponse_format: { type: "json_object" }response_format: { type: "json_schema", json_schema: { name, schema, strict: true } }strict: true rejects undeclared fields; works with Pydantic model_json_schema()
Google GeminiresponseMimeType: "application/json"Above + responseJsonSchema or SDK response_schemaSee dedicated Gemini API article
AnthropicPrompt + parseoutput_format (Claude structured output) or schema in Messages APIFields evolve with SDK; keep Schema flat

When migrating vendors, keep the Schema itself standard JSON Schema (type, properties, required, enum); SDKs only wrap the request. Do not mix OpenAPI 3.0 uppercase types (OBJECT) with JSON Schema lowercase (object).

Writing a good Schema: from Pydantic to production

Recommended flow: define types in Pydantic / Zod → export JSON Schema → tune → send to API. Put semantics in description — it enters model context and disambiguates “qty = pieces vs boxes”; type: integer alone cannot.

from pydantic import BaseModel, Field


class LineItem(BaseModel):
    name: str = Field(description="Product name")
    qty: int = Field(description="Quantity, positive integer", ge=1)
    unit_price_cents: int = Field(description="Unit price in cents", ge=0)


class Invoice(BaseModel):
    vendor: str
    currency: str = Field(description="ISO 4217, e.g. CNY")
    items: list[LineItem]
    total_cents: int

schema = Invoice.model_json_schema()
# In production add additionalProperties: false

Practical rules:

  • Prefer object root over root-level array; { "items": [...] } is more stable on some APIs.
  • Start with type / properties / required / enum, then add additionalProperties, min/max; do not dump full Draft 2020-12 — some keywords are ignored.
  • Keep nesting shallow; circular refs are rejected — flatten the Schema.
  • Split Schema vs Prompt: Schema = shape; Prompt = semantics (“extract invoice from text below…”).

OpenAI Structured Outputs example

Chat Completions supports json_schema response format since 2024. With strict: true, output should only contain Schema fields:

from openai import OpenAI
from pydantic import BaseModel

client = OpenAI()

class Invoice(BaseModel):
    vendor: str
    total_cents: int

response = client.chat.completions.create(
    model="gpt-4o-2024-08-06",
    messages=[
        {"role": "user", "content": "Extract invoice: Acme sold 2 keyboards for 398 CNY."}
    ],
    response_format={
        "type": "json_schema",
        "json_schema": {
            "name": "invoice",
            "strict": True,
            "schema": Invoice.model_json_schema(),
        },
    },
)

data = response.choices[0].message.content  # JSON string
import json
invoice = json.loads(data)

Gemini uses response_mime_type + response_json_schema — see the dedicated Gemini API article. For Anthropic, check current SDK structured output docs — same idea, official field names.

Production pipeline: generate → parse → validate → retry

Structured Output is not “one API call and done”. Fix these four steps:

  1. Generate: call model with Schema; log prompt, Schema version, raw content.
  2. Parse: JSON.parse (or SDK parsed); on failure, retry whole response — no half-parse.
  3. Schema validate: run same JSON Schema via AJV / jsonschema / Pydantic; retry or degrade on failure.
  4. Business validate: custom asserts (totals, foreign keys); human or rules engine on failure.

In development, store Schema plus 2–3 positive/negative samples in repo; use JSON Toolbox locally for structure and Diff — same mindset as REST contract tests, consumer is the LLM.

Common pitfalls: asking for “explain then JSON” while JSON Mode is on; truncation (raise max tokens or split tasks); API keys in frontend demos; Schema version drift from prompt.

How it differs from Tool Calling

Structured OutputTool Calling / MCP
ConstrainsFinal JSON reply to userTool argument JSON (inputSchema)
Side effectsNone — data onlyHost / MCP Server executes
Typical useExtract, classify, fill forms, agent handoffInventory, files, external APIs
On failureRetry or humanError in tool message → ask model again

A full agent loop often looks like: Structured Output extracts intent → Tool Calling acts → Structured Output or prose summarizes for user. Do not use Structured Output to pretend “payment API was called” — the model did not call it.

FAQ

Is “please output JSON” in the prompt enough?

No. Prompts only raise compliance odds — markdown fences, trailing commas, and field drift still happen. In production, enable at least JSON Mode; ideally pass JSON Schema through the API Structured Output channel so decoding excludes illegal tokens.

What is the difference between JSON Mode and Structured Output?

JSON Mode only guarantees valid JSON text, not field names, types, or required keys. Structured Output adds JSON Schema and filters tokens during generation — shape stabilizes so you can store or pass to the next hop directly.

Are OpenAI, Gemini, and Anthropic config fields the same?

Same concept, different names. OpenAI: response_format with json_schema and strict; Gemini: responseMimeType + responseJsonSchema; Anthropic: output_format or structured output in tools. Keep Schema standard; SDKs wrap requests.

Can Structured Output replace Tool Calling?

No. Structured Output constrains the final JSON reply; Tool Calling constrains tool argument JSON and requires the host to execute tools. Use the former for extract/classify/fill; the latter for inventory, files, MCP. Full agent chains often use both.

Do I still need to validate model output?

Yes. Constrained decoding cuts syntax errors and type drift but not semantic correctness (valid types, fabricated values). Re-run the same JSON Schema in production; retry, degrade, or review on failure.

How do I validate Schema and sample output locally?

Save JSON Schema and a few model output samples as JSON files; use JSON Toolbox in the browser for syntax and structure checks — nothing is uploaded.

Summary

To get JSON that matches JSON Schema from AI, order matters: define Schema first, enable JSON Mode / Structured Output, then write the prompt. Prompt = semantics; Schema = shape; Pydantic / Zod are author-friendly fronts; vendor APIs expose Schema channels.

Run one real invoice or support transcript end-to-end: Schema → API call → paste output into a validator. When it matches, wire database or next agent. Tool arguments still go through Tool Calling / MCP — do not merge into one API.