How to Generate Structured JSON with the Gemini API: A Complete Developer Guide

From prompt-only JSON and responseMimeType to responseSchema / JSON Schema — how Gemini structured outputs work, with Python and REST examples, Function Calling vs JSON mode, and how to validate before you ship.

The previous article Why AI Agents Can't Live Without JSON traced Tool Calling and MCP hop by hop. This one looks at the final reply the model gives a user or a downstream program: how to make Gemini emit JSON you can parse, validate, and store — not prose that merely looks like JSON.

That is Structured Outputs (controlled generation) in the Gemini docs. It shares Schema ideas with Function Calling but a different target: the former constrains the final payload; the latter constrains tool arguments. Keep them apart so an Agent does not treat “extract an invoice” and “call the payment API” as the same kind of call.

Three approaches, each stricter

Teams usually try three ways to “make Gemini output JSON”. Reliability differs by an order of magnitude:

ApproachWhat you controlWhen it is enough
Prompt only: “please output JSON”Soft constraint; markdown fences and trailing commentary still appearExploration, one-off scripts
responseMimeType: application/jsonOutput must be valid JSON textShape varies; you only need parse() to succeed
MIME + responseSchema / responseJsonSchemaFields, types, enums, and required keys are constrainedProduction extraction, forms, agent-to-agent payloads

Everyone has seen the first failure mode: a ```json fence, an extra paragraph, single quotes, a trailing comma. The second layer parses, but price may be a string and items may be missing. The third layer is this tutorial: hand JSON Schema to the API so the decoder avoids illegal paths at each token.

Constrained decoding: why Schema beats a prompt

A prompt only raises the odds that the model wants to comply. Structured Output compiles the Schema into generation: if the next token would break JSON syntax or leave the Schema (for example starting an undeclared field), its probability is suppressed. So response.text is usually a parseable object — no regex to strip fences.

Since 2025 the Gemini API complements the OpenAPI 3.0-style responseSchema with standard JSON Schema (often responseJsonSchema on the wire). Pydantic model_json_schema() and Zod exports can be sent almost as-is. Gemini 2.5 and later also tend to preserve property order from the Schema, which helps CSV columns and tables downstream.

Classification has a side path: responseMimeType: text/x.enum emits only the enum string (for example Keyboard), with no braces. Use application/json when you need an object; use the enum MIME when you need a single label.

Python: a full google-genai example

Prefer the current SDK google-genai (from google import genai). Do not mix it with the legacy google-generativeai package. With GEMINI_API_KEY set:

from google import genai
from pydantic import BaseModel, Field


class LineItem(BaseModel):
    name: str = Field(description="Product name")
    qty: int = Field(description="Quantity, positive integer")
    unit_price_cents: int = Field(description="Unit price in cents")


class Invoice(BaseModel):
    vendor: str
    currency: str = Field(description="ISO 4217, e.g. CNY")
    items: list[LineItem]
    total_cents: int


client = genai.Client()
response = client.models.generate_content(
    model="gemini-2.5-flash",
    contents="Extract an invoice from: Acme sold 2 keyboards at 199 CNY each.",
    config={
        "response_mime_type": "application/json",
        "response_schema": Invoice,
    },
)

print(response.text)      # JSON string
invoice = response.parsed  # Invoice instance (Pydantic path)
print(invoice.total_cents)

response.parsed is meaningful when response_schema is a Pydantic or SDK type. If you pass a raw JSON Schema dict (next section), json.loads(response.text) and validate yourself.

For many records use list[Invoice] or wrap invoices: list[Invoice] in an object. An array at the root is less stable on some models than always returning an object; production code usually does the latter.

response_schema vs JSON Schema

Do not guess which config key to use:

  • response_schema: a Pydantic model, Python Enum, or SDK Schema object. The SDK maps it to the on-wire OpenAPI subset.
  • response_json_schema: a JSON Schema object (dict). Use it for Invoice.model_json_schema(), Zod toJSONSchema(), and richer keywords such as additionalProperties, minimum / maximum, and prefixItems.
schema = {
  "type": "object",
  "properties": {
    "vendor": { "type": "string" },
    "currency": { "type": "string", "enum": ["CNY", "USD", "EUR"] },
    "items": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "name": { "type": "string" },
          "qty": { "type": "integer", "minimum": 1 },
          "unit_price_cents": { "type": "integer", "minimum": 0 }
        },
        "required": ["name", "qty", "unit_price_cents"],
        "additionalProperties": False
      }
    },
    "total_cents": { "type": "integer" }
  },
  "required": ["vendor", "currency", "items", "total_cents"],
  "additionalProperties": False
}

response = client.models.generate_content(
    model="gemini-2.5-flash",
    contents="Extract the invoice: …",
    config={
        "response_mime_type": "application/json",
        "response_json_schema": schema,
    },
)

Older REST docs use uppercase types in responseSchema (OBJECT, STRING, ARRAY, INTEGER). The JSON Schema path uses lowercase object / string. Do not mix the two keyword sets. Put field meaning in description: it enters the model context and decides whether qty is pieces or cases. Types alone cannot.

What the REST request looks like

On the Gemini Developer API, generateContent puts structured output under generationConfig. The key goes in x-goog-api-key or a query param — never in a frontend repo.

POST https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash:generateContent

{
  "contents": [
    {
      "role": "user",
      "parts": [{ "text": "Extract an invoice from the text: …" }]
    }
  ],
  "generationConfig": {
    "responseMimeType": "application/json",
    "responseJsonSchema": {
      "type": "object",
      "properties": {
        "vendor": { "type": "string" },
        "total_cents": { "type": "integer" }
      },
      "required": ["vendor", "total_cents"]
    }
  }
}

On success the candidate text is candidates[0].content.parts[0].text — a JSON string. Vertex AI uses the same field names; only the endpoint and GCP auth change. Images and PDFs can be inputs: the Schema constrains output, not multimodal input.

How to split work with Function Calling

Both use Schema to govern JSON, but they sit on different hops:

Structured OutputFunction Calling / Tool Calling
What is constrainedFinal reply JSONTool-argument JSON
Who runs side effectsNobody; it is just dataHost / MCP Server
Typical configresponseMimeType + Schematools[].parameters / inputSchema
On failureRetry or fall back to a humanWrite the error as a tool message and ask again

Invoice extraction, moderation labels, turning notes into a task list: Structured Output. Inventory lookup, creating a ticket, reading a repo file: tools — see the data-flow article and JSON Schema and MCP evolution. Do not use Structured Output to pretend a payment API already ran — the model did not call it.

Runtime validation and common pitfalls

Constrained decoding is not business correctness. Keep two gates:

  1. Syntax and Schema: after json.loads, validate again with the same JSON Schema (required, enum, minimum).
  2. Business invariants: for example sum(item.qty * item.unit_price_cents) == total_cents. Schema cannot express that; you write it.

Common pitfalls:

  • Unsupported keywords: dumping a full Draft 2020-12 Schema may silently ignore some keywords. Start with type / properties / required / enum / items, then add additionalProperties and min/max.
  • Array at the root: { "items": [ ... ] } as an object root is often more reliable.
  • Mixing Markdown commentary with JSON: once JSON MIME is on, do not also ask for “explain first, then JSON”.
  • Truncation: raise maxOutputTokens, or split “list first, then fill each row”.
  • Keys in the frontend: structured-output demos in the browser leak API keys. The Schema can be public; the key stays on the server.

During development, keep the Schema and two or three positive/negative samples in git. Inspect structure and Diff locally in JSON Toolbox — the same contract-testing habit as REST, with Gemini as the consumer.

FAQ

What is the difference between JSON MIME type alone and also sending a Schema?

With only responseMimeType application/json, the model tries to emit valid JSON, but field names, types, and required keys are unconstrained. Adding responseSchema or responseJsonSchema constrains tokens during decoding, so the shape is stable enough to persist or pass to the next agent.

How do I choose response_schema vs response_json_schema?

Use response_schema with a Pydantic model or SDK Schema; the SDK can expose response.parsed. Use response_json_schema for a full JSON Schema object (additionalProperties, min/max, prefixItems) or when sending Pydantic/Zod model_json_schema() as-is. Both require response_mime_type=application/json.

Can structured output replace Function Calling?

No. Structured output constrains the final JSON the user or downstream code sees. Function Calling / Tool Calling constrains tool-argument JSON and still requires the host to execute the tool. Use structured output for extraction, classification, and form-filling; use tools for weather, files, and MCP. Agent pipelines often use both.

Does the model guarantee 100% Schema compliance?

Constrained decoding slashes syntax errors and type drift, but semantic hallucinations, truncation, and ignored unsupported keywords still happen. In production, run the same Schema through a validator and retry or degrade on failure.

Are nested objects, arrays, and enums supported?

Yes. Objects, arrays, and string enums are the usual combination. For classification you can set MIME to text/x.enum so the model emits only the enum value, not a JSON object. Very deep nesting or cyclic refs may be rejected — flatten the Schema.

How do I validate Schema and sample output locally?

Save responseJsonSchema and a few model output samples as JSON files. Check syntax and structure locally in JSON Toolbox in the browser — nothing is uploaded. Use the same Schema again at runtime after you ship.

Summary

To get structured JSON from Gemini, the order is: Schema first, JSON MIME second, prompt last. The prompt owns meaning (what to extract); the Schema owns shape (what fields look like). Pydantic / Zod are author-friendly fronts; on the wire you send response_schema or response_json_schema.

Run one real invoice or a support transcript: write the Schema → call generateContent once → paste response.text into a validator. Only then wire a database or the next agent. Tool arguments still go through Function Calling / MCP — do not collapse them into one API.