Gemini APIで構造化JSONを生成する:開発者向け完全ガイド

プロンプトだけのJSON、responseMimeType から responseSchema / JSON Schema まで。Gemini の Structured Outputs、Python と REST の例、Function Calling との役割分担、公開前の検証方法。

前回の記事AI エージェント が JSON なしでは生きていけない理由Tool Calling と MCP をホップごとにトレースします。これは、最終返信このモデルは、ユーザーまたはダウンストリーム プログラムに、解析、検証、保存できる JSON を Gemini に発行させる方法を提供します。単に JSON のように見える散文ではありません。

それは、Gemini ドキュメントの構造化出力 (制御された生成) です。スキーマのアイデアは 関数呼び出し と共有されますが、ターゲットが異なります。前者は、最終ペイロード;後者は制約しますツール 引数。 エージェント が「請求書の抽出」と「支払いAPI への電話」を同じ種類の通話として扱わないように、それらを分離してください。

それぞれより厳格な 3 つのアプローチ

チームは通常、「Gemini に JSON を出力させる」ために 3 つの方法を試みます。信頼性は桁違いに異なります。

アプローチあなたがコントロールするもの十分なとき
プロンプトのみ: 「JSON を出力してください」ソフト制約。値下げフェンスと末尾のコメントが引き続き表示される探索、一回限りのスクリプト
responseMimeType: application/json出力は有効な JSON テキストである必要があります形状はさまざまです。成功するには parse() だけが必要です
MIME + responseSchema / responseJsonSchemaフィールド、型、列挙型、および必要なキーには制約があります本番環境の抽出、フォーム、エージェント間のペイロード

Everyone has seen the first failure mode: a ```json fence, an extra paragraph, single quotes, a trailing comma. The second layer parses, but price may be a string and items may be missing. The third layer is this tutorial: hand JSON Schema to the API so the decoder avoids illegal paths at each token.

制約付きデコード: スキーマがプロンプトに勝る理由

A prompt only raises the odds that the model wants to comply. Structured Output compiles the Schema into generation: if the next token would break JSON syntax or leave the Schema (for example starting an undeclared field), its probability is suppressed. So response.text is usually a parseable object — no regex to strip fences.

Since 2025 the Gemini API complements the OpenAPI 3.0-style responseSchema with standard JSON Schema (often responseJsonSchema on the wire). Pydantic model_json_schema() and Zod exports can be sent almost as-is. Gemini 2.5 and later also tend to preserve property order from the Schema, which helps CSV columns and tables downstream.

Classification has a side path: responseMimeType: text/x.enum emits only the enum string (for example Keyboard), with no braces. Use application/json when you need an object; use the enum MIME when you need a single label.

Python: 完全な google-genai の例

Prefer the current SDK google-genai (from google import genai). Do not mix it with the legacy google-generativeai package. With GEMINI_API_KEY set:

from google import genai
from pydantic import BaseModel, Field


class LineItem(BaseModel):
    name: str = Field(description="Product name")
    qty: int = Field(description="Quantity, positive integer")
    unit_price_cents: int = Field(description="Unit price in cents")


class Invoice(BaseModel):
    vendor: str
    currency: str = Field(description="ISO 4217, e.g. CNY")
    items: list[LineItem]
    total_cents: int


client = genai.Client()
response = client.models.generate_content(
    model="gemini-2.5-flash",
    contents="Extract an invoice from: Acme sold 2 keyboards at 199 CNY each.",
    config={
        "response_mime_type": "application/json",
        "response_schema": Invoice,
    },
)

print(response.text)      # JSON string
invoice = response.parsed  # Invoice instance (Pydantic path)
print(invoice.total_cents)

response.parsed is meaningful when response_schema is a Pydantic or SDK type. If you pass a raw JSON Schema dict (next section), json.loads(response.text) and validate yourself.

For many records use list[Invoice] or wrap invoices: list[Invoice] in an object. An array at the root is less stable on some models than always returning an object; production code usually does the latter.

response_schema vs JSON スキーマ

どの設定キーを使用するかを推測しないでください。

  • 応答スキーマ: Pydantic モデル、Python Enum、または SDK スキーマ オブジェクト。 SDK は、これをオンワイヤの OpenAPI サブセットにマッピングします。
  • response_json_schema: a JSON Schema object (dict). Use it for Invoice.model_json_schema(), Zod toJSONSchema(), and richer keywords such as additionalProperties, minimum / maximum, and prefixItems.
schema = {
  "type": "object",
  "properties": {
    "vendor": { "type": "string" },
    "currency": { "type": "string", "enum": ["CNY", "USD", "EUR"] },
    "items": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "name": { "type": "string" },
          "qty": { "type": "integer", "minimum": 1 },
          "unit_price_cents": { "type": "integer", "minimum": 0 }
        },
        "required": ["name", "qty", "unit_price_cents"],
        "additionalProperties": False
      }
    },
    "total_cents": { "type": "integer" }
  },
  "required": ["vendor", "currency", "items", "total_cents"],
  "additionalProperties": False
}

response = client.models.generate_content(
    model="gemini-2.5-flash",
    contents="Extract the invoice: …",
    config={
        "response_mime_type": "application/json",
        "response_json_schema": schema,
    },
)

Older REST docs use uppercase types in responseSchema (OBJECT, STRING, ARRAY, INTEGER). The JSON Schema path uses lowercase object / string. Do not mix the two keyword sets. Put field meaning in description: it enters the model context and decides whether qty is pieces or cases. Types alone cannot.

REST リクエストの内容

On the Gemini Developer API, generateContent puts structured output under generationConfig. The key goes in x-goog-api-key or a query param — never in a frontend repo.

POST https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash:generateContent

{
  "contents": [
    {
      "role": "user",
      "parts": [{ "text": "Extract an invoice from the text: …" }]
    }
  ],
  "generationConfig": {
    "responseMimeType": "application/json",
    "responseJsonSchema": {
      "type": "object",
      "properties": {
        "vendor": { "type": "string" },
        "total_cents": { "type": "integer" }
      },
      "required": ["vendor", "total_cents"]
    }
  }
}

On success the candidate text is candidates[0].content.parts[0].text — a JSON string. Vertex AI uses the same field names; only the endpoint and GCP auth change. Images and PDFs can be inputs: the Schema constrains output, not multimodal input.

関数呼び出し で作業を分割する方法

どちらも JSON を管理するためにスキーマを使用しますが、異なるホップ上にあります。

構造化された出力関数呼び出し / ツール呼び出し
制約されるもの最終返信 JSONツール引数 JSON
副作用を実行するのは誰ですか誰でもない;それは単なるデータですホスト / MCP サーバー
一般的な構成responseMimeType + スキーマtools[].parameters / inputSchema
失敗時再試行するか、人間に戻りますエラーをツールメッセージとして書き込み、再度質問してください

請求書の抽出、モデレーションラベル、メモをタスクリストに変換: 構造化された出力。インベントリの検索、チケットの作成、リポジトリ ファイルの読み取り: ツール - を参照データフローの記事そしてJSON スキーマ と MCP の進化。構造化出力を使用して、支払い API がすでに実行されているように見せかけないでください。モデルが呼び出したわけではありません。

実行時検証とよくある落とし穴

制約されたデコードはビジネス上の正確さではありません。 2 つのゲートを維持します。

  1. Syntax and Schema: after json.loads, validate again with the same JSON Schema (required, enum, minimum).
  2. Business invariants: for example sum(item.qty * item.unit_price_cents) == total_cents. Schema cannot express that; you write it.

よくある落とし穴:

  • サポートされていないキーワード: 完全なドラフト 2020-12 スキーマをダンプすると、一部のキーワードが無視される場合があります。 type/properties/required/enum/items から始めて、AdditionalProperties と min/max を追加します。
  • Array at the root: { "items": [ ... ] } as an object root is often more reliable.
  • Markdown コメントと JSON の混合: JSON MIME がオンになったら、「最初に説明してから JSON」と尋ねないでください。
  • 切り捨て: maxOutputTokens を発生させるか、「最初にリストを作成し、次に各行を埋める」ように分割します。
  • フロントエンドのキー: ブラウザーの構造化出力デモで API キーが漏洩します。スキーマはパブリックにすることができます。キーはサーバー上に残ります。

開発中は、スキーマと 2 つまたは 3 つのポジティブ/ネガティブ サンプルを git に保存します。 JSON Toolbox でローカルに構造と Diff を検査します。これは、Gemini をコンシューマとして使用する、REST と同じコントラクト テストの習慣です。

よくある質問

JSON MIME タイプのみを送信する場合と、スキーマも送信する場合の違いは何ですか?

responseMimeType application/json のみを使用すると、モデルは有効な JSON を出力しようとしますが、フィールド名、型、および必要なキーには制約がありません。 responseSchema または responseJsonSchema を追加すると、デコード中にトークンが制約されるため、形状は永続化または次のエージェントに渡すのに十分安定しています。

response_schema と response_json_schema を選択するにはどうすればよいですか?

response_schema を Pydantic モデルまたは SDK スキーマで使用します。 SDK は、response.parsed を公開できます。完全な JSON Schema オブジェクト (AdditionalProperties、min/max、prefixItems) に対して、または Pydantic/Zod model_json_schema() をそのまま送信する場合は、response_json_schema を使用します。どちらも response_mime_type=application/json が必要です。

構造化出力は 関数呼び出し を置き換えることができますか?

いいえ。構造化出力は、ユーザーまたはダウンストリーム コードに表示される最終的な JSON を制限します。 関数呼び出し / ツール呼び出し はツール引数 JSON を制約しますが、それでもホストがツールを実行する必要があります。構造化された出力を抽出、分類、フォーム入力に使用します。天気、ファイル、MCP のツールを使用します。 Agent パイプラインでは、多くの場合両方が使用されます。

モデルは 100% のスキーマ準拠を保証しますか?

制約付きデコードでは、構文エラーや型ドリフトは発生しますが、意味上の幻覚、切り捨て、およびサポートされていないキーワードの無視は引き続き発生します。本番環境では、バリデーターを介して同じスキーマを実行し、失敗した場合は再試行するか機能を低下させます。

ネストされたオブジェクト、配列、列挙型はサポートされていますか?

はい。オブジェクト、配列、文​​字列列挙は通常の組み合わせです。分類のために、MIME を text/x.enum に設定して、モデルが JSON オブジェクトではなく列挙値のみを出力するようにすることができます。非常に深い入れ子や循環参照は拒否される可能性があります。スキーマをフラット化します。

スキーマとサンプル出力をローカルで検証するにはどうすればよいですか?

responseJsonSchema といくつかのモデル出力サンプルを JSON ファイルとして保存します。ブラウザーの JSON ツールボックス で構文と構造をローカルで確認します。何もアップロードされません。出荷後の実行時に同じスキーマを再度使用します。

まとめ

To get structured JSON from Gemini, the order is: Schema first, JSON MIME second, prompt last. The prompt owns meaning (what to extract); the Schema owns shape (what fields look like). Pydantic / Zod are author-friendly fronts; on the wire you send response_schema or response_json_schema.

Run one real invoice or a support transcript: write the Schema → call generateContent once → paste response.text into a validator. Only then wire a database or the next agent. Tool arguments still go through Function Calling / MCP — do not collapse them into one API.