Gemini 3.7 Flash Explained: Pricing, API, JSON Output, and Function Calling

Google’s August 2026 Gemini 3.7 Flash — positioning, intro vs 2027 standard pricing, gemini-3.7-flash API access, Structured Output and Function Calling, and how it compares to 3.6 Flash.

On August 13, 2026, Google released Gemini 3.7 Flash — just three weeks after 3.6 Flash — calling it “our most intelligent workhorse model yet for coding and agents.” If you already use the Gemini API for structured JSON or Tool Calling, this guide covers positioning, introductory pricing, how to call the API, JSON output, and Function Calling, and ties into the JSON / Agent series on this blog.

The previous article, How to Generate Structured JSON with the Gemini API, explains responseMimeType and Schema. Earlier, Tool Calling → MCP data flow covered tool-parameter JSON. 3.7 Flash improves both: more reliable multi-step planning, sharper tool calls, and a lower introductory price.

What is Gemini 3.7 Flash

Gemini 3.7 Flash sits in Google’s Flash line — high throughput, low latency, built for production agents and coding. It is not a Pro/Ultra “deep thinking” flagship, but Google claims it beats 3.6 Flash on software engineering, knowledge work, and web development.

Three themes from the launch:

  • Coding: debugging, issue resolution, higher first-pass accuracy (FrontierCode 1.1 Main 43.6% vs 34.4% for 3.6).
  • Agents: more diligent multi-step planning and tool use — DeepSWE v1.1 65.3% vs 49.0%, AutomationBench 30.4% vs 17.0%.
  • Knowledge work: finance, legal, biosciences on long documents (GDP.pdf 34.0% vs 22.0%).

For consumers, Spark on Google AI Pro/Ultra now runs 3.7 Flash. For enterprises, the same model is on Gemini Enterprise Agent Platform and Vertex AI. Most developers still start with the Gemini API (AI Studio key).

Specs and multimodal input

ItemGemini 3.7 Flash
Model ID (API)gemini-3.7-flash
Input modalitiesText, image, video, audio, PDF
Context windowUp to ~1M input tokens
Max output~64k tokens (reasoning output billed as output)
Structured OutputresponseMimeType + Schema supported
Function CallingSupported; same tool declaration format as 3.x
Context CachingSupported; intro cache read ~$0.075 / 1M tokens

Multimodal input plus long context fits tasks like “whole annual report PDF → structured JSON” or “UI screenshot → code.” Schema constrains output shape, not what you pass in contents.

Latest pricing

3.7 Flash launched with an introductory rate — about half of 3.6 Flash’s standard price. Google has published standard pricing from 2027; budget both tiers.

TierWindowInput / 1M tokensOutput / 1M tokens
Intro2026-08-13 – 2026-12-31$0.75$3.75
StandardFrom 2027-01-01$1.50$7.50

Details (Gemini API / AI Studio; Vertex AI has its own SKUs — confirm before committing budget):

  • Reasoning in output: thinking tokens and final answer tokens both count toward output pricing.
  • Context Caching: intro cache read ~$0.075 / 1M; from 2027 ~$0.15 / 1M. Good for fixed system prompts + tool schemas in agents.
  • Batch: offline bulk jobs often have separate discounted rates.
  • Free tier: AI Studio may still offer rate-limited free usage — check the console.

Agent bills often skew toward output tokens. At intro $3.75 / 1M output, 3.7 Flash is half of 3.6 Flash’s $7.50 standard — useful for Q4 2026 agent load tests.

API access and model ID

Use the official google-genai SDK (pip install google-genai) with GEMINI_API_KEY from AI Studio. Model name: gemini-3.7-flash.

from google import genai

client = genai.Client()
response = client.models.generate_content(
    model="gemini-3.7-flash",
    contents="Explain MCP vs Function Calling in three sentences.",
)
print(response.text)

REST is the same as 2.5 / 3.6, with the model in the path:

POST https://generativelanguage.googleapis.com/v1beta/models/gemini-3.7-flash:generateContent
Header: x-goog-api-key: YOUR_KEY

{
  "contents": [{ "role": "user", "parts": [{ "text": "Hello" }] }]
}

On Vertex AI, Structured Output and tool config match; only endpoint and GCP auth differ. Keep API keys on the server — never in static frontends or MCP logs.

Structured JSON output

3.7 Flash fully supports Gemini Structured Outputs: set response_mime_type: application/json plus response_schema (Pydantic) or response_json_schema (JSON Schema dict). Same constrained decoding as 2.5 / 3.6.

from google import genai
from pydantic import BaseModel, Field

class Task(BaseModel):
    title: str
    priority: str = Field(description="low | medium | high")
    due_date: str | None = None

client = genai.Client()
response = client.models.generate_content(
    model="gemini-3.7-flash",
    contents="Extract a todo: finish budget review by next Friday, high priority.",
    config={
        "response_mime_type": "application/json",
        "response_schema": Task,
    },
)
task = response.parsed
print(task.priority)

Structured Output vs Function Calling: the former is final user/downstream JSON; the latter is tool argument JSON. See Prompt to Structured Output and Gemini JSON tutorial. 3.7 Flash is often chosen for stable complex schemas.

Still validate twice: json.loads + JSON Schema; business rules (e.g. totals) in your code. Diff samples locally in the JSON toolbox — nothing uploaded.

Function Calling / Tool Calling

Declare tools with JSON Schema parameters; the model returns function calls; your host executes and feeds results back. 3.7 Flash emphasizes more diligent multi-step tool planning — fewer bad retries, sometimes lower total tokens.

from google import genai
from google.genai import types

get_weather = types.FunctionDeclaration(
    name="get_weather",
    description="Current weather for a city",
    parameters={
        "type": "object",
        "properties": {
            "city": {"type": "string", "description": "e.g. Shanghai"}
        },
        "required": ["city"],
    },
)

client = genai.Client()
response = client.models.generate_content(
    model="gemini-3.7-flash",
    contents="Is it good for a run in Shanghai now?",
    config=types.GenerateContentConfig(
        tools=[types.Tool(function_declarations=[get_weather])],
    ),
)

for part in response.candidates[0].content.parts:
    if part.function_call:
        print(part.function_call.name, part.function_call.args)

Typical agent loop: Tool Calling for facts → Structured Output for fixed JSON → MCP for Cursor / Claude Desktop. Align tool Schema with MCP inputSchema. See JSON Schema, Function Calling, and MCP evolution.

Structured OutputFunction Calling
JSON roleFinal answer / extractionTool arguments
ExecutionParse only, no side effectsYour code / MCP server
3.7 Flash edgeLonger schemas, nested fieldsCoherent multi-step tool plans

vs 3.6 Flash and when to choose

ScenarioRecommendation
New agent / coding (Q3–Q4 2026)Default gemini-3.7-flash, use intro pricing
Stable on 3.6, no JSON/tool painMigration optional; standard prices align in 2027
Long reasoning, lowest hallucinationConsider Gemini Pro / thinking models
Mostly Structured Output3.7 Flash + local Schema validation
Heavy MCP tool loops3.7 Flash + Context Caching for tool schemas

Beyond list price, measure Structured Output pass rate and tool retries on your tasks — that is where 3.7 Flash is meant to win.

Production tips

  1. Swap model globally: set default to gemini-3.7-flash; keep 3.6 as fallback for a week.
  2. Test JSON paths separately: same prompts for Structured Output vs Tool Calling; validate with JSON Schema.
  3. Cache system + tools: fixed system prompt and tools via Context Caching.
  4. Budget at 2027 standard rates: intro output pricing doubles after Dec 31, 2026.
  5. Separate keys and schemas: schemas can be public; keys only on the server.

FAQ

What is the API model ID for Gemini 3.7 Flash?

Use gemini-3.7-flash in the Gemini API and google-genai SDK. REST path: models/gemini-3.7-flash:generateContent. Vertex AI may show region or version suffixes — follow your project docs.

When does intro pricing end?

Intro pricing is $0.75 input / $3.75 output per million tokens through December 31, 2026. From January 1, 2027, standard pricing is $1.50 / $7.50; Context Caching read rates also double.

Does 3.7 Flash support JSON Schema structured output?

Yes. Same as 2.5 / 3.6: response_mime_type=application/json with response_schema or response_json_schema. For production extraction and agent payloads, always use Schema — not prompt-only JSON.

Can I use only Function Calling or only Structured Output?

Yes, but roles differ. Structured Output for classification, forms, extraction; Function Calling for databases, email, MCP. Production agents often use both in one conversation.

Do I need code changes migrating from gemini-3.6-flash?

Usually just the model string plus regression tests. tools, responseJsonSchema, and multimodal parts stay compatible. Compare JSON samples in staging before full cutover.

How do I validate model JSON locally?

Save responseJsonSchema and outputs as JSON files; use the JSON toolbox in the browser for syntax and structural diff — nothing is uploaded.

Summary

Gemini 3.7 Flash is Google’s 2026 workhorse for coding and agents: model ID gemini-3.7-flash, intro pricing through 2026 at $0.75 / $3.75 per million input/output tokens, doubling in 2027. The API surface matches earlier Flash models; the win is more stable multi-step tools and Structured Output in production.

Try 3.7 Flash in staging with JSON + tool regression, and diff outputs against Schema locally. For configuration depth, keep using the series articles — Structured Output and Function Calling are not the same API.