Gemini 3.7 Flash 是什么?最新价格、API、JSON 输出、Function Calling 完整介绍

Google 2026 年 8 月发布的 Gemini 3.7 Flash:模型定位、Intro 限时价与 2027 标准价、gemini-3.7-flash API 接入、Structured Output 与 Function Calling 要点,及与 3.6 Flash 选型对比。

2026 年 8 月 13 日,Google 发布 Gemini 3.7 Flash——距 3.6 Flash 仅三周,官方称其为「面向编码与 Agent 的最强 workhorse 模型」。如果你已经在用 Gemini API 做结构化 JSON 或 Tool Calling,这篇把模型定位、限时价格、接入方式、JSON 输出与 Function Calling一次讲清,并接上本系列前几篇的数据流上下文。

上一篇《Gemini API 如何生成结构化 JSON》讲的是 responseMimeType 与 Schema;更早的《Tool Calling → MCP 数据流》讲的是工具参数 JSON。3.7 Flash 的升级重点正是这两类 JSON 场景:更稳的多步规划、更准的工具调用、更低的 Intro 单价。

Gemini 3.7 Flash 是什么

Gemini 3.7 Flash 属于 Google Flash 系列——高吞吐、低延迟、面向生产 Agent 与编码的中间档模型,不是 Pro/Ultra 那种「无限思考」旗舰,但在软件工程、知识工作和 Web 开发上,Google 宣称其智能已明显超过 3.6 Flash。

官方强调的三条主线:

  • 编码:调试、Issue 修复、首遍代码准确率提升(FrontierCode 1.1 Main 43.6% vs 3.6 的 34.4%)。
  • Agent:多步规划、工具调用更「较真」,DeepSWE v1.1 65.3% vs 49.0%,AutomationBench 30.4% vs 17.0%。
  • 知识工作:金融、法律、生物等长文档推理(GDP.pdf 34.0% vs 22.0%)。

对个人用户,Google AI Pro/Ultra 订阅的 Spark 个人 Agent 也已切换到 3.7 Flash。对企业,同一模型在 Gemini Enterprise Agent Platform 与 Vertex AI 上提供。开发者最常见入口仍是 Gemini API(AI Studio 密钥)。

规格与多模态能力

项Gemini 3.7 Flash
模型 ID(API)gemini-3.7-flash
输入模态文本、图片、视频、音频、PDF
上下文窗口最高约 100 万 token 输入
最大输出约 64k token(含 reasoning 输出计费)
Structured Output支持 responseMimeType + Schema
Function Calling支持,与 3.x 系列工具声明格式一致
Context Caching支持,Intro 期缓存读取约 $0.075 / 1M token

多模态输入 + 长上下文,适合「整份年报 PDF + 结构化抽取 JSON」或「截图 UI + 生成代码」一类任务。Schema 约束的是输出形状,不限制你能喂什么进 contents。

最新价格与计费要点

3.7 Flash 上线时带限时 Intro 价:约为 3.6 Flash 正式价的一半。Google 已公布 2027 年起恢复「标准价」,预算要按两档分别估算。

档位有效期输入 / 1M token输出 / 1M token
Intro 价2026-08-13 ~ 2026-12-31$0.75$3.75
标准价2027-01-01 起$1.50$7.50

补充计费细节(Gemini API / AI Studio 口径,Vertex AI 有独立 SKU,部署前请核对官方价目):

  • 输出含 reasoning:带思考链的输出 token 与最终答复 token 一并计入 output 单价。
  • Context Caching:Intro 期缓存读取约 $0.075 / 1M;2027 年起约 $0.15 / 1M。适合 Agent 固定 System Prompt + 工具 Schema 反复发送的场景。
  • Batch:大批量离线任务通常有折扣,与实时 generateContent 分开计价。
  • 免费层:AI Studio 仍可能有速率受限的免费额度,具体上限以控制台为准。

Agent 账单往往「输出 token > 输入 token」——3.7 Flash 在 Intro 期 $3.75 / 1M 输出,比 3.6 Flash 正式 $7.50 便宜一半,适合 Q4 2026 做 Agent 压测与灰度。

API 接入与模型 ID

推荐官方新 SDK google-genai(pip install google-genai),环境变量 GEMINI_API_KEY 来自 AI Studio。模型名固定为 gemini-3.7-flash——不要写成 gemini-3.7-flash-001 除非文档明确要求带后缀。

from google import genai

client = genai.Client()
response = client.models.generate_content(
    model="gemini-3.7-flash",
    contents="用三句话解释 MCP 和 Function Calling 的区别。",
)
print(response.text)

REST 入口与 2.5 / 3.6 相同,只是 path 里的 model 换成 gemini-3.7-flash:

POST https://generativelanguage.googleapis.com/v1beta/models/gemini-3.7-flash:generateContent
Header: x-goog-api-key: YOUR_KEY

{
  "contents": [{ "role": "user", "parts": [{ "text": "Hello" }] }]
}

Vertex AI 上字段名与 Structured Output / Tool 配置一致,endpoint 与 GCP 鉴权不同。Android Studio、Google Antigravity 也可选同一模型做 Agent 原型。密钥只放服务端,不要写进前端静态站或 MCP Server 日志。

JSON 结构化输出

3.7 Flash 完整支持 Gemini Structured Outputs(结构化输出):在 generationConfig / SDK config 里设 response_mime_type: application/json,再配 response_schema(Pydantic)或 response_json_schema(JSON Schema 字典)。原理与 2.5 / 3.6 相同——约束解码在生成阶段就避开非法 JSON 路径。

from google import genai
from pydantic import BaseModel, Field

class Task(BaseModel):
    title: str
    priority: str = Field(description="low | medium | high")
    due_date: str | None = None

client = genai.Client()
response = client.models.generate_content(
    model="gemini-3.7-flash",
    contents="从会议纪要提取待办:下周五前完成预算评审,高优先级。",
    config={
        "response_mime_type": "application/json",
        "response_schema": Task,
    },
)
task = response.parsed
print(task.priority)

与 Function Calling 的分工不变:Structured Output 管给用户/下游程序的最终 JSON;Tool Calling 管工具参数 JSON。详见《从 Prompt 到 Structured Output》与《Gemini JSON 教程》。3.7 Flash 在复杂 Schema(嵌套对象、enum、数组)上的稳定性是选型的主要理由之一。

落地仍建议两道闸:json.loads + 同一份 JSON Schema 再校验;业务不变量(如金额合计)自己写代码验。样例可丢进 JSON 工具箱本地 Diff,不上传服务器。

Function Calling / Tool Calling

Google 文档里的 Function Calling 与 OpenAI Tool Calling 同构:你在请求里声明 tools(名称、描述、parameters JSON Schema),模型返回 functionCall / tool 消息,宿主执行后再把结果塞回对话。3.7 Flash 的卖点之一是「更 diligently 地思考多步规划与 tool calls」——更少无效重试,Agent 总 token 可能反而下降。

from google import genai
from google.genai import types

get_weather = types.FunctionDeclaration(
    name="get_weather",
    description="返回指定城市的当前天气",
    parameters={
        "type": "object",
        "properties": {
            "city": {"type": "string", "description": "城市名,如 Shanghai"}
        },
        "required": ["city"],
    },
)

client = genai.Client()
response = client.models.generate_content(
    model="gemini-3.7-flash",
    contents="上海现在适合跑步吗?",
    config=types.GenerateContentConfig(
        tools=[types.Tool(function_declarations=[get_weather])],
    ),
)

for part in response.candidates[0].content.parts:
    if part.function_call:
        print(part.function_call.name, part.function_call.args)

典型 Agent 链路:Tool Calling 查数据 → Structured Output 汇总成固定 JSON → MCP 暴露给 Cursor / Claude Desktop。工具 Schema 与 MCP inputSchema 应对齐,否则模型填参通过、Server 端却拒收。系列文章《JSON Schema、Function Calling 与 MCP 演进》有完整时间线。

Structured OutputFunction Calling
JSON 用途最终答复 / 抽取结果工具参数
谁执行无副作用,直接 parse你的代码 / MCP Server
3.7 Flash 优势长 Schema、嵌套字段更稳多步 tool 计划更连贯

与 3.6 Flash 及竞品怎么选

场景建议
新 Agent / 编码项目(2026 Q3–Q4)默认 gemini-3.7-flash,吃 Intro 价
已稳定跑在 3.6 Flash、无 tool/JSON 痛点可暂缓迁移,但 2027 年两者标准价对齐
超长推理、最低幻觉考虑 Gemini Pro / 带 thinking 的型号,单价更高
Structured Output 为主、少工具3.7 Flash + 本地 Schema 校验足够
大量 MCP 工具、多轮 tool loop3.7 Flash + Context Caching 固定工具 Schema

第三方 benchmark 常列 Claude Sonnet、GPT 系列对照;Google 称 3.7 Flash Intro 价低于多家同档 API。选型时除标价外,应看你的任务上 Structured Output 通过率与 tool 重试次数——3.7 的改进主要体现在减少后者。

落地建议

  1. 模型字符串全局替换:配置中心把 default model 改为 gemini-3.7-flash,保留 3.6 作 fallback 一周对比错误率。
  2. JSON 双轨测试:同一批 Prompt,分别测 Structured Output 与 Tool Calling,用 JSON 校验器对照 Schema。
  3. 缓存 System + tools:Agent 系统提示与 tools 列表走 Context Caching,降 input 成本。
  4. 按 2027 标准价做预算:Intro 结束 output 翻倍,避免 Q1 2027 账单惊 surprise。
  5. 密钥与 Schema 分离:Schema 可开源;API Key 只在服务端,MCP 走 OAuth 或短期 token。

常见问题 FAQ

Gemini 3.7 Flash 的 API 模型 ID 是什么?

在 Gemini API 与 google-genai SDK 中使用 gemini-3.7-flash。REST 路径为 models/gemini-3.7-flash:generateContent。Vertex AI 控制台可能显示带区域或版本后缀,以当前项目文档为准。

Intro 价什么时候结束?之后多少钱?

Intro 价 $0.75 输入 / $3.75 输出(每百万 token)有效至 2026 年 12 月 31 日。2027 年 1 月 1 日起标准价为 $1.50 / $7.50,Context Caching 读取价也相应翻倍。

3.7 Flash 支持 JSON Schema 结构化输出吗?

支持。与 2.5 / 3.6 相同,使用 response_mime_type=application/json 配合 response_schema 或 response_json_schema。复杂抽取与 Agent 间载荷建议始终带 Schema,不要只靠提示词「请输出 JSON」。

Function Calling 和 Structured Output 能只开一个吗?

可以,但职责不同。只开 Structured Output 适合分类、填表、抽取;只开 Function Calling 适合查库、发邮件、调 MCP。生产 Agent 经常一轮对话里两者都用:工具拿事实,Structured Output 给下游固定形状。

从 gemini-3.6-flash 迁移要改代码吗?

通常只需改 model 字符串与回归测试。tools、responseJsonSchema、多模态 parts 格式与 3.6 兼容。若你依赖特定 thinking 或温度行为,建议在 staging 对比同一套 JSON 样例再切流量。

如何本地验证模型吐出的 JSON?

把 responseJsonSchema 与模型输出存成 JSON 文件,用 JSON 工具箱在浏览器本地做语法校验与结构 Diff,数据不上传。这与上线前测 REST 契约是同一思路。

总结

Gemini 3.7 Flash 是 2026 年面向编码与 Agent 的主力 Flash:模型 ID gemini-3.7-flash,Intro 价至 2026 年底为 $0.75 / $3.75(每百万 input / output token),2027 年起翻倍。API 层与旧 Flash 一致,升级点在于多步 tool 调用与 Structured Output 的实战稳定性。

建议你今天做两件事:把 staging 默认模型改成 3.7 Flash 跑一轮 JSON + Tool 回归;用 JSON 工具箱对照 Schema 与样例输出。Structured Output 与 Function Calling 的配置细节仍见系列专题,不要混成一种 API。