Short answer: as of 28 August 2026, Gemini 3.8 Flash is not publicly released. There is no model card, no documented gemini-3.8-flash API, and Google DeepMind’s Flash docs still stop at Gemini 3.7 Flash, which went GA on 13 August. The same week, Business Insider reported that employees are already trying a “Gemini 3.8 Flash Preview” on the internal coding platform Jetski, and posts on X talk about a public drop “in a few weeks.” Those are leads, not a launch note.
The previous piece, Gemini 3.7 Flash pricing, API, and JSON, covers the Flash model you can actually call. Structured JSON with the Gemini API and OpenAI vs Gemini Structured Output cover the Schema contract. This article answers three questions only: when 3.8 might ship, where quality is likely to move, and whether the JSON / Function Calling request shape will change.
Has it shipped? Status on 2026-08-28
Treat this as a checklist:
- Public product: not shipped. AI Studio / Gemini API / Vertex AI do not list a callable 3.8 Flash.
- Model ID: no documented
gemini-3.8-flash. Putting that string in production will 404. - Pricing: not disclosed. Tables that claim “3.8 at half of 3.7” have no official source.
- Current Flash workhorse: still
gemini-3.7-flash, intro priced at $0.75 / $3.75 per million input / output tokens through 2026-12-31, then $1.50 / $7.50 from 2027-01-01.
So the honest answer to “when does it launch?” is: Google has not set a public date; given the last two Flash drops, “weeks, not next year” is the better prior. The rest of this piece separates rumor from cadence so an internal Preview is not mistaken for GA.
August news and how strong the evidence is
Rank by evidence, not by social-media heat.
| Date | Claim | Strength |
|---|---|---|
| 2026-08-13 | Gemini 3.7 Flash GA, official model card and intro pricing | Confirmed |
| 2026-08-10 | A Tencent paper names “Gemini 3.8 Flash” as one LLM judge (three days before 3.7 GA) | Weak: typo or early internal name |
| Q2 2026 earnings | Sundar Pichai said Flash should land close to a monthly cadence | Confirmed (cadence, not 3.8 itself) |
| 2026-08-27 | Business Insider: staff using “Gemini 3.8 Flash Preview” on Jetski; Google declined to comment | Medium-high: named reporting, still not a launch |
| 2026-08-27 | X posts: already internal, maybe public in weeks, Fable 5-level quality | Weak: no model card, no API |
| Same period | Guesses that an anonymous Arena AI model is 3.8 | Weak: identity unverified |
The Business Insider details matter: the internal label is Gemini 3.8 Flash Preview, the surface is Jetski, and one tester said it already felt “noticeably better than 3.7 Flash” while stressing it was too early for a full review. That matches the 3.6 → 3.7 pattern — dogfood inside Google, then GA a few weeks later.
The louder X claims (“months of partner testing,” “Fable 5 quality at Flash price”) still have no DeepMind model card or API changelog behind them. The Tencent paper mentions “3.8” once. The two boring explanations remain best: a version typo, or an eval team hitting an unpublished endpoint. One word in one paper is not a ship calendar.
How to read the launch window
Google has not published a 3.8 date. The only intervals you can use are the Flash line’s own:
- Gemini 3.6 Flash: late July 2026 GA (same wave as 3.5 Flash-Lite).
- Gemini 3.7 Flash: GA on 2026-08-13, about three weeks after 3.6.
- Pichai on the Q2 call: Flash should stay close to monthly.
- A 3.8 Preview was on an internal platform by 27 August — employee trial, not “it ships next week.”
If the three-to-four-week rhythm holds, a public Developer Preview or GA is more likely in early to mid September 2026 than as a sidecar to Gemini 4. Gemini 4 pre-training was confirmed to have started on 2026-07-21 with no launch window; “3.8 before Gemini 4” is a low bar because the flagship is slower by design.
Also price in “it never ships as 3.8”: an internal Preview can fold into a quiet 3.7 refresh or come back under another name. Do not make 3.8 a milestone dependency. Keep the model string in config and leave traffic on 3.7.
3.7 Flash vs the rumored 3.8
| Item | Gemini 3.7 Flash (confirmed) | Gemini 3.8 Flash (unconfirmed) |
|---|---|---|
| Status | GA 2026-08-13 | Internal Preview reported; no public API |
| Model ID | gemini-3.7-flash | Undocumented; do not ship it |
| Intro price | $0.75 / $3.75 per million tokens through 2026-12-31 | Not disclosed |
| 2027 standard | $1.50 / $7.50 | Not disclosed |
| Context | ~1M input / ~64k output | Not disclosed; Flash usually keeps the window |
| DeepSWE v1.1 | 65.3% (3.6: 49.0%) | No official score |
| Positioning | Workhorse for coding and agents | Rumor: closer to flagship, still Flash-priced |
The empty right-hand column is the point: the only Flash you can take to an architecture review today is 3.7. Its intro price is already an order of magnitude below Claude Fable 5 at $10 / $50. Even if 3.8 closes some of that quality gap, the default model is still chosen by cost-per-task times retries — not by an anonymous Arena score.
Performance outlook (from 3.6→3.7, not invented scores)
With no public bench, do not invent 3.8 percentages. What you can extrapolate is the direction 3.6 → 3.7 already paid out, plus the internal “noticeably better than 3.7” remark:
- Agents / multi-step tools: 3.7 vs 3.6 moved DeepSWE from 49.0% to 65.3% and AutomationBench from 17.0% to 30.4%. If the next hop stays a workhorse, the prize is fewer wasted tool retries, not another context-window headline.
- Coding: FrontierCode 1.1 Main moved from 34.4% to 43.6%. Putting the Preview on Jetski (a coding surface) is a signal that Flash remains the default engine for writing and fixing code, not a chat skin.
- Long documents and extraction: knowledge-work scores such as GDP.pdf moved from 22.0% to 34.0%. For JSON extraction that usually means fewer drifted nested fields, enums, and array lengths — wait for GA and rerun the same Schema; do not write it into an SLA early.
- Latency and price: Flash’s product promise is cheap, fast, and callable thousands of times a day. If 3.8 broke that (a Pro in disguise), teams would stay on 3.7 intro pricing. The safer forecast is Flash-tier price, quality aimed at 3.7’s weak spots.
Comparing “who is smarter” against Claude Fable 5 ($10 / $50) is the wrong frame. Comparing “fewer tool round-trips on the same Schema” against 3.7 is the right one. Treat every forecast as a hypothesis until a model card exists.
API JSON preview
This is the layer the series actually cares about. From Gemini 2.5 through 3.7, Structured Output and Function Calling kept almost the same request shape: the model string and behavior changed, the field names did not. If 3.8 is a 3.7 increment, the stable preview is compatibility, not a new JSON API.
Structured Output: keep the Schema, do not bet on a new switch
Today’s 3.7 pattern will almost certainly still apply after 3.8 GA: response_mime_type=application/json plus response_schema (Pydantic) or response_json_schema (a JSON Schema dict). Constrained decoding blocks illegal tokens at generation time — the same path as the Gemini JSON guide and Prompt to Structured Output.
from google import genai
from pydantic import BaseModel, Field
# Do not point production at an undocumented 3.8 ID before GA.
# After GA you typically change this one string; Schema and checks stay put.
MODEL = "gemini-3.7-flash"
class ExtractedTask(BaseModel):
title: str
priority: str = Field(description="low | medium | high")
due_date: str | None = None
client = genai.Client()
response = client.models.generate_content(
model=MODEL,
contents="Extract the action item: finish the budget review by next Friday, high priority.",
config={
"response_mime_type": "application/json",
"response_schema": ExtractedTask,
},
)
print(response.parsed)
Reasonable hopes for 3.8: a higher first-pass rate on nested objects, optional fields, enums, and arrays; less leaked chain-of-thought when thinking and JSON are both on. Do not hope for full JSON Schema 2020-12 (Gemini’s subset limits on additionalProperties, prefixItems, and friends are in the comparison piece), or for collapsing “final answer JSON” and “tool argument JSON” into one knob.
Function Calling: the tool Schema stays the contract
The tool side is still tools plus a parameters JSON Schema, a functionCall from the model, then host execution and a follow-up turn. 3.7’s pitch was more diligent multi-step plans; if the 3.8 dogfood really is “noticeably better,” look first for fewer wrong tools, fewer dropped required fields, fewer empty loops on the same tool. MCP inputSchema still has to match the declaration, or the model “succeeds” and the server rejects — that is independent of 3.8. See Tool Calling → MCP data flow.
from google import genai
from google.genai import types
lookup_order = types.FunctionDeclaration(
name="lookup_order",
description="Return shipment status for an order id",
parameters={
"type": "object",
"properties": {
"order_id": {"type": "string", "description": "Order id, e.g. GW-20260828"}
},
"required": ["order_id"],
},
)
client = genai.Client()
response = client.models.generate_content(
model="gemini-3.7-flash",
contents="Has GW-20260828 shipped yet?",
config=types.GenerateContentConfig(
tools=[types.Tool(function_declarations=[lookup_order])],
),
)
| JSON job | Today (3.7) | 3.8 preview |
|---|---|---|
| Final answer / extraction | responseMimeType + Schema | Field names likely unchanged; watch first-pass rate |
| Tool arguments | tools[].parameters | Fewer retries and type drifts; do not loosen the contract |
| MCP | inputSchema aligned with the tool | Independent of the model bump; keep local validation |
Keep two gates: json.loads plus the same JSON Schema, then your own invariants (money, enums, timezones). Drop sample outputs into the JSON toolbox for a local Diff — no upload. 3.7 and a future 3.8 should share the same fixtures, or you cannot tell whether it actually got more stable.
What to do now
- Keep production on 3.7 Flash:
gemini-3.7-flashat intro pricing, default model in config, not scattered literals. - Do not reserve a fictional 3.8 ID: an internal Preview name is not an SDK argument. Add a fallback row only after AI Studio / the DeepMind card lists it.
- Finish the JSON fixtures first: the same prompts through Structured Output and Tool Calling, checked locally against Schema. When 3.8 GAs, rerun once with a new model string.
- Cache System + tools: put the agent system prompt and tool list on Context Caching. 3.7 intro cache reads are about $0.075 / 1M tokens.
- Budget at 2027 standard rates: even if 3.8 ships with a new intro price, 3.7 itself doubles on 2027-01-01. Do not let a Q1 invoice depend on a rumored cheaper SKU.
- Watch official channels only: the DeepMind Flash card, Gemini API changelog, Vertex model list. X leaks and anonymous Arena models are not a schedule.
FAQ
When will Gemini 3.8 Flash be released?
Google has not published a date. Given the roughly three-week 3.6→3.7 gap and the CEO’s near-monthly Flash cadence, a public window in early to mid September 2026 is the better prior if the internal Preview goes well. That is not a commitment and must not be a project deadline.
Can I call gemini-3.8-flash today?
No. As of 28 August 2026 there is no public API, SDK model ID, or DeepMind model card. Use gemini-3.7-flash in production.
Does an internal Preview mean GA is imminent?
No. Business Insider confirmed staff saw Gemini 3.8 Flash Preview on Jetski; Google declined to comment. Dogfood can GA in weeks, or it can be renamed or folded back into 3.7. Read it as “in progress,” not “scheduled.”
Will 3.8 be stronger and cheaper than 3.7?
There is no official bench or price list. 3.7’s gains over 3.6 sat in multi-step tools and coding; one internal tester called the 3.8 Preview “noticeably better than 3.7,” which does not generalize. Pricing is more likely to stay Flash-tier than to jump to flagship list prices.
Will Structured Output / Function Calling change API on 3.8?
The 2.5→3.7 habit is: new model ID, same request fields — response_mime_type, response_schema / response_json_schema, tools.parameters. Treat a Flash increment as compatible first, then regression-test the same Schema on GA day.
Should I wait for 3.8 before adopting Structured Output?
No. The JSON Schema contract, a second validation pass, and error replay do not depend on a Flash minor version. Get the pipeline green on 3.7; 3.8 is a model-string swap.
Takeaways
Gemini 3.8 Flash has not launched. On 28 August 2026 the confirmed facts are: 3.7 Flash is in production, intro pricing still holds, and an internal 3.8 Preview is already circulating. Flash cadence over the last two months points to “weeks around September,” not “wait for Gemini 4.” Performance and price have no model card yet; the only honest forecast follows the 3.6→3.7 agent and coding vector.
The JSON story on the API is unchanged: Structured Output owns the final shape, Function Calling owns tool arguments, both use JSON Schema. Fixtures and local validation today beat a leak calendar. When 3.8 appears on the official model list, rerun the same Schema once.