After Claude Sonnet 5.5, the Price List Did Not Move — Why JSON Extraction Still 400s. From between_tools to JSON Schema

As of 29 September 2026: Sonnet 5.5 (28 Sep) keeps $2/$10. thinking.disabled is a 400; use between_tools. Tool-less JSON extraction still needs adaptive thinking plus Structured Output — not thinking off.

Up front: Sonnet 5.5 keeps every line of the price list and changes the request contract. On 28 September 2026 Anthropic shipped Claude Sonnet 5.5 (claude-sonnet-5-5). Input / output stay $2 / $10; cache reads stay $0.20. Anthropic says output is 30%+ faster and most work costs about 30% less — that saving is tokens per task, not the sticker. What breaks production is the old switches copied from Sonnet 5: thinking: {"type": "disabled"} returns 400; tool_choice values any and tool return 400; a non-default temperature, top_p, or top_k is also 400. The quieter trap is JSON extraction: on a request with no tools, between_tools means the model does not think first, and Structured Output still goes wrong. Anthropic’s prompting page files that case under “Reasoning tasks with JSON output.”

Written as of 29 September 2026, against the model page, What’s new, and the migration guide still current that day. This site already has why Opus 5.5 will not let you force JSON with tool_choice (24 September). This piece does not rewrite those four failures. It only answers what a Sonnet 5 → 5.5 JSON pipeline has to change on top. Haiku 5.5 is still “in the coming weeks.” This article does not borrow that date.

What shipped on 28 September

Claude Sonnet 5.5 is the second model in the Claude 5.5 family. Official positioning: the speed-and-intelligence middle; everyday coding, bug fixes, documents and spreadsheets. The ID is claude-sonnet-5-5 on the Claude API, Google Cloud, Microsoft Foundry, and Claude Platform on AWS; Bedrock uses anthropic.claude-sonnet-5-5. Retirement is not sooner than 28 September 2027. The tokenizer matches Sonnet 5, so the same text produces the same token counts.

Prices per million tokens: $2 in, $10 out, $2.50 for a 5-minute cache write, $4 for a 1-hour cache write, $0.20 for a cache read. Batch is half price. Context is 1M; synchronous max output is 128K. Message Batches can reach 300K with output-300k-2026-03-24. The minimum cacheable prompt drops from Sonnet 5’s 1,024 tokens to 512. Default effort is high — Opus 5.5 defaults to medium. Omit effort and the request runs at high. That is not silent compatibility with Sonnet 5.

The launch frames Sonnet 5.5 as the faster, cheaper complement to Opus 5.5. This article does not pick a model from a leaderboard. What matters is that the price list stayed put while thinking, tool_choice, and the JSON-extraction contract moved.

Five hard failures, one table

The docs list the requests that become 400 on 5.5. The first three share a root with Opus 5.5 and Fable 5.1. The replacement on the first row is different:

Old request (Sonnet 5)On 5.5What to send instead
thinking: {"type": "disabled"} or enabled plus budget_tokens400 invalid_request_error pointing at between_toolsTurn off up-front thinking with {"type": "between_tools"}; for reasoning, omit thinking or send adaptive
tool_choice: {"type": "any"} or {"type": "tool", "name": "..."}400; the token-counting endpoint applies the same checkauto (or none) plus strict: true, or Structured Output
Replay a thinking block after editing system / tools / an earlier message400 by default on accounts created after 31 August 2026Keep the conversation append-only; change instructions with a mid-conversation system message
computer_20251124 on the Claude API or Google Cloud400computer_toolset_20260801; Bedrock still accepts the old tool
An advisor set to Opus 4.8 / 4.7 or Sonnet 5400Use an advisor 5.5 accepts (including Opus 5.5, Fable / Mythos 5.1, or 5.5 itself)

One more change does not fail the request but goes quiet: longer notes between tool calls arrive as thinking blocks. At the default display: "omitted" the text is empty. A UI that streamed those sentences as a progress bar goes silent. For progress under adaptive thinking, set thinking.display (updates needs a beta header). between_tools returns the summary text, and that type rejects a display field.

Sampling is no longer a knob. A non-default temperature, top_p, or top_k is a 400. Requests that used a lower temperature to “stabilize JSON” now have to stabilize the schema instead.

disabled is gone; between_tools is the floor

Sonnet 5 turned thinking off with disabled. On 5.5 that line is a 400 whose message points at between_tools. The lowest setting looks like this:

{
  "model": "claude-sonnet-5-5",
  "thinking": { "type": "between_tools" },
  "output_config": { "effort": "high" }
}

between_tools needs no beta header and is accepted on every platform that offers the model. It turns off up-front thinking only. Longer progress notes between tool calls still return as thinking blocks. Echo them unmodified; the model gets the full note, not the summary. With no tools, the response is usually text only — the old disabled shape.

The limits are strict. It is accepted at low / medium / high. Pair it with xhigh or max and you get 400. Add display, budget_tokens, or block_binding and you also get 400. A mid-conversation change to output_config.effort is a 400 — per-turn effort needs adaptive thinking. When server-side fallback lands on Sonnet 5, between_tools is translated to disabled there.

So 5.5 does not force thinking to stay at full depth. It removes the old disabled switch. Once that switch is gone, people who want thinking off must send a new type. People who want stable JSON often should not turn it off — see the next section.

For tool-less JSON extraction, do not turn thinking off

Anthropic’s prompting page has its own section for this: give Sonnet 5.5 a JSON task that needs a few steps (totals, a rule, a ranking) and at low / medium it often answers without thinking first. With Structured Output, the visible text can only be JSON, so the work can happen only in thinking. Skip thinking and accuracy drops.

The documented order is:

  1. Use Structured Output when you can. The body is then an object that matches the schema. You do not JSON.parse chat prose. See what Structured Output is.
  2. Use adaptive thinking, not between_tools. On a request with no tools, between_tools means no thinking first. A “think before you answer” line in the prompt has no effect there. Splitting “answer, then JSON” into two requests scored well in their tests and cost too much latency and tokens to be a default.
  3. End the system prompt with Think the problem through before you answer. At high, accuracy approaches xhigh for a modest extra in output tokens. At low / medium it also rises, at a larger token cost.
  4. Leave room in max_tokens for thinking. With Structured Output at low / medium, the model occasionally thinks until the cap. Treat a response whose stop_reason is max_tokens as a failure even if the text is valid JSON, and retry.

Without Structured Output, the model often works in the prose and puts the JSON last. Parsing the whole reply fails. The documented fallback: read only text blocks, try a parse from each { or [, and keep the last complete value. Do not take the span from the first { to the last } — a draft can sit in the middle. See why JSON.parse fails. That is a fallback, not a contract.

Forced tools are a 400, same as Opus 5.5

The 5.5 family shares the same error:

tool_choice: type "tool" and "any" are not supported for this model.

The replacement is still: keep tool_choice on auto, set strict: true on the tool, and pin parameters with JSON Schema. Put the final answer on Structured Output. If you want the model to call a tool instead of answering in prose, say in the prompt when the tool applies. The prompt influences which tool auto picks. It does not replace the schema. The long version is in why Opus 5.5 will not let you force JSON with tool_choice and why Tool Calling depends on JSON Schema.

5.5 adds one tolerance note: it occasionally calls a tool with the wrong letter case, or a parameter under a slightly different name. Do not treat that as fatal. Accept the call when the match is unique, or return is_error: true with the exact name. Keep the schema strict. Name matching can be one notch looser.

The price list did not move; default effort is high

List prices match Sonnet 5. The official “about 30% cheaper” is fewer tokens per task, not the menu. Independent write-ups already show the other side: when a task does not bound its own output, the new model can cost more. Run the same schema through an extraction sweep before you change routing.

Effort levels are recalibrated. The same high is not the same amount of thinking as on Sonnet 5. The docs start at high for everyday work; medium for well-specified agentic coding and tool loops; medium or low for chat and latency-sensitive work. Changing top-level effort busts the prompt cache. Per-turn changes need per-message effort (beta) under adaptive thinking. between_tools will not let effort move mid-conversation.

Do not mix defaults with Opus 5.5: that model defaults to medium; Sonnet 5.5 defaults to high. Swap only the model string and both the bill and the latency jump. Thinking blocks also do not travel freely. 5.5 can read blocks from Sonnet 5, Opus 4.8, Haiku 4.5, and earlier models. It does not read Opus 5, Opus 5.5, Fable, or Mythos. No other model reads Sonnet 5.5 blocks. Move a conversation from Opus 5.5 onto Sonnet 5.5 and the API drops the reasoning. The request still returns 200; dropped blocks are not billed.

Temperature, cache, computer use, advisor

A non-default temperature / top_p / top_k is a 400. Pipelines that “tightened JSON” with sampling now have to tighten the schema. The cache floor is 512 tokens, so short system prompts cache more easily. Changing top-level effort still invalidates that cache.

On the Claude API and Google Cloud, computer use only accepts computer_toolset_20260801. Bedrock still accepts computer_20251124. The advisor tool (beta) will not let a 5.5 executor pair with Opus 4.8 / 4.7 or Sonnet 5. Advisors that 5.5 accepts return encrypted advisor_redacted_result blocks; the client cannot read the advice text.

To change a tool schema mid-conversation, inline-tools-2026-09-15 can carry a full definition in a mid-conversation system message without editing the top-level tools array. That is the same rule as prefix-bound thinking: append history. Compaction (compact-2026-09-04) can swap in a signed summary while keeping thinking valid, under the conditions on Anthropic’s Compaction page.

Five checks before you change the model string

  1. Search for thinking. Remove disabled and budget_tokens. To turn off up-front thinking, send between_tools at high or below. For JSON reasoning tasks, use adaptive thinking.
  2. Search for tool_choice. Replace every any / named tool with auto. Turn on strict and tighten input_schema.
  3. Search for temperature / top_p / top_k. Delete non-defaults. Stabilize JSON with the schema, not with sampling.
  4. Set effort explicitly. Do not inherit the default high. Shared routing with Opus 5.5 has two different defaults.
  5. Replay and cache. On post-31 August accounts, do not edit tools or system in history. Coming from Opus 5.5, expect thinking blocks to be dropped.

Inspect the schema with local JSON tools

Before you set the model to claude-sonnet-5-5, lay out three texts in the browser: the old input_schema, one arguments object you used to get with disabled or a forced tool_choice, and the schema you plan to put on Structured Output.

  • JSON validator — is the grammar legal; if you have a schema, check required fields and extra keys together.
  • JSON formatter — expand a one-line tool definition and see whether additionalProperties is set.
  • JSON Diff — compare a sample from thinking-off with the smallest object the strict schema allows.

Nothing leaves the browser. Stabilize the contract, then change the model string. 5.5 will change effort and reject disabled and the old tool_choice. Your field names and required list should not loosen with it.

FAQ

Can I ship by only changing the model to claude-sonnet-5-5?

Not as a default. If the request still has thinking.disabled, budget_tokens, tool_choice any/tool, a non-default temperature, or computer_20251124 on the Claude API, you get 400.

Is between_tools the new disabled?

Only on requests with no tools. With tools, progress still arrives as thinking blocks. It rejects display / budget_tokens / mid-conversation effort changes, and it cannot pair with xhigh or max. Do not use it for JSON reasoning tasks.

If the price list did not move, why would the bill move?

Default effort is high, and the levels were recalibrated. The official 30% saving is tokens per task, not the sticker. When a task does not bound output, independents have seen a higher bill. Run the same schema yourself.

Is this the same as Opus 5.5 on 22 September?

Forced tools, bound thinking blocks, and the old computer tool share a root. What Sonnet 5.5 adds is disabled→between_tools, default effort high, locked sampling, advisor pairings, and “do not turn thinking off for tool-less JSON.” Migrate the two lines separately.

Without Structured Output, how do I still collect JSON?

Read only text blocks and parse the last complete JSON value. Do not parse the whole reply. Retry when stop_reason is max_tokens. That is a fallback. Prefer a schema when you can.

Has Haiku 5.5 shipped?

As of 29 September 2026 the official line is still “in the coming weeks.” Do not pre-assign these breaking changes to a tier that is not out.

Summary

Sonnet 5.5 holds the price list still and turns disabled into a 400. Copy a Sonnet 5 request and you hard-fail on thinking, tool_choice, and temperature. To turn off up-front thinking, send between_tools. For stable JSON, use adaptive thinking plus Structured Output, or auto plus strict plus JSON Schema. On a tool-less reasoning task, turning thinking off is a measured accuracy drop.

Flatten the schema, a sample arguments object, and the strict contract in a local validator first, then switch to claude-sonnet-5-5. The sticker can stay. The field contract should not loosen with it.