Why AI Agents Need JSON Even More After the OpenAI Agents API: Harness, Tool Calling, JSON Schema

As of 18 September 2026: the Agents API public beta hosts the Codex harness. Once you no longer write the loop, almost everything you still own is JSON — tool Schema, arguments, tool_result, MCP, session events. A loose contract just lets the hosted loop run bad parameters faster.

Up front: the Agents API hosts the loop. It does not host the contract. On 10 September 2026, OpenAI put the Agent Harness that drives Codex in front of developers as a public beta. Model scheduling, context compaction, subagents, and sandbox lifetime moved out of your process and into beta.agents.sessions. What you still hold is almost all JSON: the JSON Schema on each function tool, arguments, the string in tool_result, MCP inputSchema, and the session event stream. A hosted harness running the loop is not the same as someone else validating your fields. Loosen the contract, and the hosted loop just runs bad parameters more often.

This piece is written as of 18 September 2026. This site already has why agents cannot live without JSON, why Tool Calling depends on JSON Schema, whether JSON Schema becomes the standard contract, MCP / Skills / Tools / Subagents, and what MCP is. This article only answers why JSON matters more after the Agents API, not less.

What actually shipped on 10 September

OpenAI’s line is: the same harness and infrastructure that power Codex, as a hosted cloud agent for developers. The public docs put it under the beta.agents namespace; requests carry OpenAI-Beta: agents=v1. The harness has no separate fee. You pay for model tokens, tools, and sandbox time.

Creating a session means submitting one JSON document: model, instructions, tool list, environment, input. Official samples use gpt-6-astra. Tools can be MCP, custom functions, or built-in retrieval. The environment can be none, openai_hosted, or a sandbox you bring — Blaxel, Cloudflare, Daytona, E2B, Modal, Vercel, and similar. Multi-agent is a flag: multi_agent.enabled plus max_concurrent_subagents.

This is not another chat endpoint that says “please output JSON.” The Responses API is still there. The Agents SDK is still there. What the Agents API takes is the loop itself: who picks the next hop, when context is compacted, when a subagent is spawned. Field names may still move during the beta. The split is already stable: OpenAI runs the harness; you supply the tool contract and the business result.

Three doors: Responses, Agents SDK, Agents API

As of September 2026, OpenAI leaves three ways to build an agent side by side. Before you mix them, ask where the loop runs:

DoorWhere the loop runsWhere state livesWhat you still write
Responses APIYour appHistory you assemble / ConversationsModel calls, tool returns, the whole loop
Agents SDKYour processSDK sessions plus your storageApprovals, deploy, and you can still change the loop
Agents APIOpenAI’s hosted Codex harnessServer-side session / turn / itemTool definitions, function results, environment choice; you cannot change the loop

One-shot completion still belongs on Responses. If you need to own approvals and persistence, use the SDK. If you want multi-day work, compaction, subagents, and a sandbox operated for you, use the Agents API. All three still describe tool parameters with JSON Schema. The difference: the first two still let you patch the loop; the third only lets you patch the contract and the return payload.

What an Agent Harness is — and what it will not sign

A harness is the runtime between the model and side effects: read events, pick tools, feed results, compact context, keep a long job alive. The Codex harness is open source. The Agents API is OpenAI operating that same logic and versioning it with the models. The launch note calls out automatic compaction, Tool search, Programmatic Tool Calling, and parallel subagents.

It will not sign any of these:

  • whether a customer_id must exist, or must be a UUID;
  • whether your function may accept extra keys;
  • whether an MCP server’s inputSchema is tight or slack;
  • whether the output you send back is an object, a string, or a chat paragraph.

Those remain JSON Schema plus a check you run yourself. A hosted harness raises how long a loop can run and how much it can parallelize. It does not raise whether this hop’s parameters are legal. Treating those as the same thing is the first mix-up this article splits apart.

Why a hosted loop means more JSON hops

When you write the loop, bad JSON usually dies on your side: parse fails, fields do not match, you stop. Once the loop is hosted, failure is deferred, copied, and sent down more channels:

HopPayloadWho produces itWho must validate
Create sessionagent / tools / environment JSONYour appYou, before you submit
Function definitionJSON Schema (parameters)Your appYou: tighten required / additionalProperties
Model issues a callarguments objectHosted harness + modelYou: validate again before executing
Return a resulttool_result.output stringYour appYou: stringify a legal value
MCPJSON-RPC + inputSchemaServer / harnessThe server and your allow list
Event streamagent.session.* JSON eventsThe hosted serviceYou: branch on type; do not parse it as chat prose

Add Tool search loading definitions on demand, Programmatic Tool Calling chaining calls in code, and subagents each carrying their own context — one user task now makes more JSON round-trips than a single Function Calling hop. Hosting hides those hops. Hidden is not the same as optional to validate. For a hop-by-hop map see from Tool Calling to MCP.

Tool Calling: function tools are still JSON Schema

Function tools on the Agents API reuse the Responses API shape. What you put on agent.tools is not a paragraph. It is a name, a description, and a JSON Schema:

{
  "type": "function",
  "name": "get_customer",
  "description": "Look up a customer by ID.",
  "parameters": {
    "type": "object",
    "properties": { "customer_id": { "type": "string" } },
    "required": ["customer_id"],
    "additionalProperties": false
  }
}

The official sample fills required and sets additionalProperties to false. That is not a formatting habit. If the agent may add one extra key, that key can become a path, a SQL fragment, or a delete. The schema is the contract the model sees while decoding, and the contract you should run again before executing. For strict mode, ajv, and a second-pass pipeline see why Tool Calling depends on JSON Schema.

The description still helps the model pick a tool. It does not replace types, enums, and required fields. The smarter the harness, the more willingly it picks a “close enough” tool from a long list. Close enough is what the schema is there to reject.

requires_action and tool_result: the return path is JSON too

When the model needs your function, the session pauses on agent.session.requires_action. Pending work lives in required_actions. A function_call item in history is not enough on its own. A typical pending call looks like this:

{
  "type": "function_call",
  "turn_id": "turn_123",
  "call_id": "call_123",
  "name": "get_customer",
  "arguments": { "customer_id": "123" }
}

The docs present arguments as an object. Do not wrap it in a chat reply and scrape it with JSON.parse — that is the wrong channel, covered in why JSON.parse fails. Validate the object against the same schema, run the function, then post agent.session.input.tool_result to the session events endpoint with the original turn_id and call_id.

On success, success: true and output as a string or a supported content array. Objects go through JSON.stringify first. On failure, success: false and an error the model can read. Do not send stacks, secrets, or a whole database row back. If the process dies after the function ran but before the result reached OpenAI, key the side effect by session / turn / call and stay idempotent: on restart, read pending actions before running it again.

Function handlers always run in your application, even when the session has a sandbox. The harness will not execute get_customer for you. If you are offline, the hop stays blocked. That is one of the few synchronous points in a hosted loop that still belongs entirely to you — and the hop where the JSON has to be right.

Tool search and Programmatic Tool Calling

A large tool list burns tokens and breaks cache if every schema sits in context. The Agents API loads functions eagerly by default. Rare ones can set defer_loading: true, with {"type": "tool_search"} on agent.tools. The model finds the definition, then calls it. That adds a hop where “the definition is JSON too”: the schema that search returns must match the function you actually implemented. Do not advertise a wide contract and execute a narrow one.

Programmatic Tool Calling lets supported models write a short program that runs eligible tools in parallel or in a chain, then brings only filtered results back into context. That cuts the cost of filling the window on every hop. It raises the requirement that the intermediate JSON is legal. If types drift in the middle, later filters and merges go wrong inside a harness you cannot see. The SDK already has a fix that encodes structured errors as JSON. That path eats schema, not prose.

MCP and subagents: more schemas, more JSON

Add an MCP server to agent.tools and the harness discovers tools, calls them, and feeds results back. Unlike functions, those calls do not pass through your application. HTTP connects from OpenAI by default; you can also connect from the environment, or spawn stdio inside the sandbox. What you still control is allowed_tools, whether a failed init fails the turn (required: true), and how tight the server’s own inputSchema is.

MCP messages are still JSON-RPC. A slack schema means the hosted harness will fire more requests you never see. That is not “the protocol made you safe.” It is “the loop moved farther away.” For the protocol layers see what MCP is; for the boundary with Skills and Subagents see the 2026 agent stack.

Each subagent keeps its own context; the parent merges. Parallel work cuts latency and also fans out many arguments objects. If the merge is still a schema-free essay, you only postponed “parse the chat” to the last hop. Conclusions that enter a program should still use Structured Output or a result schema you define — not another scrape of prose. See what Structured Output is.

Four things you still validate locally

After the harness is hosted, the list does not get shorter. It gets narrower:

  1. Tool schemas. Fill required, set additionalProperties: false, tighten enums. Do not lean on the description to stop side effects.
  2. Arguments before execute. Even if the vendor already applied the schema, run the same document again in your process. Wrong types, missing fields, extra keys stop here.
  3. The output you send back. Make legal JSON, then stringify. Errors go out as success: false. Do not hand the model a raw internal exception.
  4. Keep events and chat on separate channels. Branch on event.type. Do not treat a whole SSE stream as one JSON value. Structured answers to the user go through Structured Output, not JSON.parse on an assistant sentence.

On the security side: a string inside arguments can be an injection, not “the type matched, so run it.” See malicious JSON and prompt injection. Whether JSON Schema becomes the cross-vendor contract is the standard-contract piece — the Agents API does not weaken that claim. It pushes the claim onto the only layer you can still change.

Inspect the contract with local JSON tools

Before you hand work to a hosted session, look at three texts in the browser: the tool schema, a sample arguments object, and the output you plan to send back.

  • JSON validator — is the grammar legal; if you have a schema, check fields, required, and extra keys together.
  • JSON formatter — expand a one-line tool_result and see whether you serialized a whole database row.
  • JSON Diff — compare the arguments the model sent with the smallest object the schema allows.

Nothing leaves the browser. It is the right place to sit a failed required_actions payload, a parameters document, and one stringified result next to each other. Stabilize the contract, then let the hosted harness run for days.

FAQ

Does the Agents API mean I can stop writing JSON Schema?

The opposite. Once the loop is hosted, the schema is the main contract you still hold. Function parameters, MCP inputSchema, and the output you return are still JSON.

How should I choose among Agents API, Agents SDK, and Responses?

One-shot calls go on Responses. If you need to own the loop, approvals, and storage, use the SDK. If you want long jobs, compaction, subagents, and a sandbox operated by OpenAI, use the Agents API. All three still take JSON Schema for tool parameters.

Arguments are already an object. Do I still call JSON.parse?

Do not parse the surrounding chat again. Treat arguments as an object, as the docs do, and validate it with the same JSON Schema. Scraping arguments out of prose is the wrong channel.

Why must tool_result be stringified?

The docs want output as a string or a supported content array. Make legal JSON, then stringify, so you do not mix a second encoding with “looks like an object, is actually a string.”

Do MCP tools pass through my application?

Not by default. The harness talks to the server. What you tighten is the server’s own inputSchema, allowed_tools, and any approval for irreversible actions inside that server.

Will field names change during the beta?

They might. This article follows public docs as of 18 September 2026. The split will not: the harness runs the loop; you provide the JSON contract. If a field is renamed, the duty to validate stays on your side.

Summary

The Agents API cuts the work of “how to run an agent to completion.” It raises the weight of “every JSON hop must be right.” What shipped on 10 September is the Codex harness: sessions, compaction, tool search, programmatic calls, subagents, sandboxes. It will not check what a customer_id should look like, and it will not turn your tool_result into a legal string for you.

Wiring an agent into a program in 2026 still follows the same order: tools on JSON Schema, final answers on Structured Output, chat prose is not an API. What changed is that once the loop is hosted, the only place you can still patch is the contract. Check the schema, the arguments, and the return payload in a local validator first, then hand the work to a hosted session. Models will change. The harness will take new versions. Your field contract should not loosen with them.