Up front: in 2026, AI agent security is not “will the model say the wrong thing?” It is “can untrusted structured data become a real tool invocation?” An injected chatbot at worst emits a misleading reply. An injected agent at worst sends mail, edits files, queries a database, or triggers payment. JSON is both the contract and the payload. Prompt injection often does not live in the chat box — it lives in ticket fields, web pages, mail bodies, and API responses, all arriving as JSON.
Written as of 14 September 2026. We already covered why Tool Calling depends on JSON Schema, the agent JSON data flow, JSON Schema as contract, and MCP / Skills / Tools / Subagents. This piece only answers what malicious JSON, indirect injection, over-privileged Tool Calling, and loose Schemas each risk — and what the Host should enforce. It is a defensive guide. It does not provide reproducible attack steps or ready-to-use injection text.
One table: chatbot vs agent attack surface
What the user types is only a slice of what an agent ingests. In the 2026 default stack the model also reads tool results, MCP resources, page excerpts, ticket fields, and mail bodies — almost always as JSON, or as strings wrapped in JSON. Industry lists often file this under Prompt Injection, Insecure Output Handling, and Excessive Agency. You do not need the taxonomy first. You need to know who is speaking and who is executing.
| Dimension | Chatbot | Agent that can call tools |
|---|---|---|
| Failure mode | A wrong or steered reply | A real side effect: write, request, mutate data |
| Untrusted input | The current user message | User text + external JSON + tool results + page/mail fields |
| Executor | The model only emits text | Host / MCP Server runs arguments |
| Contract | The prompt (soft) | JSON Schema + Host validation (hard) |
| Smallest fix | Prompt tweaks, refusal policy | Tighter Schema, tool allowlist, independent authz |
One sentence: the model can be fooled; the Host must not execute along with it. The security boundary sits between generating tool_calls and actually running them. Validation, authorization, and audit belong there — not “treat model output as a trusted RPC.”
Malicious JSON: extra fields, type confusion, nested payloads
Malicious JSON is often well-formed and still dangerous. A strict parser rejects trailing commas or comments. A loose parser, string concatenation, or “treat it as text then JSON.parse” just delays the failure. In agent systems the damage starts when the Host already treats the value as an object.
Three shapes to defend first (patterns only — not a how-to):
| Pattern | What it looks like | If the Host does not validate | Defense |
|---|---|---|---|
| Extra fields | A business object with undeclared keys | Merged into config, or forwarded to the next tool | additionalProperties: false; drop unknown keys |
| Type confusion | A number/array arrives as a string or object | Authz branches misfire, or a whole blob becomes an argument | Pin type, enum, format |
| Nested payload | A string field that embeds more JSON or a long brief | Inner text enters the prompt and is read as instructions | Length limits; validate inner JSON; never treat raw text as system speech |
A tool-parameter contract that “works” because it barely constrains anything, then the tightened version. These are defensive samples, not attack material:
{
"type": "object",
"properties": {
"payload": { "type": "object" }
}
}
That Schema is almost no contract: any key, any nesting passes. In production, allowlist fields and forbid extras:
{
"type": "object",
"properties": {
"order_id": { "type": "string", "pattern": "^A-[0-9]{4,8}$" },
"amount": { "type": "number", "minimum": 0, "maximum": 100000 },
"note": { "type": "string", "maxLength": 200 }
},
"required": ["order_id", "amount"],
"additionalProperties": false
}
When you review samples, use this site’s JSON validator against that Schema, and JSON Diff to compare “arguments the model emitted” with “the smallest object you allow.” Extra keys and flipped types are the first alarm.
One more miss: do not merge untrusted objects into internal config or structures that sit on a prototype chain. In a JavaScript Host, unknown keys in a config object can mean more than “one extra field.” The fix is the same: project an allowlisted object from the Schema, then pass that down.
Prompt injection: instructions hidden in structured data
Direct injection is the user talking at the chat box, trying to override the system prompt. Indirect injection matches how 2026 agents actually run: instructions hide in data the model will read later — ticket titles, mail bodies, page paragraphs, PDF excerpts, JSON strings returned by another tool. Models do not naturally separate “rules the developer wrote” from “sentences inside the data.”
Structured data makes this quieter. A customer note is just customer_note in an API. Once it is in context, it shares a token stream with the system prompt. If the Host concatenates tool results into the next turn as-is, an external site or sender is effectively writing prompts for the agent.
Defense is not “please ignore malicious instructions” as another sentence. Prompts reduce risk; they are not a boundary. More stable controls:
- Label provenance — external text enters as a distinct role or wrapper (for example
tool/untrusted), never folded into system. - Shrink the visible surface — summarize instead of pasting; send only the fields you need, not the whole JSON document.
- Gate high-impact actions — transfers, deletes, outbound send need a policy engine or a human, independent of the model.
- Treat tool results as data, not instructions — Schema-validate returns before they re-enter the prompt.
Same instinct as how Apple Intelligence protects personal data: what leaks or gets abused is often a field you put in JSON. In the model, a field can be read as speech; in tool arguments, it can be read as an action.
Tool Calling: privilege abuse that looks like a valid call
The danger is not “can the model emit JSON?” It is whether the Host treats the model as an already-authorized caller. This is the classic confused deputy: the model proposes; permission belongs to the user session, tenant, and service account. “Call export_orders” is not the same as “this user may export the table.”
Over-privilege in 2026 often looks like this:
- A single
run_command/http_requestwith almost-arbitrary strings; - Tools that run as a service account broader than the end user;
- The Host checks JSON syntax but not “may this user touch this resource?”;
- User or page JSON is passed straight through as tool arguments, skipping your Schema.
A declaration that is too open, then one you can actually review:
{
"name": "fetch_url",
"description": "Fetch any URL and return text.",
"parameters": {
"type": "object",
"properties": {
"url": { "type": "string" }
},
"required": ["url"]
}
}
{
"name": "fetch_public_doc",
"description": "Fetch an allowlisted documentation URL.",
"parameters": {
"type": "object",
"properties": {
"doc_id": {
"type": "string",
"pattern": "^doc-[a-z0-9-]{1,32}$"
}
},
"required": ["doc_id"],
"additionalProperties": false
}
}
The second is not “smarter.” It turns an open capability into an auditable identifier. The Host resolves doc_id against its own allowlist, fills the origin, and applies timeouts and size caps. The model never sees an arbitrary URL, so there is one fewer path to an untrusted source.
Before execute, ask three questions: is this tool on the session allowlist? Did arguments pass a Schema and authz check independent of the model? On failure, do you reject — or feed error detail back as fresh injectable text? Parameter errors and validation covers the pipeline. For security, add: fail closed, and do not let the model retry with new arguments in an unbounded loop.
A Schema that is too loose is a vulnerability
Structured Output and Tool Calling both use JSON Schema to block illegal tokens — the 2026 default contract, see what Structured Output is. Schema guarantees shape, not safe meaning.
Loose contracts usually look like:
type: objectwith noproperties, oradditionalPropertiesleft true;- An action name that should be an enum, written as any
string; - An identifier field that can hold a page-long brief;
- Vendor
strictmode covering a subset, while the Host assumes “the model side already blocked everything.”
Stack two gates: model-side Schema reduces nonsense; the Host validates again with the same (or stricter) Schema, then projects into an internal type. Vendor Structured Outputs do not replace your authz. Field differences between vendors are themselves risk — see the Structured Output comparison. A migration that loosens the contract widens the attack surface.
When Schema is a security control, prefer enum, const, pattern, maxLength, minimum / maximum, required, additionalProperties: false. If you need free text, isolate it, cap it, and never map free text to a tool name or a URL.
MCP and external tools: where the trust boundary sits
MCP solves cross-process discovery and invocation. It does not solve “is this Server benign?” What MCP is already draws the JSON-RPC line: model API on one side, Host ↔ Server on the other. Security needs a second line: tool descriptions, Resource text, and return JSON from a third-party MCP Server are untrusted input.
The risk is not the protocol date. It is trust placed on the wrong layer:
description/inputSchemafromtools/listpasted into the system prompt;resultfromtools/callentering the next turn without validation;- Too many Servers on one Host, overlapping names or capabilities, execution anyway;
- Treating “the user allowed this Server” as “every sentence it returns is an instruction.”
Skills have the same class of problem: a Skill is a how-to, not an authz layer. Loading SKILL.md from an untrusted repo adds a workflow the model will follow. See what the four layers own — do not replace Schema with a Skill, or permission isolation with a Subagent. A Subagent can narrow tools; the parent still decides which secrets it receives.
Defense checklist: validate, allowlist, least privilege
Tighten along the execution path, outside-in — do not start with the prompt:
- Write the contract first. One JSON Schema per tool: complete
required,additionalProperties: false, identifiers viapattern/enum. - Validate again on the Host. Do not trust “the model already followed the Schema.” Use ajv or equivalent; fail closed.
- Project; do not pass through. Copy allowlisted fields into an internal DTO, then call the business API.
- Minimize tools. Prefer
fetch_public_docoverfetch_url; prefer read-only over write. - Authorize as the user, not as the model. Check session, tenant, and resource ACL inside the tool implementation.
- Gate outbound side effects. Send, transfer, delete, production deploy: policy engine or human confirm.
- Downgrade external text. Tool results, pages, and mail are not system. When needed, a read-only Subagent returns a summary to the parent.
- Audit arguments. Log tool name, validated parameters, who authorized, whether it was denied.
Prompts still help: say that data is not instructions, list forbidden actions. They are an assist layer. Ten more prompt sentences will not replace a missing Schema or ACL.
Review contracts with local JSON tools
Before you ship, look at three files together: the tool parameters / MCP inputSchema, one “happy path” arguments object, and one deliberately ugly sample (extra keys, wrong types, oversized strings). No real secrets, no real user data.
- JSON validator — is the sample legal JSON, and does it satisfy your Schema?
- JSON Diff — what did the model add beyond the minimal object?
- Tree viewer — is nesting deeper than it should be; does a string field hide another structure?
Nothing leaves the browser. That fits reviewing a contract about to go to production, and comparing arguments from a failed call. Stabilize field names and required, then wire the Host or MCP Server.
FAQ
Can JSON Schema stop prompt injection?
Not by itself. Schema limits the shape and range of tool arguments, which lowers the chance that an arbitrary string becomes an arbitrary action. Indirect injection happens when text enters context — you still need provenance, downgrade, and high-impact gates.
The model already uses Structured Output / strict mode. Must the Host validate?
Yes. Vendor constraints apply at generation time, and each vendor supports a different Schema subset. Security decisions must happen before execution, with your own validator and authz.
What is wrong with passing user JSON straight to a tool?
You skip the contract. Extra fields, type confusion, and nested text arrive unchanged in code that has side effects. Validate, then project; pass only allowlisted fields.
Can MCP Server output be used as the system prompt?
No. Descriptions from tools/list, Resource text, and tools/call results are untrusted data. Validate them, then re-enter the dialogue at a lower privilege.
What is the minimum viable agent security baseline?
Strict JSON Schema, Host re-validation, a tool allowlist, per-user authz, and gates on outbound side effects. Without those five, do not connect the agent to production data.
How do I check locally whether tool JSON is dangerous?
Put Schema and argument samples in the JSON toolbox for validation and Diff. Confirm required, additionalProperties, and length/enum limits before wiring the Host or MCP Server.
Summary
The agent attack surface sits between structured data and tool execution: malicious JSON slips past the contract, prompt injection steers intent, Tool Calling turns intent into side effects, and a loose Schema opens the door for all three. The 2026 default stack (Tools, MCP, Skills, Subagents) makes agents more useful — and makes “fool one invocation” worth more than “fool one reply.”
Defend inside-out: make JSON Schema a hard contract; Host validates, projects, and authorizes; keep tools small; do not treat external text as instructions. Prompts assist; they are not the boundary. Validate samples locally before you ship — models can change; field names, required, and “who may execute” should not.