Earlier posts in this series set the stage: the evolution of JSON Schema, Function Calling, and MCP explains why they exist; JSON data flow from Tool Calling to MCP traces where bytes move; whether JSON Schema is becoming the standard Agent contract covers ecosystem convergence. This article focuses on a practical question: why Tool Calling almost inevitably depends on JSON Schema, and how to classify and validate parameter vs. type errors.
When a model picks a tool and fills parameters, the host cannot “trust luck”—it must fail-fast against the same Schema before execution. One hallucinated argument can delete data, send the wrong email, or poison the next turn. Bottom line: JSON Schema is the only parameter contract understood by model APIs, MCP, and host runtimes alike; validate after parse and before execute, and feed structured errors back for retry.
Why Tool Calling depends on JSON Schema
Tool Calling (same data flow as Function Calling) means: the model chooses a tool and outputs JSON arguments that match a contract. Three parties must agree:
- Model APIs: OpenAI, Gemini, and Anthropic Tools APIs describe
parameterswith JSON Schema; some vendors also constrain decoding with Schema. - MCP: each Tool’s
inputSchemais JSON Schema; Hosts often pass it through or trim to a subset when mapping to model APIs. - Host programs: need machine-readable, versionable, CI-checkable contracts—ajv, Python jsonschema, etc. beat “JSON format in the prompt” by orders of magnitude.
Without Schema, hosts regex-parse or prompt-parse arguments—that breaks at Agent scale. Schema gives shape (which fields), types, and constraints (enum, minimum, pattern)—everything you need syntactically before calling HTTP/DB/MCP. Business rules (“does priority=high violate SLA?”) still need code; Schema blocks most hallucinations at the syntax layer.
User intent → model reads JSON Schema in tools[]
→ outputs tool_calls[].function.arguments (JSON string)
→ host JSON.parse + Schema validate
→ only then call MCP / HTTP / DB
Three places Schema sits on the call chain
| Stage | Schema role | Typical failure |
|---|---|---|
| Tool registration (tools / MCP list) | Tells the model what tools exist and what args they need | Invalid Schema, draft mismatch, misleading description |
| Model output (tool_calls.arguments) | Constrains generated parameter JSON | Missing required, wrong types, invented fields |
| Tool result (messages) | Optional: constrain result shape before context | Non-JSON response, field drift |
Vs. Structured Output: Structured Output constrains the final user-facing JSON reply; Tool Calling Schema constrains execution parameters. You can share one Schema source (Pydantic / Zod), but validate Tool arguments on every tool_calls before execute.
Parameter errors: missing, extra, wrong names, syntax
Parameter errors mean JSON may parse (or fails before parse) but violates Schema keys and required rules:
| Error | Example | Schema keyword | Mitigation |
|---|---|---|---|
| Missing required | Schema needs title, args only have priority | required | Feed error back; clarify required in description |
| Extra fields | Model invents urgent: true | additionalProperties: false | OpenAI strict often enforces; otherwise strip or reject |
| Wrong key spelling | titel vs title | properties keys | Consistent naming; strong descriptions |
| JSON syntax | Trailing comma, single quotes | (parse layer) | JSON.parse first; Structured Output reduces syntax errors |
| Empty arguments | {} but Schema has required | required, minProperties | Zero-arg tools: explicit properties: {} |
// Schema fragment
{
"type": "object",
"properties": {
"ticket_id": { "type": "string", "description": "Ticket ID" },
"note": { "type": "string" }
},
"required": ["ticket_id"],
"additionalProperties": false
}
// Model output (missing ticket_id) → validation fails
{ "note": "Please handle ASAP" }
Type errors: mismatches, enum, nesting, coercion
Type errors: fields exist but values’ JSON type or format violates Schema:
| Error | Example | Common cause |
|---|---|---|
| Primitive type | limit: "10" should be number | Models often stringify numbers |
| enum violation | priority: "urgent", enum is low/medium/high | Description didn’t list allowed values |
| Nested array/object | Expected tags: [], got string | Schema too complex for model subset |
| format string | email fails format: email | Hallucinated email or date formats |
| oneOf/anyOf | Polymorphic arg matches no branch | Over-complex Schema for target API |
Coercion: some validators coerce "10" to 10. In production Agents, prefer coercion off—silent fixes hide systematic drift. If you must coerce, document it and lock behavior in CI samples.
// Type error example
Schema: { "limit": { "type": "integer", "minimum": 1, "maximum": 100 } }
Model: { "limit": " fifty " } // string, not numeric → fail
Validation: syntax, arguments, strict mode
1. Validate the Schema itself
Before registering tools, meta-validate parameters / inputSchema (draft 2020-12, etc.). JSON Toolbox in the browser works locally—don’t ship invalid Schema to model APIs.
2. Validate arguments against Schema
After JSON.parse(arguments), validate with the same Schema used at registration:
- JavaScript / TypeScript: ajv (mind draft and
strictoptions) - Python: jsonschema, Pydantic (
model_validateafter JSON parse) - Codegen: Zod / Pydantic → JSON Schema single source
3. Vendor strict mode
OpenAI strict: true requires a stricter subset (e.g. all objects with additionalProperties: false). That reduces model-side errors but does not replace host validation—dialects differ by vendor; see the Contract article.
4. Sample-driven CI
Per tool: valid argument samples + intentional failures in CI. Schema changes are breaking API changes—version them.
Pipeline de ponta a ponta e feedback de erros
Pipeline mínimo (estende o artigo sobre fluxo de dados):
1. tools/list or static register → validate each inputSchema syntax
2. On tool_calls → JSON.parse(arguments)
├─ parse fail → tool message "JSON syntax error: …" → model retry
└─ parse ok → ajv/jsonschema validate
├─ fail → structured errors (missing, type, enum) → feed back
└─ ok → execute + optional business rules
3. Tool result → optional result Schema before append to messages
4. Log: schema version, raw arguments, error codes (no secrets)
Error feedback must be machine-readable: “ticket_id is required” beats “bad params, retry”. Many frameworks format validation errors as JSON in the tool role for self-correction.
A execução ainda precisa de autenticação e idempotência – o esquema garante a forma, e não “este ticket_id pertence ao usuário”.
Recomendações práticas
- Fonte de esquema único: Pydantic / Zod → MCP inputSchema + ferramentas OpenAI.
- As descrições são prompts: elas orientam enum e exigem conformidade - revise o esquema como o código da API.
- Esquema simples, validação estrita: corte a profundidade oneOf/$ref para o subconjunto da API de destino; falha rapidamente, sem soluções silenciosas.
- Dois pontos de verificação obrigatórios: após tool_calls antes da execução; depois de MCP retornar antes do contexto (se os resultados alimentarem o modelo).
- Valide localmente primeiro: cole Schema + sample arguments em JSON Toolbox antes da produção.
- Separado da Saída Estruturada: Esquema de resposta do usuário vs. Esquema de ferramentas – não mescle.
Perguntas frequentes
O Tool Calling pode pular o JSON Schema e usar linguagem natural para parâmetros?
Protótipos sim; produção não. A linguagem natural não pode fail-fast ou versão CI; os modelos omitem campos e tipos de desvio. APIs convencionais e MCP são padronizadas como Schema.
Os argumentos são uma string ou um objeto?
A maioria das APIs de conclusão de bate-papo usa uma string JSON – os hosts JSON.parse são validados. Algumas APIs mais recentes retornam objetos; de qualquer forma, valide com o mesmo esquema.
Quantas tentativas em caso de falha na validação?
Freqüentemente, 1–3 com feedback estruturado de erros e, em seguida, esclareça ou escale. A repetição infinita queima tokens e pode gerar alucinações.
ajv vs Pydantic?
Hosts Node/TS: ajv no JSON Schema diretamente. Python com modelos Pydantic: gere Schema + model_validate em tempo de execução. Mesma fonte do esquema voltado para o modelo.
Com o modo restrito ativado, ainda valida no host?
Sim. rigoroso reduz erros de modelo; isso não impede resultados MCP sujos, desvios de esquema/código ou violações de regras de negócios.
Como validar esquema e argumentos localmente?
Cole o esquema e o JSON de amostra em JSON Toolbox — validação local do navegador, nada carregado.
Resumo e próximas etapas
Tool Calling depende do JSON Schema porque é o contrato de parâmetro compartilhado e verificável para modelos, MCP e hosts. Classificar erros de parâmetros (ausentes, extras, nome errado, sintaxe) e erros de tipo (tipos, enum, aninhamento); interceptar antes de executar e alimentar erros estruturados para autocorreção.
Próximo: escolha uma ferramenta real (por exemplo, criação de ticket), escreva Esquema + amostras válidas/inválidas, valide localmente em JSON Toolbox e conecte o Agent. Ordem da série: evolução → fluxo de dados → contrato → este artigo (validação).