In 2023, ChatGPT plugins gave the world a first look at models calling APIs. In 2024, Function Calling became standard across vendors. In 2025, Anthropic released MCP, and IDEs like Cursor and Claude Desktop adopted it — on the same arc, JSON evolved from a data-exchange format into the Agent's type system and handshake protocol.
If you are building RAG pipelines, automation workflows, or Copilot-style products, you will eventually encounter three terms: JSON Schema (structural constraints), Function Calling (model picks tools and fills parameters), and MCP (Model Context Protocol — standardized tool connectivity). This article is for backend, platform, and AI application developers. It walks through why each layer emerged, what problem it solves, how they fit together, and how to choose in practice.
Why Agents Need Structured Interfaces
The core pattern of early LLM applications was: users asked questions → the model generated natural language → manually copied the results for execution. This is enough for chat scenarios, but it cannot reliably drive database writing, sending emails, checking inventory, etc.Repeatable and auditableautomated tasks.
The ReAct (Reason + Act) mode under the pure Prompt project allows the model to write "Action: search(query=...)" in the text, and the host program uses regular parsing - it can run, but it is fragile: bracket nesting, quote escaping, and multi-language mixing will cause parsing failure. What the production environment requires isMachine readable, verifiable, versionablecontract instead of relying on luck to parse Markdown.
JSON meets three requirements: it exists in large amounts in LLM training data, it can be read by both humans and programs, and it has a mature Schema verification ecosystem. As a result, JSON Schema has become the de facto standard for describing "what shape of data the model should output"; Function Calling also incorporates "which function to call and what parameters to pass" into the same JSON structure.
Technology evolution timeline
| stage | representative ability | Core pain points | Solution |
|---|---|---|---|
| 2022–2023 Early | Plain text + Prompt template | Output unparsable, hallucinatory parameters | Few-shot example constraint format |
| 2023 mid | ReAct / Toolformer ideas | Regular parsing action is unstable | Agreed JSON block, still rely on Prompt |
| End of 2023–2024 | OpenAI Function Calling | The formats are not uniform among manufacturers | API level tools parameters, JSON Schema description |
| 2024 | Structured Outputs | The model may still miss fields | Server-side constraint decoding, forcing compliance with Schema |
| End of 2024–2025 | MCP (promoted by Anthropic and others) | N×M integration: every IDE × every tool | Unified Host ↔ Server protocol, pluggable tools |
| 2025–2026 | Agent SDK + MCP Ecosystem | Permissions, auditing, multi-tenancy | OAuth, stdio/SSE transport, tool discovery |
The essential changes to this line are:Translate "what the model wants to do" from natural language into typed structured messages, and then safely executed by the host program or MCP Server.
JSON Schema: Agent’s “type system”
JSON Schema was originally used for API documentation and configuration verification (OpenAPI, Kubernetes CRD, etc.). In the Agent scenario, it assumes two types of responsibilities:
- Tool input parameters: Description
search_productsrequiresquery(string) andlimit(integer, default 10) - Model output: For example, when extracting entities, classification labels, approval conclusions, etc., fixed fields must be returned for downstream consumption.
Typical tool parameters Schema
{
"type": "object",
"properties": {
"city": {
"type": "string",
"description": "City name, e.g. Beijing or Shanghai"
},
"unit": {
"type": "string",
"enum": ["celsius", "fahrenheit"],
"description": "Temperature unit"
}
},
"required": ["city"]
}The description field is particularly important: it enters the context of the model and helps the model understand when to call and what the semantics of each parameter are - Schema also servesvalidatorandPrompt.
Structured Outputs and Schema
If only Schema is written into Prompt, the model may still have extra fields or type errors. Structured Outputs / JSON mode provided by OpenAI, Google, etc. will constrain tokens during the decoding stage so that the output strictly conforms to the Schema. This is a necessity for the "Invoice OCR → Structured JSON → Accounting System" type of pipeline.
Suggestions during the development stage: first use tools such as JSON toolboxLocal verification Schema syntax, and then use the sample payload to verify whether required and enum intercept illegal input as expected.
Function Calling: Handshake between models and tools
Function Calling (also called Tool Use, Tools API by various vendors) defines the communication between the model and the host.a round of handshakes:
- The host sends the tool list (name, description, parameters Schema) to the model along with messages
- The model does not directly execute the code, but returns
tool_calls: selected tool name + JSON parameter string - The host executes the real function (check DB, adjust HTTP), and stuff the result back into the conversation as
toolrole message - The model generates answers visible to the end user based on the results
Comparison with ReAct text mode
| Dimensions | ReAct text | Function Calling |
|---|---|---|
| Parameter format | Free text, needs to be parsed | JSON, API native fields |
| Multiple tools in parallel | Disaster | Support multiple tool_calls at a time |
| Model fine-tuning alignment | weak | Vendor training for tool format |
| Observability | Need to create a log by yourself | Standard message structure, easy to track |
Function Calling does not eliminate the Agent framework (LangChain, AutoGen, Cursor Agent, etc.), but becomes the interface between the framework and the model.thin protocol layer——The framework is responsible for orchestration, retry, and memory; the model API is responsible for "deciding which tool to call."
MCP: pluggable tool ecosystem
Function Calling solves the problem of "how to declare the call on the model side." But when the number of tools grows and the sources are dispersed (GitHub, Slack, Postgres, browsers, file systems), new problems arise:
- Each Host (IDE, Chat client, self-built Agent) must write an adaptation for each tool.
- Permissions, credentials, stdio/HTTP transport methods are independent of each other
- Users cannot "install an MCP Server and make it available everywhere"
Model Context Protocol (MCP)Open sourced by Anthropic in late 2024, it is positioned as a standard protocol between Host and Tool Provider. The analogy relationship is roughly:
| analogy | Web era | Agent era |
|---|---|---|
| Capability description | OpenAPI/JSON Schema | MCP Tool definition (including inputSchema) |
| Runtime connection | HTTP REST | MCP transport such as stdio / SSE |
| client | Browser, SDK | MCP Host (Cursor, Claude Desktop…) |
| plug-in market | npm, Chrome extension | MCP Server registry |
MCP core concepts
- Host: The application that initiated the connection (such as Cursor IDE)
- Client: MCP client in Host, maintaining session with Server
- Server: Processes that expose tools, resources, and prompts (such as filesystem-mcp, github-mcp)
- Capabilities: The tool list is dynamically discovered instead of hard-coded in Prompt.
MCP Tool's inputSchema itself is JSON Schema. Therefore, MCP does not replace Function Calling, but standardizes "tool implementation"; Host may still use MCP toolsmappingFunction Calling format for the model API.
How the three work together
Use a logical layer to understand the relationship between the three:
┌─────────────────────────────────────────────┐
│ 用户 / 业务系统 │
└─────────────────────┬───────────────────────┘
▼
┌─────────────────────────────────────────────┐
│ Agent Host(编排、权限、记忆) │
│ ┌─────────────┐ ┌─────────────────────┐ │
│ │ LLM API │◄──►│ Function Calling │ │
│ │ (推理) │ │ (tool_calls 消息) │ │
│ └─────────────┘ └─────────────────────┘ │
│ ▲ │ │
│ │ JSON Schema ▼ │
│ ┌────────┴────────┐ ┌──────────────────┐ │
│ │ 输出 Schema │ │ MCP Client │ │
│ │ (Structured │ │ ──stdio/SSE──► │ │
│ │ Outputs) │ │ MCP Server(s) │ │
│ └─────────────────┘ └──────────────────┘ │
└─────────────────────────────────────────────┘- JSON Schema: Cross-cutting each layer - tool parameters, MCP inputSchema, model structured output
- Function Calling: Model ↔ Host calling syntax
- MCP:Host ↔ Tool bus for the external world
Small scripts may only have Function Calling + a few local functions; enterprise-level Agent platforms often use MCP clusters + unified Schema registry + audit logs.
Complete call link example
The user asked: "What is the temperature in Shanghai today? By the way, check my json-schema related warehouse on GitHub."
- HostPull available tools from MCP:
get_weather,github_search_repos - Converted to a tools array of the model API, each with JSON Schema parameters
- ModelReturn two tool_calls, the parameters are both legal JSON
- HostCall weather Server and github Server via MCP to collect JSON results
- Results are returned as tool messages; the model synthesizes natural language replies
- If you need to write it into the work order system, then useOutput SchemaConstraint final JSON:
{ "summary", "temperature", "repo_count" }
If any step parameter does not conform to the Schema, the host can reject it and request the model to retry before execution - this is difficult to do with text ReActfail-fast.
Selection comparison and best practices
| scene | suggestion |
|---|---|
| Single backend + 3 or fewer tools | Function Calling + handwritten Schema is enough |
| IDE/Desktop Copilot, tools continue to grow | Prioritize MCP Server and reduce Host customization integration |
| The downstream system only needs JSON, not natural language. | Structured Outputs + Strict Schema |
| Multi-model vendors (OpenAI + Claude + open source) | Schema is decoupled from tool definition and manufacturer API, and middle-layer conversion |
| Compliance and Audit | Record each tool_calls and Schema version, prohibit undefined tools |
Schema design points
- Field
descriptionclearly writes business semantics, which can reduce miscalls better thantypealone. requiredBetter to be strict than loose; usedefaultor explicitly nullable for optional fields- Use string + description instead for large enumerations to avoid the
enumlist being too long and occupying the context - Schema is incorporated into Git version management, and Code Review is done the same as API changes.
FAQ
What is the relationship between JSON Schema and Function Calling?
Function Calling defines how the model declares and calls tools; JSON Schema describes the structural constraints of tool parameters and model output. Most APIs directly use a subset of JSON Schema as the parameters definition of tools.
Do I still need MCP with Function Calling?
Function Calling solves the calling protocol between the single-shot model and the host program; MCP solves how tools are discovered, authorized, connected and reused across processes. Complex Agents usually overlap the two: MCP provides the tool ecosystem, and Function Calling is the calling syntax on the model side.
Will MCP replace OpenAPI?
Will not be completely replaced. OpenAPI describes the HTTP API contract; MCP is oriented to the tool connection between Agent runtime and IDE. REST services can still use OpenAPI, and the Agent side can be accessed through MCP Server packaging.
Why does Agent output also require JSON Schema constraints?
Structured output facilitates program parsing, verification and downstream pipeline consumption, reduces missing fields or type errors caused by "free play" of the model, and improves the reliability of automated tasks.
Which layer should you learn first when developing an Agent?
Recommended sequence: JSON Schema basics → Single-tool Function Calling → Multi-step Agent orchestration → Introduce MCP on demand to connect to external systems. Each layer solves problems of different granularity.
How to locally verify the JSON Schema used by Agent?
You can use the verification function of the JSON toolbox to verify locally in the browser whether the Schema syntax matches the sample data, and the data is not uploaded to the server.
Summary and next steps
The transformation of AI Agent from "being able to chat" to "being able to do things" relies on locking uncertainty into structured boundaries layer by layer: JSON Schema defines the shape, Function Calling defines how the model reaches out, and MCP defines how the tool connects to the ecosystem. The three are not substitutes for each other;Different layers on the same stack.
Next step suggestion: take a real business tool (check orders, send notifications), write JSON Schema for it → connect to Function Calling and run through a single round → then evaluate whether it is worth encapsulating it as an MCP Server for multiple Host reuse. Schema and sample data can be verified locally in the JSON toolbox before going online.