Tool Calling Internals: How LLMs Decide to Call a Function
After this lesson, you will be able to:
- Explain what tool calling actually is at the token level — a special-token wrapper around a structured payload, not a separate model
- Read and write JSON-Schema function descriptions and predict how the model will parse them
- Compare the three dominant tool-call formats: ChatML / Hermes tags, Llama 3 python_tag syntax, and the OpenAI tool_calls JSON field
- Reason about parallel tool calls, chain-of-tools loops, and why they need different training data
- Enforce valid tool calls at decode time with grammar-constrained generation (Outlines, xgrammar, lm-format-enforcer)
- Recognize the four canonical tool-calling failure modes and the defenses for each
- Place MCP (Model Context Protocol) inside the larger picture: protocol for tool servers vs API field for tool descriptions
Before You Start
#What is tool calling, really?
A pretrained language model only ever does one thing: given a prefix of tokens, predict the next token. Everything else (chat formatting, system prompts, RAG, agents) is a convention layered on top of that single capability. Tool calling is one such convention. It works like this:
- The orchestrator (your code) sends the model a prompt that includes descriptions of available tools, usually in JSON-Schema form.
- The model generates text that includes a structured payload like
{"name": "get_weather", "arguments": {"city": "Paris"}}wrapped in special tokens or a designated field. - The orchestrator parses that payload, recognizes it as a tool call, and executes the real function
get_weather(city="Paris")in your runtime. - The orchestrator appends the tool's result back into the conversation as a new message and asks the model to continue.
- The model generates a final natural-language response using the result.
The crucial point: the model never executes anything. It only emits a string the orchestrator agrees to interpret as a command. The "tool" is whatever real code your orchestrator decides to run when it sees the structured payload.
bash in Claude Code, the same one ChatGPT uses to run a plugin, the same one a customer-support bot uses to look up your order. Different surface, identical guts.#Why this matters
Three things shifted in the year before this lesson was written:
-
Every production agent is built on tool calling. Anthropic's computer-use API, Claude Code, the MCP ecosystem, Cursor, Devin, ChatGPT Operator, Replit Ghostwriter, GitHub Copilot Workspace — all of them are tool-calling loops with different prompts. There is no separate "agent model"; the agent is the loop around the tool call.
-
Models are now trained for it. Llama 3.1 / 3.2, Claude 3.5+, GPT-4o, Mistral Large 2, Gemini 1.5+, Granite Code 8B, Qwen 2.5 — all of these ship with tool calling baked into the base model. You do not need to fine-tune a model to use tools any more.
-
The protocol layer is consolidating. OpenAI's tool-call API, Anthropic's tool blocks, and the open Model Context Protocol (MCP) are converging on the same shape: JSON-Schema in, JSON tool-call out, JSON-RPC for the server side.
If you want to build anything that reaches outside the chat window, tool calling is the only bridge. Everything else is plumbing.
#The text-level mechanics
What does a tool call actually look like to the model? It depends on which format the model was trained on. Here are the three you will meet in practice.
#Format 1: Hermes / ChatML tags
<tool_call>...</tool_call> wrapper inside the assistant turn.<|im_start|>system
You are a helpful assistant with access to functions. Call them
in JSON inside <tool_call> tags.
<tools>
[{"name": "get_weather", "description": "Current weather for a city",
"parameters": {"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]}}]
</tools>
<|im_end|>
<|im_start|>user
What's the weather in Paris?
<|im_end|>
<|im_start|>assistant
<tool_call>{"name": "get_weather", "arguments": {"city": "Paris"}}</tool_call>
<|im_end|>
<tool_call> is just two adjacent tokens it has seen thousands of times in its supervised-fine-tuning corpus, always followed by a JSON object, always terminated by </tool_call>. The orchestrator scans the generated text for that wrapper, parses the inner JSON, and dispatches.#Format 2: Llama 3 python_tag syntax
<|python_tag|> token followed by a Python-style call:<|start_header_id|>assistant<|end_header_id|>
<|python_tag|>get_weather.call(city="Paris")<|eom_id|>
<|eom_id|> ("end of message") tells the orchestrator the model is yielding control. The orchestrator parses the call (function.method(kwargs) syntax, not JSON), runs it, and appends a tool-response message.python_tag token simply tells the orchestrator "what follows is a structured call, not prose."#Format 3: OpenAI / Anthropic API fields
The frontier API providers hide the special tokens from you entirely. You send a request like:
{
"model": "claude-sonnet-4-7",
"messages": [{"role": "user", "content": "Weather in Paris?"}],
"tools": [{
"name": "get_weather",
"description": "Current weather for a city",
"input_schema": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]
}
}]
}
And get back a structured response:
{
"content": [
{"type": "tool_use", "id": "toolu_01abc",
"name": "get_weather", "input": {"city": "Paris"}}
],
"stop_reason": "tool_use"
}
tools list into an in-context prompt (often using a Hermes-style tag set), runs the model, parses the special tokens out of the raw output, and hands you a clean JSON object. The convenience is real, but the underlying mechanism is identical to the open formats. There is no separate "tool-call subsystem" inside the model.A model emits {name: 'get_weather', argments: {'city': 'NYC'}} — note the typo in 'argments'. The orchestrator validates this against the JSON Schema {required: ['arguments']}. Pass or fail?
#How tool calling is trained into the model
A base model does not know what a tool call is. The behavior is taught during supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF) or direct preference optimization (DPO). The recipe, simplified:
- Synthetic conversation generation. Hire annotators or use a stronger model to generate thousands of dialogues where the right answer is "call a tool". Each dialogue includes a tool catalog in the system prompt, a user question, and an assistant turn that emits a correct tool call in the target format.
- SFT on (prompt, correct tool call) pairs. The model learns the surface pattern: when the system prompt declares tools and the user asks something that maps to one, emit the tagged JSON.
- Schema-conditioned generation. Examples include diverse JSON Schemas (nested objects, enums, optional fields) so the model generalizes to schemas it has not seen.
- DPO on (good_call, bad_call) preference pairs. The "bad" examples include malformed JSON, hallucinated parameters, wrong tool selection, and calls when a tool was not needed. The preference signal pushes the model toward well-formed calls and away from these specific failure modes.
- Reward shaping for argument correctness. Some pipelines explicitly check argument types against the schema and only count an output as "good" if it round-trips through a validator.
The Toolformer paper (Schick et al. 2023) showed that you can do most of this in a self-supervised way: insert tool calls into a corpus, keep the ones that improve next-token loss on the surrounding text, train on the kept examples. The Gorilla paper (Patil et al. 2023) extended this to a massive catalog of real APIs, training a Llama variant to invoke 1,600+ ML model APIs by name and signature.
#Function description format: what the model actually sees
The format you use to describe a tool dramatically affects whether the model picks the right one. The de facto standard is JSON Schema:
{
"name": "get_weather",
"description": "Get the current weather conditions for a city. Returns temperature in Celsius, humidity, wind speed, and a short condition string. Use this when the user asks about weather, temperature, rain, or conditions in a specific location.",
"parameters": {
"type": "object",
"properties": {
"city": {
"type": "string",
"description": "The city name, e.g. 'Paris' or 'San Francisco'. Country optional but recommended for disambiguation."
},
"units": {
"type": "string",
"enum": ["metric", "imperial"],
"description": "Units for the temperature. Default 'metric'."
}
},
"required": ["city"]
}
}
Four pieces matter, and three of them are prose:
name: a stable identifier the model emits verbatim. Snake_case or camelCase are both fine; the model will copy whatever you give it.descriptionat the top level: the most important field. This is the prose the model uses to decide whether to call this tool at all. A description that is vague ("gets weather data") will lose to a description that gives the model concrete trigger phrases ("Use this when the user asks about weather, temperature, rain, or conditions"). Treat this like prompt engineering — because it is.parametersschema: structural constraints. Used by both the model (to format arguments correctly) and the validator (to reject malformed output).descriptionon each parameter: tells the model how to fill the slot. Concrete examples beat type annotations."e.g. 'Paris' or 'San Francisco'"is worth ten words of theory.
<tools>...</tools> tag the model has been trained on. The model has no notion of "API parameters" — it just sees text describing functions and learns that the right response is structured JSON matching the schema.You have two tools: 'search_internal_docs' (description: 'searches the company wiki') and 'search_web' (description: 'general web search'). User asks 'what is our PTO policy?' The model picks search_web. Why?
#Parallel tool calls
Modern frontier models can emit multiple tool calls in a single assistant turn. This is critical for latency: if a user asks "what's the weather in Paris and the stock price of AAPL?", a sequential agent makes two round-trips (call weather, wait, call stock, wait). A parallel-capable model emits both calls at once, the orchestrator runs them in parallel, and the user sees one combined response.
Raw token sequence in the Hermes format:
<|im_start|>assistant
<tool_call>{"name": "get_weather", "arguments": {"city": "Paris"}}</tool_call>
<tool_call>{"name": "get_stock", "arguments": {"ticker": "AAPL"}}</tool_call>
<|im_end|>
tool_use content blocks in the same message. The orchestrator dispatches both, awaits both, and feeds both tool_result blocks back in the next user message — paired by tool_use_id.<tool_call> tags in the same assistant turn during SFT will not emit them at inference, no matter how cleverly you prompt. Claude 3+ and GPT-4o/4-Turbo support this; older models often do not. Llama 3.1 supports it; Llama 3.0 does not.Why bother training models with parallel tool calls instead of just looping serially?
#Structured generation: making bad output literally unsamplable
Training the model to emit valid tool calls works most of the time. But "most of the time" is not good enough when an agent makes 100 tool calls per task and one malformed call breaks the whole chain. The robust fix is to enforce structure at decode time — to never let the model sample an invalid token in the first place.
The mechanism is called grammar-constrained decoding, and it works at the logit level:
- Compile the JSON Schema (or any context-free grammar) into a finite-state machine.
- At each decode step, ask the FSM: "given the tokens I have generated so far, which token IDs are legal next?"
- Mask the model's logits so all illegal tokens are pushed to -inf before the softmax.
- Sample from the remaining valid distribution.
{"name": "get_weather", "arguments": {"ci and the schema demands the next character be one of the letters t-z (to spell city), the decoder will mask every other token down to -inf. The model literally cannot sample a typo.The major implementations:
- Outlines (dottxt-ai): converts Pydantic models, JSON Schema, or regex into FSMs. Compiles once, masks logits at each step. Works on top of any HuggingFace model.
- xgrammar (mlc-ai): high-performance C++ engine that handles JSON Schema and EBNF grammars. Used inside vLLM and SGLang for structured-output serving at scale.
- lm-format-enforcer: similar approach, integrates with vLLM, AutoGPTQ, llama-cpp-python.
- Guidance / Instructor: developer-facing wrappers — Instructor takes a Pydantic model, generates the schema, validates the output, retries on failure. Less rigorous than logit masking (it retries instead of constraining), but very ergonomic.
response_format: {"type": "json_schema"} and Anthropic's tool-call enforcement both run a constrained decoder behind the scenes. If you are calling the API, your tool-call JSON is essentially guaranteed to be syntactically valid.Run that and watch each failure mode get caught by exactly the layer responsible for it: bad JSON by the parser, missing fields by the schema validator, unknown tools by the registry, well-formed calls by the dispatcher. This is the entire backend of a tool-calling system in 80 lines of Python. Everything you see in MCP, in the OpenAI SDK, in Claude's runtime, is a more elaborate version of the same loop.
What is the fundamental difference between Anthropic's MCP and OpenAI's tool-call API?
#Common production patterns
Tool calling is a primitive. Real agents compose it into one of three shapes.
#Single-turn tool call
The simplest loop:
user → model → tool_call → orchestrator runs tool → tool_result → model → final answer
Two model invocations, one tool execution. This is what a typical "look up something for the user" assistant does. The orchestrator code is roughly 30 lines.
#Chain-of-tools (ReAct loop)
The agent loop:
while not done:
response = model.complete(messages)
if response.tool_calls:
for call in response.tool_calls:
result = run(call)
messages.append(tool_result(call.id, result))
else:
return response.content # final answer
tool_use content blocks. This is what every general-purpose agent (Claude Code, Cursor, Devin) actually runs. The ReAct paper (Yao et al. 2022) gave this pattern its name; today nobody calls it ReAct any more, but the structure is universal. See track-09-agents/react-pattern for the historical write-up.#MCP server farms
mcp__filesystem__read, mcp__postgres__query, mcp__github__create_pr, …), passes that catalog to the model, and lets the same ReAct loop drive a much larger surface area. This is how Claude Desktop, Cursor, and Zed compose tools from independent vendors.#Tool execution pitfalls and defenses
Four failure modes show up in production at high enough frequency to deserve names.
#1. Argument hallucination
{city: string} and the model emits {city: "Tokyo", language: "ja"}. Sometimes catastrophic: the model emits {user_id: "<UNKNOWN>"} because it could not actually find a user_id.additionalProperties: false is set, the validator rejects the call and you either retry with the error message in context ("the field language is not allowed; valid fields are city") or use constrained decoding so the typo is unsamplable.#2. Wrong tool selection
search_web when it should have picked search_internal_docs, or create_pr when the user asked to describe what a PR would do, not actually open one. Tool descriptions are pattern-matched against the user message; descriptions that overlap semantically will get confused.clarify(question) that lets the model ask the user before committing to a destructive action.#3. Infinite loops
The model keeps calling the same tool with slight argument variations, never producing a final answer. Common when the tool result is ambiguous or when the model thinks the user is still waiting for more detail.
max_iterations cap (most production agents stop at 20-50 tool calls). Detect repeated (name, hash(arguments)) pairs and inject a system message: "you have called this tool with these arguments already; the result was X. Either use that result or call a different tool." Time budget on total wall-clock.#4. Capability surface mismatch
delete_file(path="/etc/passwd") because in training data filesystems usually allow delete; your sandbox does not. Calls web_search when you only wired up search_docs.{"error": "denied"} but {"error": "delete is disabled in this sandbox; you can only read files via read_file()"}. The model will incorporate that into its next attempt.#Tool-augmented training
A short tour of the research that built today's tool-calling models.
- Toolformer (Schick et al., Meta, Feb 2023) — self-supervised tool insertion. Take a text corpus, run a teacher model to propose tool calls at every position, keep only the calls whose result lowers cross-entropy on the next ~10 tokens. Train on the augmented corpus. The model learns when a tool helps, without human annotation.
- Gorilla (Patil et al., Berkeley, May 2023) — large-scale API training. Built a corpus of 1,600+ ML APIs (HuggingFace, TorchHub, TensorFlow Hub), generated training pairs of
(user_intent, correct API call), fine-tuned Llama. Showed that with retrieval-augmented training, models can correctly call APIs they have never seen during training, just from their documentation. - ToolLLM / ToolBench (Qin et al., Tsinghua, Jun 2023) — 16,000 real-world REST APIs, multi-step task generation, DFS-based annotation of correct trajectories.
- Llama 3.1 / 3.2 (Meta, Jul/Sep 2024) — first major open base model with tool calling baked in via the
<|python_tag|>token. Supports parallel calls, built-in tools (brave_search,wolfram_alpha,code_interpreter), and custom user-defined tools. - Granite Code 8B Instruct + Tools (IBM, 2024). Instruction-tuned code model with native tool calling for enterprise integrations.
- Claude 3.x / 4.x (Anthropic). Tool use as a first-class content block (
tool_use/tool_result), parallel calls supported since 3.5, computer use (clicking, typing, screen reading) released October 2024.
#Modern API patterns: a head-to-head
The two big providers expose tool calling with nearly identical semantics under slightly different field names.
| Concept | OpenAI | Anthropic |
|---|---|---|
| Tool list (request) | tools: [{type:"function", function:{name, description, parameters}}] | tools: [{name, description, input_schema}] |
| Tool call (response) | message.tool_calls: [{id, function:{name, arguments}}] | content: [{type:"tool_use", id, name, input}] |
| Tool result (request) | role:"tool", tool_call_id, content | role:"user", content: [{type:"tool_result", tool_use_id, content}] |
| Parallel calls | Yes, default on | Yes, default on |
| Force a specific call | tool_choice: {function:{name}} | tool_choice: {type:"tool", name} |
| Structured response (no tools) | response_format: {json_schema} | tool with one-call output via tool_choice |
Functionally the same; cosmetically different. Both convert the JSON Schema into an in-context prompt under the hood. Both run constrained decoding on the tool-call output to guarantee well-formed JSON. Both ship parallel calls. The interesting differences are:
- Anthropic uses content blocks, so a single assistant message can interleave prose and multiple tool calls in one ordered list. OpenAI keeps prose in
message.contentand tool calls in a paralleltool_callsarray. - OpenAI lets the model emit
parallel_tool_calls: falseas a request-level flag if you want to force serial calling. Anthropic does not have a native off-switch (you would settool_choiceper round instead). - Anthropic ships
cache_controlon tool definitions so a long tool catalog can sit in the prompt cache.
If you are starting from scratch in 2026, write your orchestrator against an abstraction layer (LiteLLM, Anthropic SDK with its OpenAI-compat shim, or your own thin wrapper) and treat the field-name differences as a serialization concern.
#A tiny JSON-mode FSM, by hand
{"city": "<string>"} — exactly enough to constrain get_weather arguments.The state set is tiny and the alphabet is small, but everything important about Outlines, xgrammar, lm-format-enforcer, and OpenAI's JSON-mode is in those 40 lines. Real implementations:
- operate at the token level (LLaMA tokenizer has ~128k tokens; the FSM precomputes which tokens are legal in each state),
- compile JSON Schema into a much larger FSM (one state per
$refboundary, one alphabet per regex pattern, etc.), - use bit-vectors so the per-step legality check is O(1) instead of O(vocab),
- handle the BPE merge ambiguity where one token might cross a state boundary.
None of those are conceptual changes. They are engineering optimizations on the same idea: at every decode step, mask out tokens the grammar forbids, then sample.
#A worked end-to-end, in token order
Putting everything together — here is what actually happens when a user asks Claude "what's the weather in Paris?" with one tool available:
- Your orchestrator sends the API a
messageslist with the user message plus atoolslist containing theget_weatherschema. - Anthropic's server converts
toolsinto a system-prompt insertion (<tools>...</tools>) using the format Claude was trained on. - The model decodes tokens. At some point its logits over the next-token distribution favor the
<tool_use_start>special token. A constrained decoder confirms valid emission and the model continues into the structured payload. - The decoder is now generating inside the schema for
get_weather. JSON-mode is on; the only sampleable tokens are those legal at the current FSM state. - At the end of the JSON, the model emits a stop token. Anthropic's server packages the tool call into a
tool_usecontent block and stops generation withstop_reason: "tool_use". - Your orchestrator parses the response, validates
inputagainst the schema once more for safety, runsget_weather(city="Paris"), and gets back{"temp_c": 14.2, ...}. - Your orchestrator sends a new request, this time with the assistant's
tool_useblock appended and a new user message containing a matchingtool_resultblock keyed bytool_use_id. - The model decodes again, now with both the original question, its own tool call, and the tool's result in context. It emits a normal assistant message, no tool tags this time, that uses the result: "It's 14.2°C and cloudy in Paris right now."
- Your orchestrator returns that text to the user.
Two model invocations, one tool call. Everything between is plumbing. This is the loop. Once you have written it once, every agent framework on the planet looks the same.
#Forward reference
track-09-agents — tool-use, react-pattern, mcp, building-an-agent, code-agents, computer-use-agents, multi-agent-systems — assumes the mechanics in this lesson. If a single pattern matters, it is this:catalog (JSON Schema) ──▶ model ──▶ structured call ──▶ runtime ──▶ result ──▶ model
(constrained ↓
decoding) (loop or stop)
Memorize that pipeline; everything else is decoration.
Key Takeaways
- Tool calling is a learned text convention: special tokens or a designated JSON field wrap a structured payload the orchestrator interprets as a function call. The model never executes anything itself.
- Three dominant formats: Hermes / ChatML tags (<tool_call>...</tool_call>), Llama 3 <|python_tag|>fn.call(...)<|eom_id|>, and API-level tool_calls / tool_use fields. Provider APIs serialize tool catalogs into one of these formats under the hood.
- JSON Schema descriptions are part schema, part prompt engineering. The top-level description chooses whether the tool fires; per-parameter descriptions choose how it is filled. Concrete trigger phrases beat type theory.
- Parallel tool calls and chain-of-tools require training data with those patterns. Claude 3+, GPT-4o, Llama 3.1+ all support parallel calls; older models do not.
- Constrained decoding (Outlines, xgrammar, lm-format-enforcer) makes malformed JSON unsamplable by masking illegal tokens at each step. In 2026, validate-and-retry is the slow path; logit masking is the fast path.
- Four failure modes: argument hallucination (defense: schema validation), wrong tool selection (defense: better descriptions, fewer tools), infinite loops (defense: max_iterations), capability mismatch (defense: action-shaped errors).
- Tool count > ~10 degrades selection accuracy. Route, prune, or retrieve.
- MCP is a protocol for tool servers (host ↔ external process). OpenAI / Anthropic tool fields describe functions to the model. MCP plugs into the model's tool-call surface; it does not replace it.