Function Calling: How Tool Use Works Under the Hood
After this lesson, you will be able to:
- Understand that a tool call is just the AI generating JSON text token by token. There is no magic 'function calling mode', and why this matters for debugging
- Handle tricky situations: parallel tool calls (multiple tools at once), partial outputs, and malformed arguments that break your JSON parser
- Know the differences between Anthropic and OpenAI tool call formats so you can work with both
- Add Pydantic validation to catch bad tool arguments before execution, returning clear error messages the AI can learn from
- Apply tool result truncation strategies so large tool outputs do not blow up your context window
Before You Start
Lesson scope note: This lesson covers the under-the-hood mechanics — constrained decoding, parallel tool execution with asyncio, and Pydantic validation. For the conceptual 7-step flow and tool schema design, see Tool Use & Function Calling (Lesson 2).
This lesson pulls back the curtain on how tool calls actually work under the hood. It is more technical than the previous tool-use lesson, but understanding these internals is what separates developers who can debug agent failures from those who just stare at error logs. You will feel empowered after this one.
#The 2 AM JSONDecodeError
When developers first learn function calling, they picture it as magic: the model "knows" to call a function. But there is no magic. The model generates token sequences. Some token sequences happen to look like function calls because the model was trained on millions of examples of function calls. The provider then parses those tokens into a structured object your code can execute.
Once you understand the generation mechanics, debugging becomes simple. "JSONDecodeError"? The model generated malformed JSON tokens. "Missing required field"? The model assigned near-zero probability to the field name token. "Wrong argument type"? The model saw similar training examples with a string where your schema requires an integer.
#How Tool Calls Are Generated Token by Token
The LLM does not have a special "function calling mode." It generates tokens. The structured tool call output is just a particular shape of token sequence that the provider has trained the model to produce when it decides to use a tool.
Here is what happens at inference time:
#User Message Arrives
get_weather(location: string, units: string). The model receives: the conversation history, the system prompt, and the tool definitions — all as one combined context.#Model Decides: Tool or Text?
At the first output token, the model has two paths: generate natural language ("The weather in Tokyo is...") or generate a tool use signal. Because the model has been fine-tuned on examples of tool use, it assigns high probability to the tool use signal tokens when the question requires real-time data it does not have.
#Tool Use Signal Tokens
stop_reason: "tool_use" and a content block of type: "tool_use". The content block contains: id (a unique identifier), name (the tool name), and input (the arguments as JSON). These come back as structured objects, not raw text — the provider parses them before returning the API response.#JSON Argument Generation
input field, the model generates tokens one by one: {, ", l, o, c, a, t, i, o, n, ", :, , ", T, o, k, y, o, ", }. The grammar constraint at each step restricts the vocabulary: after {, only " or } are valid JSON; after "location", only : is valid; after "Tokyo", only ", ,, or } are valid. This is why JSON arguments are usually well-formed even without explicit JSON mode.#Your Code Takes Over
get_weather) and input ({"location": "Tokyo"}). Your code, not the model, actually calls the weather API. The model has stopped generating. It is waiting. Your code fetches the weather data and sends it back as a tool result in the next message.#Model Resumes
With the tool result in context, the model generates a final natural language response: "The current weather in Tokyo is 18°C and partly cloudy, with a chance of rain in the afternoon." The tool call never appears in the user-facing response — only the synthesized answer does.
#Constrained Decoding and JSON Mode
Normal LLM sampling works by assigning probabilities across the entire vocabulary (~100,000 tokens) and sampling from that distribution. Constrained decoding restricts this vocabulary at each step based on the current parse state.
#How Grammar-Constrained Sampling Works
At each token position, a parser tracks the current JSON parse state. The state determines which tokens are valid next:
- After
{: valid next tokens are"(start a key) or}(close object) - After
"key":: valid next tokens are"(string value),[(array),{(object), or digits (number) - After
"key": "val": valid next tokens are,(another field) or}(close object)
Tokens outside these valid sets are assigned probability zero before sampling. The model cannot generate invalid JSON — the probability mass is redistributed across the valid tokens at each step.
Token step 14:
Current parse state: inside object, after key "location", expecting value
Valid next tokens: " (string start), [ (array start), { (object start), 0-9 (number start)
Model's raw probability distribution:
"T" -> 0.42 (part of "Tokyo")
"S" -> 0.18 (could be "San Francisco")
"N" -> 0.11 (could be "New York")
...
"]" -> 0.003 (invalid here — zeroed out)
"}" -> 0.002 (invalid here — zeroed out)
Wait — "T" is not in the valid set either (expecting the opening quote first).
After zeroing invalid tokens and renormalizing, the only options starting here
are the quote character. The model generates: "
Try it! Send a tool call request to the Anthropic API with a tool that has a required "city" parameter of type "string" and an optional "units" parameter of type "string" (enum: ["celsius", "fahrenheit"]). Ask "What is the weather?" without specifying a city. Watch how the model still generates valid JSON but might guess a city. That is constrained decoding in action — valid structure, potentially wrong content.
#Why the Model Can Still Hallucinate Within Valid JSON
Constrained decoding ensures valid JSON structure. It does not ensure correct values. The model can generate:
{
"location": "Tokio", // Misspelled city — valid JSON, wrong value
"units": "kelvin", // Not in the enum — valid JSON, unsupported value
"date": "yesterday" // Not a valid date format — valid JSON, wrong format
}
Grammar constraints enforce syntax. Your Pydantic validation enforces semantics.
#The Outlines Library for Custom Constrained Generation
outlines library implements constrained decoding with arbitrary JSON schemas:import outlines
from pydantic import BaseModel
from typing import Literal
class WeatherArgs(BaseModel):
location: str
units: Literal["fahrenheit", "celsius"] = "celsius"
model = outlines.models.transformers("mistralai/Mistral-7B-v0.1")
generator = outlines.generate.json(model, WeatherArgs)
# The generator will ONLY produce output that parses to WeatherArgs
result = generator("What is the weather in Tokyo?")
# result is guaranteed to be a valid WeatherArgs instance#Provider Formats: Anthropic vs OpenAI
The two dominant providers use different formats for the same concept. This matters when switching providers or using a library like LiteLLM that abstracts both.
#The Same Tool in Both Formats
A simple calculator tool defined for both providers:
Anthropic Format
# Tool definition
anthropic_tools = [
{
"name": "calculate",
"description": "Evaluate a mathematical expression and return the precise numeric result.",
"input_schema": {
"type": "object",
"properties": {
"expression": {
"type": "string",
"description": "Mathematical expression, e.g. '847 * 293' or '(10000 * 1.05**10)'"
},
"precision": {
"type": "integer",
"description": "Decimal places in the result. Default: 2."
}
},
"required": ["expression"]
}
}
]
# What Anthropic returns when the model calls the tool
# response.content[0] is a ContentBlock with type="tool_use"
anthropic_response_example = {
"type": "tool_use",
"id": "toolu_01A09q90qw90lq917835lq9",
"name": "calculate",
"input": {
"expression": "847 * 293",
"precision": 0
}
}OpenAI Format
# Tool definition
openai_tools = [
{
"type": "function",
"function": {
"name": "calculate",
"description": "Evaluate a mathematical expression and return the precise numeric result.",
"parameters": {
"type": "object",
"properties": {
"expression": {
"type": "string",
"description": "Mathematical expression, e.g. '847 * 293' or '(10000 * 1.05**10)'"
},
"precision": {
"type": "integer",
"description": "Decimal places in the result. Default: 2."
}
},
"required": ["expression"]
}
}
}
]
# What OpenAI returns when the model calls the tool
# response.choices[0].message.tool_calls[0]
openai_response_example = {
"id": "call_abc123",
"type": "function",
"function": {
"name": "calculate",
"arguments": '{"expression": "847 * 293", "precision": 0}'
# Note: arguments is a STRING, not parsed JSON — you must json.loads() it
}
}Key differences:
| Anthropic | OpenAI | |
|---|---|---|
| Tool key | input_schema | parameters (nested under function) |
| Tool type wrapper | None | "type": "function" wrapper required |
| Arguments format | Parsed dict | JSON string (must json.loads()) |
| Response access | response.content (list) | response.choices[0].message.tool_calls |
| Tool ID key | id | id (same) |
#The Abstraction Layer: LiteLLM
LiteLLM normalizes these differences with a single unified interface:
import litellm
import json
# Same call works for both providers
response = litellm.completion(
model="anthropic/claude-opus-4-5", # or "gpt-4o" — same code
messages=[{"role": "user", "content": "What is 847 * 293?"}],
tools=openai_tools, # LiteLLM uses OpenAI format, converts for Anthropic
)
# Same response format regardless of provider
tool_calls = response.choices[0].message.tool_calls
for tc in tool_calls:
name = tc.function.name
args = json.loads(tc.function.arguments)
print(f"Tool: {name}, Args: {args}")⚡ Playground: Tool Use → — write a prompt and watch the model decide which tool to call and with what arguments.
#Parallel Tool Calls
Modern LLM APIs allow the model to request multiple tools in a single response. This is a major latency optimization for tasks that require independent information.
User: "What is the weather in New York and San Francisco, and what is 847 * 293?"
Without parallel tool calls:
Step 1: get_weather(New York) — wait 800ms
Step 2: get_weather(San Francisco) — wait 800ms
Step 3: calculate(847 * 293) — wait 50ms
Total: ~1,650ms + 3 LLM calls
With parallel tool calls:
Step 1: [get_weather(New York), get_weather(San Francisco), calculate(847 * 293)] — all at once
Total: ~800ms (slowest tool) + 2 LLM calls
#Handling Parallel Tool Calls with asyncio
import asyncio
import anthropic
import json
from typing import Any
client = anthropic.Anthropic()
async def execute_tool_async(tool_name: str, tool_input: dict) -> Any:
"""Execute a single tool call asynchronously."""
# Route to the appropriate tool function
if tool_name == "get_weather":
return await get_weather_async(tool_input["location"])
elif tool_name == "calculate":
return await calculate_async(tool_input["expression"])
elif tool_name == "search_web":
return await search_web_async(tool_input["query"])
else:
return {"error": "UnknownTool", "message": f"Tool '{tool_name}' is not registered"}
async def handle_parallel_tool_calls(response) -> list[dict]:
"""Execute all tool calls in a single response in parallel."""
tool_use_blocks = [
block for block in response.content
if block.type == "tool_use"
]
if not tool_use_blocks:
return []
# Launch all tool executions simultaneously
tasks = [
execute_tool_async(block.name, block.input)
for block in tool_use_blocks
]
# Wait for all to complete (or fail)
results = await asyncio.gather(*tasks, return_exceptions=True)
# Build tool result messages
tool_results = []
for block, result in zip(tool_use_blocks, results):
if isinstance(result, Exception):
tool_results.append({
"type": "tool_result",
"tool_use_id": block.id,
"content": json.dumps({
"error": type(result).__name__,
"message": str(result)
}),
"is_error": True
})
else:
tool_results.append({
"type": "tool_result",
"tool_use_id": block.id,
"content": json.dumps(result)
})
return tool_resultstool_use blocks. You execute all of them with asyncio.gather. You must return all results together in a single user message before the model continues — you cannot return them one at a time.#Handling Malformed Arguments
Even with grammar constraints, models generate wrong arguments. Missing required fields, wrong types, and invalid values all happen in production. The correct response is always: validate first, return a structured error, let the model self-correct.
#Pydantic Validation for Every Tool
from pydantic import BaseModel, Field, ValidationError
from typing import Optional, Literal
import json
# Define the argument schema with Pydantic
class CalculateArgs(BaseModel):
expression: str = Field(
description="Mathematical expression to evaluate",
min_length=1,
max_length=200
)
precision: Optional[int] = Field(
default=2,
ge=0,
le=10,
description="Decimal places in the result"
)
class GetWeatherArgs(BaseModel):
location: str = Field(min_length=2, max_length=100)
units: Literal["fahrenheit", "celsius"] = "celsius"
# Map tool names to their argument schemas
TOOL_SCHEMAS: dict[str, type[BaseModel]] = {
"calculate": CalculateArgs,
"get_weather": GetWeatherArgs,
}
def validate_and_execute(tool_name: str, raw_input: dict) -> dict:
"""Validate tool arguments and execute, returning structured errors on failure."""
if tool_name not in TOOL_SCHEMAS:
return {
"error": "UnknownTool",
"message": f"Tool '{tool_name}' is not registered. Available tools: {list(TOOL_SCHEMAS.keys())}"
}
schema = TOOL_SCHEMAS[tool_name]
try:
validated_args = schema(**raw_input)
except ValidationError as e:
# Return structured error — the model can read this and self-correct
errors = []
for error in e.errors():
errors.append({
"field": ".".join(str(loc) for loc in error["loc"]),
"error": error["msg"],
"received": error.get("input"),
"expected": error.get("ctx", {}).get("expected")
})
return {
"error": "ValidationError",
"message": f"Invalid arguments for tool '{tool_name}'",
"schema": schema.model_json_schema(),
"validation_errors": errors
}
# All arguments valid — execute the tool
try:
if tool_name == "calculate":
result = eval(validated_args.expression, {"__builtins__": {}})
return {"result": round(float(result), validated_args.precision)}
elif tool_name == "get_weather":
return {"weather": f"72°F, sunny in {validated_args.location}"} # Mock
except Exception as e:
return {
"error": "ExecutionError",
"message": str(e)
}
# Example: model passes wrong type
result = validate_and_execute("calculate", {"expression": 847, "precision": 2})
print(json.dumps(result, indent=2))
# {
# "error": "ValidationError",
# "message": "Invalid arguments for tool 'calculate'",
# "validation_errors": [{"field": "expression", "error": "Input should be a valid string", ...}]
# }#The Self-Correction Loop
ValidationError, pass it back to the model in the tool result. The model almost always corrects itself on the next attempt:Attempt 1 — Model generates:
{"tool": "calculate", "input": {"expression": 847, "precision": 2}}
# expression is an int, should be a string
Validation catches this, returns:
{"error": "ValidationError", "field": "expression",
"error": "Input should be a valid string", "received": 847}
Attempt 2 — Model self-corrects:
{"tool": "calculate", "input": {"expression": "847 * 293", "precision": 2}}
# Correct — model learned from the structured error
The key is making the error message informative. A raw Python traceback tells the model nothing useful. A structured error with the field name, what was received, and what was expected gives the model exactly what it needs to correct the argument.
#Tool Result Truncation
Tool results can be massive. A web search might return 50,000 tokens of HTML. A database query might return 1,000 rows. Adding these directly to the context window is expensive and often degrades reasoning quality (more noise, less signal).
Your tool returns a 40,000 token document. Your agent has a 200K context window. After 3 similar tool calls, you are at 120K tokens. After 5 more steps, you hit 200K. What breaks?
Context overflow is a hard failure. The API either returns an error (if you exceed the hard limit) or silently truncates the oldest messages (if your framework does truncation). Either way, the agent loses access to the tool results from early in the conversation — the ones that might contain the most important information.
#Truncation Strategies
Three strategies for keeping tool results manageable:
import json
from typing import Any
def truncate_tool_result(
result: Any,
max_tokens: int = 2000,
strategy: str = "tail"
) -> str:
"""Truncate a tool result to fit within a token budget."""
result_str = json.dumps(result) if not isinstance(result, str) else result
# Rough approximation: 4 chars ≈ 1 token
max_chars = max_tokens * 4
if len(result_str) <= max_chars:
return result_str
if strategy == "tail":
# Keep the end — often more relevant for search results
truncated = result_str[-max_chars:]
return f"[Truncated {len(result_str) - max_chars} characters from start]\n{truncated}"
elif strategy == "head":
# Keep the beginning
truncated = result_str[:max_chars]
return f"{truncated}\n[Truncated {len(result_str) - max_chars} characters from end]"
elif strategy == "summary":
# Use an LLM to summarize — most expensive but highest quality
summary = summarize_with_llm(result_str, target_tokens=max_tokens)
return f"[Summary of {len(result_str)//4} token result]\n{summary}"
elif strategy == "paginate":
# Return page 1 with instructions to request more
page = result_str[:max_chars]
remaining_pages = len(result_str) // max_chars
return f"{page}\n[Result continues — {remaining_pages} more pages available. Call tool with page=2 to continue.]"
return result_str[:max_chars]When to use each strategy
| Strategy | Use When |
|---|---|
head | Tool results are front-loaded (SQL query results, ordered lists) |
tail | Tool results get more specific toward the end (log files, search results ranked by relevance) |
summary | Result contains essential information spread throughout (documents, code files) |
paginate | The tool supports pagination and the agent needs specific pages |
#The Full Tool Handler
Combining validation, parallel execution, and truncation into a production-ready handler:
import asyncio
import anthropic
import json
from pydantic import BaseModel, ValidationError
from typing import Any
client = anthropic.Anthropic()
async def run_agent_with_full_tool_handling(
user_message: str,
tools: list,
max_result_tokens: int = 2000,
max_retries: int = 3
) -> str:
"""
Production agent loop with:
- Parallel tool execution
- Pydantic argument validation
- Tool result truncation
- Structured error responses
- Retry counting
"""
messages = [{"role": "user", "content": user_message}]
retry_counts: dict[str, int] = {}
while True:
response = client.messages.create(
model="claude-opus-4-5",
max_tokens=2048,
tools=tools,
messages=messages
)
if response.stop_reason == "end_turn":
return next(
block.text for block in response.content
if hasattr(block, "text")
)
# Collect all tool calls from this response
tool_use_blocks = [b for b in response.content if b.type == "tool_use"]
if not tool_use_blocks:
break
messages.append({"role": "assistant", "content": response.content})
# Execute all tool calls in parallel
async def process_one_tool(block) -> dict:
tool_id = block.id
tool_name = block.name
# Track retries per tool
retry_counts[tool_id] = retry_counts.get(tool_id, 0)
if retry_counts[tool_id] >= max_retries:
return {
"type": "tool_result",
"tool_use_id": tool_id,
"content": json.dumps({
"error": "MaxRetriesExceeded",
"message": f"Tool '{tool_name}' failed after {max_retries} attempts"
}),
"is_error": True
}
# Validate and execute
result = validate_and_execute(tool_name, block.input)
# Truncate if the result is large
result_str = json.dumps(result)
if len(result_str) > max_result_tokens * 4:
result_str = truncate_tool_result(result, max_tokens=max_result_tokens)
is_error = "error" in result if isinstance(result, dict) else False
if is_error:
retry_counts[tool_id] = retry_counts.get(tool_id, 0) + 1
return {
"type": "tool_result",
"tool_use_id": tool_id,
"content": result_str,
"is_error": is_error
}
tool_results = await asyncio.gather(
*[process_one_tool(block) for block in tool_use_blocks]
)
messages.append({"role": "user", "content": list(tool_results)})
return "Agent loop ended without final response"#Quick Reference: Tool Call Lifecycle
Tool Call Lifecycle
#Key Takeaways
- A tool call is token generation with grammar constraints. The model generates JSON token by token; constrained decoding restricts the vocabulary at each step to maintain valid JSON structure; this is why you get well-formed JSON but not necessarily correct values
- Parallel tool calls require one round-trip, not N. When the model requests multiple tools in a single response, execute all of them with asyncio.gather and return all results together in one message; this cuts latency from N× to 1× the slowest tool
- Pydantic validation before execution prevents ghost failures. Validate every tool argument against a typed schema before touching the tool logic; return structured errors with field names and expected types so the model can self-correct; never pass raw tracebacks
- Truncate tool results before they inflate context. A 40,000 token tool result that gets added 5 times will overflow most context windows; budget 2,000 tokens per result, choose a truncation strategy based on where the signal lives in the result
#Quick Check
Grammar-constrained decoding ensures which property of tool call outputs?