Calling an LLM API from Python: Messages, Retries, Streaming and Cost
ChatSession, that handles all three. It comes after Production Python, Testing with pytest and Building APIs with FastAPI, because it uses classes, generators, logging, retries, mocks and Pydantic together. Python Projects and the capstone use it next.By the end you can send a request and read the reply, tokens and stop reason; keep the history yourself and keep it bounded; retry only the errors worth retrying and wait as the server asks; stream a reply with a generator; get JSON you can trust; count tokens and cost; and test all of it without calling a model.
After this lesson, you will be able to:
- Send a request with a system prompt and a message list, and read the reply text, token usage and stop reason
- Explain why the API remembers nothing, and keep a conversation by resending a history you trim
- Sort API errors into worth retrying and your own bug, and retry with backoff, jitter and the server's retry-after
- Stream a reply with a generator and return validated JSON, repairing a bad reply once
- Compute cost as tokens times a price you look up, and keep keys out of code and logs
- Test LLM code with an injected mock client so no test calls the network
Before You Start
#What an LLM API Call Is
A chat model on the web is an HTTP service. You send one request, a JSON document with a list of messages, and you get one response, a JSON document with the reply text and a count of the tokens used. A token is a chunk of text, often a word or part of one, and it is the unit everything is measured and billed in. The service keeps no memory of earlier requests.
user for what you said and assistant for what the model said. A separate system field carries standing instructions, such as the persona or the rules for answers. Reading the response is just as plain.| Part | What it is | Where you find it |
|---|---|---|
model | Which model answers | Request. Read the id from the environment and check the provider's current model list |
max_tokens | The most the reply may contain | Request. Required |
system | Standing instructions | Request |
messages | The conversation so far | Request. A list of role and content dicts |
content | The reply, as a list of blocks | Response. The text is content[0].text |
usage | input_tokens and output_tokens | Response |
stop_reason | Why the reply ended | Response. end_turn or max_tokens here |
Other providers' chat APIs have the same basic shape, a list of role and content messages in and text with token counts out. The field names differ, so check the provider's documentation rather than assuming.
#The mock client
anthropic package, so the code you write against it runs against the real service after one change. Replies come from a short keyword table, so every printed output is the same on every run. Token counts are len(text.split()), a stand-in; real tokenizers cut text differently. The cell below defines the whole mock. Read it once and run it, and the rest of the lesson reuses it. The cells share one Python session, so run them in order, and run this one again after a page reload.import os
from contextlib import contextmanager
from dataclasses import dataclass
from types import SimpleNamespace
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-sonnet-5-5") # check the provider's current model list
class APIStatusError(Exception): # the server answered with an error status
def __init__(self, message, status_code):
super().__init__(message)
self.status_code = status_code
class RateLimitError(APIStatusError): # HTTP 429; retry_after = seconds the server asks you to wait
def __init__(self, message="rate limited", retry_after=1.0):
super().__init__(message, 429)
self.retry_after = retry_after
class APIConnectionError(Exception): # no answer at all: the network failed
pass
@dataclass
class Usage:
input_tokens: int
output_tokens: int
@dataclass
class TextBlock:
text: str
type: str = "text"
@dataclass
class Message:
content: list
usage: Usage
stop_reason: str
model: str
role: str = "assistant"
REPLIES = { # the first keyword found in the last user message wins
"invalid": '{"label": "positive", "confidence": 0.93}',
"classify": 'Sure, here you go: {"label": "positive", "confidence": 0.93}',
"loss": "Loss measures how far a prediction is from the true answer.",
}
DEFAULT_REPLY = "I am a mock model, so I only know a few canned answers."
def count_tokens(text):
return len(text.split()) # a stand-in: real tokenizers cut text differently
class MockClient:
def __init__(self, fail_first=0, error=None):
self.calls = [] # every request received, so a test can inspect it
self.fail_first = fail_first # the first N calls raise `error`
self.error = error or RateLimitError()
self.messages = SimpleNamespace(create=self._create, stream=self._stream)
def _create(self, *, model, max_tokens, messages, system=""):
self.calls.append({"model": model, "max_tokens": max_tokens, "system": system, "messages": list(messages)})
if len(self.calls) <= self.fail_first:
raise self.error
last = messages[-1]["content"].lower()
words = next((r for k, r in REPLIES.items() if k in last), DEFAULT_REPLY).split()
stop = "max_tokens" if len(words) > max_tokens else "end_turn"
words = words[:max_tokens]
used = count_tokens(system) + sum(count_tokens(m["content"]) for m in messages)
return Message([TextBlock(" ".join(words))], Usage(used, len(words)), stop, model)
@contextmanager
def _stream(self, **kwargs):
message = self._create(**kwargs) # errors surface on entering the with block
pieces = (w if i == 0 else " " + w for i, w in enumerate(message.content[0].text.split()))
yield SimpleNamespace(text_stream=pieces, get_final_message=lambda: message)
client = MockClient()
reply = client.messages.create(
model=MODEL,
max_tokens=100,
system="You are a concise tutor.",
messages=[{"role": "user", "content": "What is a loss function?"}],
)
print(reply.content[0].text)
print(reply.usage)
print(reply.stop_reason, reply.role)Output:
Loss measures how far a prediction is from the true answer.
Usage(input_tokens=10, output_tokens=11)
end_turn assistant
MockClient(fail_first=2) makes the first two calls raise RateLimitError, which lets you test retry code. calls records every request, so a test can check what was sent. A real client has neither. Everything else is the shape you will meet again.anthropic package, a key in ANTHROPIC_API_KEY and a network connection, so it has no Run button.import os
from anthropic import Anthropic
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-sonnet-5-5") # check the provider's current model list
client = Anthropic() # reads ANTHROPIC_API_KEY from the environment
reply = client.messages.create(
model=MODEL,
max_tokens=100,
system="You are a concise tutor.",
messages=[{"role": "user", "content": "What is a loss function?"}],
)
print(reply.content[0].text)
print(reply.usage.input_tokens, reply.usage.output_tokens, reply.stop_reason)MockClient() with two imports. Nothing else changes.from anthropic import Anthropic, APIConnectionError, APIStatusError, RateLimitError # replaces the mock's classes
client = Anthropic() # replaces MockClient()#One Call, Token Limits and Stop Reasons
max_tokens caps the reply. If the model reaches the cap before it finishes, the reply stops mid-sentence and stop_reason says max_tokens instead of end_turn. Nothing is raised, so a program that ignores stop_reason will treat a cut-off answer as a complete one.client = MockClient()
question = [{"role": "user", "content": "What is a loss function?"}]
for limit in (100, 4):
reply = client.messages.create(model=MODEL, max_tokens=limit, messages=question)
print(limit, reply.stop_reason, "|", reply.content[0].text)Output:
100 end_turn | Loss measures how far a prediction is from the true answer.
4 max_tokens | Loss measures how far
stop_reason and either raise the limit or treat max_tokens as a failure.#The Script That Works Once
messages list, appends each question and each answer, and sends the whole list every time. The questions are short. Before you run it, predict what the call sends.Each question is about 2 tokens and each answer about 12. The loop resends the whole conversation on every call. How does the number of input tokens per call change over five turns?
client = MockClient()
messages = [] # the whole conversation, resent on every call
typed = sent = 0
for question in ["What is a loss function?", "Say more.", "And again?", "Once more?", "Last one?"]:
messages.append({"role": "user", "content": question})
reply = client.messages.create(model=MODEL, max_tokens=100, messages=messages)
messages.append({"role": "assistant", "content": reply.content[0].text})
typed += count_tokens(question)
sent += reply.usage.input_tokens
print(f"you typed {count_tokens(question)} tokens, the call sent {reply.usage.input_tokens}")
print("typed in total:", typed, "| sent in total:", sent)Output:
you typed 5 tokens, the call sent 5
you typed 2 tokens, the call sent 18
you typed 2 tokens, the call sent 33
you typed 2 tokens, the call sent 48
you typed 2 tokens, the call sent 63
typed in total: 13 | sent in total: 167
You typed 13 tokens in total and the five calls sent 167. The fifth call carries 63 tokens to deliver a two-token question. A provider bills for input tokens, so a long chat gets more expensive with every turn, and in the end the history no longer fits in the model's context window. That is failure one.
Failure two is the busy day. The same loop, against a service that answers 429 (too many requests):
client = MockClient(fail_first=1) # a busy day: the server rejects the first request
answers = []
try:
for question in ["What is a loss function?", "Say more."]:
reply = client.messages.create(model=MODEL, max_tokens=100, messages=[{"role": "user", "content": question}])
answers.append(reply.content[0].text)
except RateLimitError as err:
print("crashed:", err.status_code, err, "| wait", err.retry_after, "s")
print("answers saved:", answers)Output:
crashed: 429 rate limited | wait 1.0 s
answers saved: []
retry_after, which the script ignored. A 429 is not a bug in your code. It means come back soon.Failure three is the key. The two lines below are how most first scripts start.
client = Anthropic(api_key="sk-example-not-a-real-key") # works, and ends up in Git historyA key in source is a key in every copy of the repository. Environment Variables & Secrets explains why a deleted commit does not undo it. The rest of this lesson fixes the three failures in order, inside one class.
#History Is Your Job
ChatSession holds the system prompt and the history, and sends both on every call. It records a turn only after the call has succeeded.class ChatSession:
def __init__(self, client, system="", model=MODEL, max_tokens=300):
self.client, self.system, self.model, self.max_tokens = client, system, model, max_tokens
self.history = [] # the API remembers nothing, so this list is the conversation
def send(self, prompt):
messages = self.history + [{"role": "user", "content": prompt}]
reply = self.client.messages.create(
model=self.model, max_tokens=self.max_tokens, system=self.system, messages=messages
)
text = reply.content[0].text
self.history = messages + [{"role": "assistant", "content": text}] # only after success
return text
client = MockClient()
chat = ChatSession(client, system="You are a concise tutor.")
chat.send("What is a loss function?")
chat.send("Say more.")
for number, call in enumerate(client.calls, start=1):
print("call", number, "sent roles:", [m["role"] for m in call["messages"]], "| system:", call["system"])Output:
call 1 sent roles: ['user'] | system: You are a concise tutor.
call 2 sent roles: ['user', 'assistant', 'user'] | system: You are a concise tutor.
system field was sent both times. Standing instructions go in system, not in the message list, and they do not take part in trimming.self.history. The class builds a new messages list and assigns it only after create returns. A tempting shortcut appends the user message first and the answer later, and it breaks the first time a call fails.Session.send appends the user message to history before calling the client. The first call fails with a network error, the caller catches it and tries again, and the server now receives the question twice. Change send so a failed call leaves the history exactly as it was.
first try failed server received roles: ['user']
#Keeping History Bounded
The fix for growing input is to send less. The simplest policy is a token budget for the history: drop the oldest user and assistant pairs until what is left fits. Dropping in pairs keeps the conversation starting with a user message, as chat APIs expect. The function works on a copy, so the full transcript stays with the caller for your own logs and display.
def trim_history(history, budget):
"""Drop the oldest user/assistant pairs until the history fits in `budget` tokens."""
history = list(history) # work on a copy: the full transcript stays with the caller
while history and sum(count_tokens(m["content"]) for m in history) > budget:
del history[:2]
return history
QUESTIONS = ["What is a loss function?", "Say more.", "And again?", "Once more?", "Last one?"]
def tokens_sent_per_call(budget):
client, history, sent = MockClient(), [], []
for question in QUESTIONS:
messages = trim_history(history, budget) + [{"role": "user", "content": question}]
reply = client.messages.create(model=MODEL, max_tokens=100, messages=messages)
history += [{"role": "user", "content": question}, {"role": "assistant", "content": reply.content[0].text}]
sent.append(reply.usage.input_tokens)
return sent
print("no trimming :", tokens_sent_per_call(budget=10_000))
print("budget of 30:", tokens_sent_per_call(budget=30))Output:
no trimming : [5, 18, 33, 48, 63]
budget of 30: [5, 18, 17, 32, 32]
With a budget of 30 tokens, the third call sends 17 tokens instead of 33, and the later calls stop growing. The price is that the model has forgotten the oldest turns. If the first message said "my name is Asha", a trimmed session no longer knows it. Other policies exist, such as summarising old turns with another call or pinning the first message, and each trades cost against what the model remembers. Choose by what your conversation needs.
len(text.split()) stand-in. For a real budget use the token counts the response reports in usage, or the provider's token-counting feature, and check its documentation for the details.#Errors: What to Retry and What Is Your Bug
When a call fails, the right reaction depends on why. A short taxonomy covers almost everything.
| Failure | What it means | What to do |
|---|---|---|
| 429, rate limited | You sent too much too fast | Wait as the server says, then retry |
| 5xx, server error | The service had a problem | Retry with backoff, a bounded number of times |
| Connection error or timeout | No answer arrived | Retry with backoff, a bounded number of times |
| 400, 404, 422 | Your request is wrong | Fix the code. Retrying repeats the mistake |
| 401, 403 | Your key is missing, wrong or not allowed | Fix the configuration. Do not retry |
RateLimitError and APIStatusError carry a status_code, and APIConnectionError has none. In the real SDK, RateLimitError is a subclass of APIStatusError, and a timeout is an APITimeoutError, a subclass of APIConnectionError. The same two checks work for both.#Retry with Backoff and retry-after
retry_after, it waits that long instead of guessing. It also takes a sleep argument, so a test can pass a fake that records the wait.import logging
import random
import sys
import time
log = logging.getLogger("llm")
def show_logs():
logging.basicConfig(level=logging.INFO, stream=sys.stdout, format="%(levelname)s %(message)s", force=True)
def retry_after(err):
"""Seconds the server asked us to wait, or None. The mock keeps it on the error object; the real SDK keeps it in the response headers."""
seconds = getattr(err, "retry_after", None)
response = getattr(err, "response", None)
if seconds is None and response is not None:
seconds = response.headers.get("retry-after")
return float(seconds) if seconds is not None else None
def is_transient(err):
"""Worth trying again: rate limit, server error, no connection. A 400, 401 or 404 is a bug in your request."""
if isinstance(err, APIConnectionError):
return True
return isinstance(err, APIStatusError) and (err.status_code == 429 or err.status_code >= 500)
def call_with_retry(fn, attempts=4, base=0.5, sleep=time.sleep):
for n in range(1, attempts + 1):
try:
return fn()
except (APIStatusError, APIConnectionError) as err:
if not is_transient(err) or n == attempts:
raise # not worth retrying, or out of attempts: the caller sees the real error
delay = retry_after(err) or base * 2 ** (n - 1) * random.uniform(0.5, 1.0)
delay = min(delay, 60) # never sleep for as long as a server says
log.warning("%s on attempt %d/%d, waiting %.2fs", type(err).__name__, n, attempts, delay)
sleep(delay)Before you run the demo, predict the wait.
A 429 arrives with retry_after of 3.0 seconds, and the loop was given base=0.5. How long does call_with_retry wait before the second attempt?
show_logs()
random.seed(7) # only so the printed jitter is the same on every run
request = dict(model=MODEL, max_tokens=100, messages=[{"role": "user", "content": "What is a loss function?"}])
waits = [] # a fake sleep that records instead of waiting
busy = MockClient(fail_first=2) # two 429s that say "wait 1.0 s"
print(call_with_retry(lambda: busy.messages.create(**request), sleep=waits.append).content[0].text)
print("calls:", len(busy.calls), "waits:", waits)
waits.clear()
offline = MockClient(fail_first=2, error=APIConnectionError("network down")) # no retry-after: back off
call_with_retry(lambda: offline.messages.create(**request), sleep=waits.append)
print("calls:", len(offline.calls), "waits:", [round(w, 2) for w in waits])
bug = MockClient(fail_first=5, error=APIStatusError("bad request", 400)) # your bug: do not retry
try:
call_with_retry(lambda: bug.messages.create(**request), sleep=waits.append)
except APIStatusError as err:
print("gave up at once:", err.status_code, "after", len(bug.calls), "call")Output:
WARNING RateLimitError on attempt 1/4, waiting 1.00s
WARNING RateLimitError on attempt 2/4, waiting 1.00s
Loss measures how far a prediction is from the true answer.
calls: 3 waits: [1.0, 1.0]
WARNING APIConnectionError on attempt 1/4, waiting 0.33s
WARNING APIConnectionError on attempt 2/4, waiting 0.58s
calls: 3 waits: [0.33, 0.58]
gave up at once: 400 after 1 call
retry_after, so it backed off: the waits are random within a window that doubles, which is why they differ. The third gave up at once: a 400 is a bug in the request, and trying again four times only hides it behind a delay.Two cautions from Production Python still apply. Retry only calls that are safe to repeat. And bound the retries, because an unbounded loop turns a lasting outage into a hung program.
#What the real SDK already does
Anthropic() retries up to two more times on 429, 5xx and connection errors, honours retry-after and also retries timeouts. Because it already retries, your own loop sits on top: if you set both, an unlucky call can wait several rounds of both. Pick one layer for each policy. Setting the SDK's max_retries=0 and keeping the loop above gives you the logging and the policy in your own code, and it is also what makes the loop testable.The SDK also needs a timeout, because a connection that is open but silent would otherwise wait for a very long time.
from anthropic import Anthropic
client = Anthropic(max_retries=0, timeout=30.0) # seconds; our own loop does the retrying
# or for one request only:
reply = client.messages.create(
model=MODEL, max_tokens=100, timeout=10.0,
messages=[{"role": "user", "content": "What is a loss function?"}],
)APITimeoutError, which is_transient treats like any connection error.#Streaming with a Generator
text_stream is a generator of text pieces. At the end, get_final_message() returns the complete message with its usage.client = MockClient()
request = dict(model=MODEL, max_tokens=100, messages=[{"role": "user", "content": "What is a loss function?"}])
with client.messages.stream(**request) as stream:
for piece in stream.text_stream:
print(piece, end="|") # a real UI would paint each piece as it arrives
final = stream.get_final_message()
print()
print(final.usage, final.stop_reason)Output:
Loss| measures| how| far| a| prediction| is| from| the| true| answer.|
Usage(input_tokens=5, output_tokens=11) end_turn
client.messages.stream(...) makes no request by itself. The request is sent when the with block starts, so a 429 on a stream is raised there, before any piece has arrived. That matters for retries: once pieces have reached the screen you cannot quietly restart, because the text would repeat. Retrying the opening of a stream is safe, and retrying after output has begun is a decision for your UI. The SDK's own retries cover the opening.ChatSession.stream further down. A generator that holds a with block closes it properly if the caller stops early, which is the point of Generators & Context Managers.#Structured Output: Ask for JSON, Then Validate
Programs want data, not paragraphs. The common approach is to ask for JSON in the prompt, parse the reply and validate it. The mock's reply to a classification request shows the first problem: models often wrap the JSON in chat.
import json
def parse_label(text):
"""Turn a reply into {"label", "confidence"} or raise ValueError that says what is wrong."""
start, end = text.find("{"), text.rfind("}")
if start == -1 or end < start:
raise ValueError("no JSON object found")
try:
data = json.loads(text[start : end + 1])
except json.JSONDecodeError as err:
raise ValueError(f"not valid JSON ({err.msg})") from err
if data.get("label") not in ("positive", "negative"):
raise ValueError(f"label must be 'positive' or 'negative', got {data.get('label')!r}")
confidence = data.get("confidence")
if not isinstance(confidence, (int, float)) or not 0 <= confidence <= 1:
raise ValueError("confidence must be a number from 0 to 1")
return data
client = MockClient()
reply = client.messages.create(
model=MODEL, max_tokens=100, messages=[{"role": "user", "content": "Classify: I loved this film."}]
)
text = reply.content[0].text
print(repr(text))
try:
json.loads(text)
except json.JSONDecodeError as err:
print("json.loads fails:", err)
print(parse_label(text))
for bad in ['{"label": "maybe", "confidence": 0.5}', '{"label": "positive", "confidence": 7}', "I cannot decide."]:
try:
parse_label(bad)
except ValueError as err:
print("rejected:", err)Output:
'Sure, here you go: {"label": "positive", "confidence": 0.93}'
json.loads fails: Expecting value: line 1 column 1 (char 0)
{'label': 'positive', 'confidence': 0.93}
rejected: label must be 'positive' or 'negative', got 'maybe'
rejected: confidence must be a number from 0 to 1
rejected: no JSON object found
json.loads rejects the reply because of the words in front of the braces. The function parse_label takes the text between the first { and the last }, parses it and then checks the fields. Valid JSON is not the same as a valid answer: the second rejected reply parses fine and still has a confidence of 7. The validator raises a ValueError whose message says what is wrong, and that message is what lets ChatSession repair a bad reply once, by sending the error back and asking again.parse function, because Pydantic's ValidationError is a ValueError.from typing import Literal
from pydantic import BaseModel, Field
class Sentiment(BaseModel):
label: Literal["positive", "negative"]
confidence: float = Field(ge=0, le=1)
chat.ask_json("Classify: I loved this film.", Sentiment.model_validate_json)model_validate_json expects the text to be JSON only, so a reply with chat around it fails and goes through the repair round trip. Providers also offer ways to constrain the output format. Check your provider's documentation. Validation in your own code stays necessary either way, because validation is what you can test.#Tokens and Cost
usage. Cost is arithmetic on it: tokens times the price per token, kept separately for input and output because they are priced differently. The numbers below are invented so the arithmetic is checkable. Prices change, so take real ones from the provider's pricing page and keep them in one table you can update.PRICES = {"input": 2.0, "output": 10.0} # per million tokens. INVENTED numbers: copy the real ones from your provider's pricing page
def token_cost(input_tokens, output_tokens, prices):
return (input_tokens * prices["input"] + output_tokens * prices["output"]) / 1_000_000
client = MockClient()
question = [{"role": "user", "content": "What is a loss function?"}]
usage = client.messages.create(model=MODEL, max_tokens=100, messages=question).usage
print(usage)
print(f"one call: {token_cost(usage.input_tokens, usage.output_tokens, PRICES):.7f}")
# Input tokens per call in the five-turn chat above: untrimmed against a history budget of 30.
untrimmed, trimmed = [5, 18, 33, 48, 63], [5, 18, 17, 32, 32]
for label, sent in (("untrimmed", untrimmed), ("trimmed", trimmed)):
per_1000_chats = token_cost(sum(sent) * 1000, 0, PRICES) # input side only
print(f"{label:9} {sum(sent)} input tokens per chat, {per_1000_chats:.3f} per 1000 chats")Output:
Usage(input_tokens=5, output_tokens=11)
one call: 0.0001200
untrimmed 167 input tokens per chat, 0.334 per 1000 chats
trimmed 104 input tokens per chat, 0.208 per 1000 chats
The second half repeats the five-turn chat from the trimming section: the same conversation costs about 38 percent less on the input side with a history budget, in the invented units. Cost per call is tiny, and the lesson of the numbers is the shape: input grows with history, so history policy is a cost decision.
#Keys and Logs
ANTHROPIC_API_KEY when you write Anthropic() with no arguments, which is why the earlier examples contain no key. In the version tested, the client builds without a key and the first request raises a TypeError that begins "Could not resolve authentication method", so a missing key fails loudly and early. Environment Variables & Secrets covers .env, .gitignore and rotating a leaked key.Logging needs one extra rule. The retry loop logs the error type and the wait, which is useful. Do not log request headers, the client object's configuration or whole prompts by default: a prompt may hold a user's private text, and a header holds the key. Log the model, the token counts, the status and how long the call took.
#Putting It Together: ChatSession
trim_history, call_with_retry, token_cost and your parse function from the earlier cells. The class keeps the full transcript and sends a trimmed copy. It records a turn and its tokens only after a call succeeds, and stream does the same once the stream has finished.class ChatSession:
def __init__(self, client, system="", model=MODEL, max_tokens=300, history_budget=200, attempts=4, sleep=time.sleep):
self.client, self.system, self.model, self.max_tokens = client, system, model, max_tokens
self.history_budget, self.attempts, self.sleep = history_budget, attempts, sleep
self.history = [] # the full transcript; only a trimmed copy is sent
self.tokens = {"input": 0, "output": 0} # what this session has used so far
def _request(self, prompt):
messages = trim_history(self.history, self.history_budget) + [{"role": "user", "content": prompt}]
return dict(model=self.model, max_tokens=self.max_tokens, system=self.system, messages=messages)
def _record(self, prompt, reply): # called only after a call succeeded
self.history += [{"role": "user", "content": prompt}, {"role": "assistant", "content": reply.content[0].text}]
self.tokens["input"] += reply.usage.input_tokens
self.tokens["output"] += reply.usage.output_tokens
def send(self, prompt):
request = self._request(prompt)
reply = call_with_retry(lambda: self.client.messages.create(**request), self.attempts, sleep=self.sleep)
self._record(prompt, reply)
return reply.content[0].text
def stream(self, prompt):
request = self._request(prompt)
with self.client.messages.stream(**request) as stream:
yield from stream.text_stream
final = stream.get_final_message()
self._record(prompt, final) # an abandoned stream never gets here, so it leaves no half turn
def ask_json(self, prompt, parse):
try:
return parse(self.send(prompt))
except ValueError as err: # one repair attempt: show the model what was wrong
return parse(self.send(f"That reply was invalid: {err}. Reply with JSON only."))
def cost(self, prices):
return token_cost(self.tokens["input"], self.tokens["output"], prices)
show_logs()
waits = []
client = MockClient(fail_first=1) # the first call hits a rate limit
chat = ChatSession(client, system="You are a concise tutor.", sleep=waits.append)
print(chat.send("What is a loss function?"))
print("".join(chat.stream("Say more.")))
print(chat.ask_json("Classify: I loved this film.", parse_label))
print("calls made:", len(client.calls), "| waits:", waits)
print("history:", len(chat.history), "messages")
print("tokens:", chat.tokens)
print(f"cost: {chat.cost(PRICES):.6f}")Output:
WARNING RateLimitError on attempt 1/4, waiting 1.00s
Loss measures how far a prediction is from the true answer.
I am a mock model, so I only know a few canned answers.
{'label': 'positive', 'confidence': 0.93}
calls made: 4 | waits: [1.0]
history: 6 messages
tokens: {'input': 74, 'output': 32}
cost: 0.000468
usage fields the real SDK returns.anthropic package (version 1.11.0) on Python 3.12 and 3.13, pointed at a local stub HTTP server. The stub returns the documented Messages API JSON, an event stream for streaming and, on request, a 429 with a retry-after header, a 5xx, a 400, a 401, a dropped connection and a slow answer.system as a top-level field, the messages list), response parsing, the exception classes and their status_code, reading retry-after from the response headers, the SDK's own retries and timeout behaviour, streaming and get_final_message(), and Pydantic as the parse function. It did not call the live service, because there was no API key and no network permission. Real model replies, real limits and real token counts are things only your own key can show.#Testing LLM Code Without the LLM
ChatSession receives its client as an argument, so a test passes the mock. The mock also records every request, so a test can check what was sent as well as what came back.trim_history, the retry helpers, parse_label, token_cost and ChatSession, in that order, as chat.py. Next to it, test_chat.py:# test_chat.py
import pytest
from chat import (
APIConnectionError, APIStatusError, ChatSession, MockClient, RateLimitError, parse_label, trim_history,
)
@pytest.fixture
def waits():
return [] # filled by the fake sleep, so no test ever waits
def make_session(client, waits, **options):
return ChatSession(client, system="You are a concise tutor.", sleep=waits.append, **options)
def test_history_is_resent_and_system_is_sent(waits):
client = MockClient()
chat = make_session(client, waits)
chat.send("What is a loss function?")
chat.send("Say more.")
second = client.calls[1]
assert [m["role"] for m in second["messages"]] == ["user", "assistant", "user"]
assert second["system"] == "You are a concise tutor."
def test_trim_drops_oldest_pairs_and_keeps_the_copy():
history = [{"role": "user", "content": "one two three"}, {"role": "assistant", "content": "four five"},
{"role": "user", "content": "six"}, {"role": "assistant", "content": "seven"}]
assert trim_history(history, budget=3) == history[2:]
assert len(history) == 4 # the caller's list is untouched
def test_retries_a_429_and_waits_as_told(waits):
client = MockClient(fail_first=2)
assert make_session(client, waits).send("What is a loss function?").startswith("Loss")
assert len(client.calls) == 3
assert waits == [1.0, 1.0]
def test_gives_up_after_the_last_attempt(waits):
client = MockClient(fail_first=10)
with pytest.raises(RateLimitError):
make_session(client, waits, attempts=3).send("hello")
assert len(client.calls) == 3
def test_a_400_is_not_retried(waits):
client = MockClient(fail_first=10, error=APIStatusError("bad request", 400))
with pytest.raises(APIStatusError):
make_session(client, waits).send("hello")
assert len(client.calls) == 1
def test_a_failed_call_leaves_no_half_turn(waits):
chat = make_session(MockClient(fail_first=10, error=APIConnectionError("down")), waits, attempts=2)
with pytest.raises(APIConnectionError):
chat.send("hello")
assert chat.history == []
def test_stream_pieces_join_to_the_reply_and_update_history(waits):
chat = make_session(MockClient(), waits)
assert "".join(chat.stream("What is a loss function?")).startswith("Loss measures")
assert len(chat.history) == 2 and chat.tokens["output"] == 11
def test_ask_json_accepts_a_good_reply_without_a_second_call(waits):
client = MockClient()
chat = make_session(client, waits)
assert chat.ask_json("Classify: I loved this film.", parse_label)["label"] == "positive"
assert len(client.calls) == 1
def test_ask_json_sends_the_error_back_once(waits):
client = MockClient()
chat = make_session(client, waits)
def picky(text):
if text.startswith("Sure"): # the mock's first reply has chatter in front
raise ValueError("JSON only, no chatter")
return parse_label(text)
assert chat.ask_json("Classify: I loved this film.", picky)["label"] == "positive"
assert len(client.calls) == 2
assert "JSON only, no chatter" in client.calls[1]["messages"][-1]["content"]
def test_cost_is_tokens_times_price(waits):
chat = make_session(MockClient(), waits)
chat.send("What is a loss function?")
prices = {"input": 2.0, "output": 10.0}
expected = (chat.tokens["input"] * 2.0 + chat.tokens["output"] * 10.0) / 1_000_000
assert chat.cost(prices) == pytest.approx(expected)
def test_an_abandoned_stream_leaves_no_half_turn(waits):
chat = make_session(MockClient(), waits)
pieces = chat.stream("What is a loss function?")
next(pieces)
pieces.close()
assert chat.history == []pytest -v on it with Python 3.12 (the header lines with paths and versions are left out).$ pytest -v
collected 11 items
test_chat.py::test_history_is_resent_and_system_is_sent PASSED [ 9%]
test_chat.py::test_trim_drops_oldest_pairs_and_keeps_the_copy PASSED [ 18%]
test_chat.py::test_retries_a_429_and_waits_as_told PASSED [ 27%]
test_chat.py::test_gives_up_after_the_last_attempt PASSED [ 36%]
test_chat.py::test_a_400_is_not_retried PASSED [ 45%]
test_chat.py::test_a_failed_call_leaves_no_half_turn PASSED [ 54%]
test_chat.py::test_stream_pieces_join_to_the_reply_and_update_history PASSED [ 63%]
test_chat.py::test_ask_json_accepts_a_good_reply_without_a_second_call PASSED [ 72%]
test_chat.py::test_ask_json_sends_the_error_back_once PASSED [ 81%]
test_chat.py::test_cost_is_tokens_times_price PASSED [ 90%]
test_chat.py::test_an_abandoned_stream_leaves_no_half_turn PASSED [100%]
============================== 11 passed in 0.01s ==============================
sleep, so none waits. The tests check the behaviour you would otherwise discover in production: the history is resent, a 429 is retried with the server's wait, a 400 is not retried, a failed call leaves no half turn, and a bad JSON reply is repaired once.#A note on determinism and temperature
temperature setting that made output more or less random, and a low value made it more repeatable but never guaranteed it identical. Check the current documentation for your model, because some newer models do not accept sampling settings at all: the SDK version tested here has no temperature argument on messages.create. The practical rule does not depend on it. Tests of your code use the mock, and anything that checks a real model's text should check properties, such as valid JSON with a label in the allowed set, and not exact words.#Common Mistakes
#Try It: A Token Budget
token_budget to this ChatSession. Before each call, estimate what the request will send. If the tokens used so far plus that estimate would go over the budget, raise BudgetExceeded with a message that says what was used and what is needed. After a successful call, add the call's input and output tokens to tokens_used. A rejected call must leave the history unchanged. The starter includes the mock so the cell runs on its own.Tests · With the TODOs left in, all four questions are answered and tokens_used stays 0. With a correct solution the output is: ok, tokens used so far: 16, then ok, tokens used so far: 47, then stopped: 47 used, the next call needs about 33, budget 60, then messages kept: 4.
self.estimate(messages) already counts the system prompt and every message. The check goes before create, because a call that has been made has been billed. When it works, set token_budget=None and confirm all four questions go through.cost instead of tokens, using token_cost and a price table, and add a remaining() method.#Going further
Prompting, meaning how to word instructions so a model does what you want, is taught in the NLP & Transformers track. Feeding a model your own documents is the RAG & Knowledge Systems track. Letting a model call tools in a loop is the AI Agents & Agentic AI track. Each of them sends requests with the code in this lesson. Two further topics are worth knowing by name: caching repeated prompt prefixes to cut cost, and running many requests as a batch. Look both up in your provider's documentation.
#Recap
max_tokens limit, an optional system prompt and the full message list, and each response carries the reply text, token usage and a stop_reason. Because the service remembers nothing, the conversation is a list you keep, record after success and trim to a budget.retry-after, and your own bugs (400, 401, 404), which you fix. A stream is a generator you consume as pieces arrive. Structured output means asking for JSON, parsing it, validating it and repairing a bad reply once.Cost is tokens times a price you look up. Keys come from the environment and stay out of logs. A client passed into your class lets a mock stand in for the model, so your tests are fast, free and exact.
A ChatSession has answered two questions. A third call is made with chat.send('Say more.'). How many messages does the third request contain, with no trimming?