Agentic Workflow Patterns
After this lesson, you will be able to:
- Walk through Anthropic's 5 canonical workflow patterns — prompt chaining, routing, parallelization, orchestrator-workers, evaluator-optimizer — and pick the right one per task
- Distinguish 'workflows' (deterministic LLM compositions) from 'agents' (autonomous loops), and prefer workflows whenever they suffice
- Implement an orchestrator-workers pattern in 50 lines that splits a task across parallel sub-agents and aggregates their outputs
- Recognize when to escalate from workflow patterns to autonomous agents, and the cost of doing it prematurely
Before You Start
#Workflows vs Agents
#The Agent-vs-Workflow Decision Tree
Before reaching for any pattern, ask the question that Anthropic puts at the front of the post: does this task actually need LLM-driven control flow? Most of the time the answer is no, and you've just saved yourself days of debugging.
New task to build
│
▼
Can a single LLM call solve it?
│ │
│ yes │ no
▼ ▼
use 1 prompt Do I know the steps in advance?
│ │
│ yes │ no
▼ ▼
Linear sequence? Need LLM to plan
│ │ at runtime?
│ yes │ no │
▼ ▼ ▼
prompt Routing / Orchestrator-
chaining parallel- workers
ization │
▼
Quality target very high
with verifiable criteria?
│ │
│ yes │ no
▼ ▼
Evaluator- Autonomous
optimizer agent (ReAct +
loop tools, open loop)
#The Five Patterns
#1. Prompt Chaining
Decompose into a sequence of LLM calls; each step's output feeds the next. Optionally gate with programmatic checks between steps.
Step 1: Generate marketing copy from brief
↓
(gate: does it mention required claims? else retry step 1)
↓
Step 2: Translate copy to Spanish
↓
(gate: does length stay within 10% of original?)
↓
Step 3: Format as HTML email
# Prompt chaining — code-pseudocode
def prompt_chain(brief):
copy = llm("write marketing copy for: " + brief)
if not contains_required_claims(copy): # programmatic gate
copy = llm("rewrite, including claim X: " + copy)
es = llm("translate to Spanish: " + copy)
if abs(len(es) - len(copy)) / len(copy) > 0.1: # gate
es = llm("retranslate concisely: " + copy)
return llm("format as HTML email: " + es)#2. Routing
A first LLM call classifies the input; route to specialized handlers.
┌─→ refund_handler (cost-optimized model)
input ──┤
├─→ technical_handler (premium model + tools)
└─→ billing_handler (mid-tier model + DB access)
# Routing — code-pseudocode
HANDLERS = {
"refund": (cheap_model, refund_tools, refund_prompt),
"technical": (premium_model, tech_tools, tech_prompt),
"billing": (mid_model, billing_tools, billing_prompt),
}
def route(query):
category = llm("classify into {refund|technical|billing}: " + query)
model, tools, system = HANDLERS[category]
return model.complete(system=system, user=query, tools=tools)#3. Parallelization
Split the input across multiple parallel LLM calls, then aggregate.
Two flavors
- Sectioning: split independently solvable subtasks, run in parallel.
- Voting: run the same task multiple times with different temperatures or prompts; pick majority answer or use the variations.
input → split → [worker1, worker2, worker3] → aggregator → output
# Parallelization — code-pseudocode (sectioning + voting flavors)
# Sectioning: distinct subtasks, run concurrently
async def sectioning(doc):
summary, sentiment, entities = await asyncio.gather(
llm("summarize: " + doc),
llm("sentiment: " + doc),
llm("extract entities: " + doc),
)
return {"summary": summary, "sentiment": sentiment, "entities": entities}
# Voting: same task N times, aggregate
async def voting(question, n=5):
answers = await asyncio.gather(*[llm("answer: " + question, temp=0.7) for _ in range(n)])
return majority_or_aggregate(answers)#4. Orchestrator-Workers
A central LLM (orchestrator) plans and dispatches subtasks to worker LLMs; integrates their results.
input → orchestrator (plans 3 subtasks dynamically)
↓
[worker1, worker2, worker3] (each handles a subtask)
↓
orchestrator (integrates outputs into final result)
# Orchestrator-workers — code-pseudocode
async def orchestrator_workers(task):
plan = json.loads(llm("plan subtasks as JSON list of {name, instruction}: " + task))
worker_outputs = await asyncio.gather(
*[llm(step["instruction"], system=f"You are worker {step['name']}") for step in plan]
)
return llm("integrate these findings: " + json.dumps(dict(zip(
[s["name"] for s in plan], worker_outputs
))))#5. Evaluator-Optimizer
One LLM produces output; a second LLM evaluates and feeds back; the first revises. Loop until evaluator approves or max iterations.
input → optimizer (attempt) → evaluator (score + feedback)
↑ ↓
└─────── reflection ──────┘
# Evaluator-optimizer — code-pseudocode
def evaluator_optimizer(task, max_iters=3):
output = llm("attempt task: " + task)
for _ in range(max_iters):
verdict = json.loads(llm("evaluate, return {pass: bool, notes: str}: " + output))
if verdict["pass"]:
return output
output = llm(f"revise given feedback {verdict['notes']!r}: " + output)
return output # best-effortThe five patterns share a common structure: a deterministic skeleton that the developer wrote, with LLM calls plugged in at specific slots. None of them lets the LLM decide what step to run next. The moment you give that authority to the LLM, you've crossed the line into an autonomous agent, which is a separate design point with its own tradeoffs (see the decision tree above).
You're building a system that classifies support tickets, routes urgent ones to a 'fast' summarizer pipeline (1 LLM call), and routes complex ones through a 'researcher' pipeline (orchestrator with 3 workers). What's the outer pattern?
#Implementation: Orchestrator-Workers
Tests · Verify the orchestrator dispatches 1-3 workers in parallel. Verify the final review integrates content from each worker. Total time should be ~3x worker latency, not 9x (parallelism).
#Decision Rubric
You're building a customer-support chatbot that needs to handle billing, technical, and refund queries. The categories are well-defined and each needs different tools. Best pattern?
A team is using an autonomous agent (ReAct + 12 tools) for invoice processing. The structure is always: extract fields, validate, route to approver, post to ERP. Cost is 4x their target; debugging is painful. What's the most surgical fix?
| Task pattern | Recommended workflow |
|---|---|
| Linear pipeline (research → draft → translate → format) | Prompt chaining |
| Distinct categories needing specialized handling | Routing |
| Independent subtasks that can run in parallel | Parallelization |
| Subtasks discovered at runtime | Orchestrator-workers |
| Iterative quality improvement with clear criteria | Evaluator-optimizer |
| None of the above | Autonomous agent (escalate carefully) |
#Key Takeaways
- Workflows are deterministic LLM compositions with predictable cost/latency; agents are autonomous loops with higher ceiling but less predictability
- Anthropic's 5 patterns: prompt chaining, routing, parallelization, orchestrator-workers, evaluator-optimizer — covering the vast majority of production agentic systems
- Always pick the simplest pattern that works. Escalating prematurely costs you predictability without quality gain
- Parallelization gives you both quality (voting) and speed (sectioning) for the same compute budget
- Mix patterns with clean boundaries, not nested chaos
#Quick Check
What's the key difference between a workflow and an autonomous agent?