Reflexion & Self-Critique
After this lesson, you will be able to:
- Explain why adding a self-critique step lifts agent quality 10-30% on hard tasks — at the cost of 30% latency and 2x token usage
- Implement Reflexion-style verbal RL: agent attempts → critic evaluates → reflection note added to memory → retry
- Pick when self-critique is worth the cost (code generation, math reasoning, multi-hop QA) vs when ReAct alone is enough
- Avoid the common pitfalls: critic that always agrees, infinite reflection loops, reflections that don't change behavior
Before You Start
#The Core Idea
#Anatomy of Reflexion: Three Roles, One Memory
Shinn 2023 names three distinct roles inside a Reflexion agent. Conflating them is the most common implementation mistake.
| Component | Job | Common implementation |
|---|---|---|
| Actor | Attempt the task. Produces an action / output. | The "main" LLM call, usually a capable model (Claude Opus, GPT-4o). Sees the task plus all prior reflections in its context. |
| Evaluator | Score the actor's output. Pass/fail, possibly with structured feedback. | Either (a) external code (unit tests, regex match, ground-truth comparison), or (b) an LLM critic prompt that returns JSON like {score, passing, notes}. |
| Self-reflection | Write a verbal lesson learned from the failed attempt. Compresses (action, evaluation) into one or two sentences of actionable advice. | A separate LLM call with a reflection-specific prompt. Output is appended to memory. |
#The Reflexion Loop
1. Actor: attempt the task → action_1
2. Evaluator: score the result
- Pass → return result
- Fail → continue
3. Self-reflection: write a lesson on why action_1 failed
4. Memory: append the reflection to a persistent reflection list
5. Actor (retry): attempt the task with reflections in context → action_2
6. Repeat from step 2 (with max retries)
A worked example on a math problem makes the roles concrete:
Trial 1
Task: "What is the smallest positive integer divisible by 1..10?"
Actor: "It's 2 * 3 * 5 * 7 = 210." (uses unique primes <= 10)
Eval: FAIL (expected 2520; got 210)
Reflect: "I only multiplied prime factors. I must use the LCM, which means
including the highest power of each prime <= 10: 2^3, 3^2, 5, 7
= 8 * 9 * 5 * 7 = 2520."
Trial 2 (memory now contains the trial-1 reflection)
Actor: "LCM(1..10) = 2^3 * 3^2 * 5 * 7 = 2520."
Eval: PASS
In the math example above, suppose Trial 1's reflection had been only 'I made an arithmetic error.' What's the most likely Trial 2 outcome?
#Implementation: Code-Gen Reflexion
Tests · Verify Reflexion catches cases where the first attempt forgets case-insensitivity or punctuation handling. Verify it converges within 3 attempts on most palindrome cases.
#When to Use Self-Critique
| Setting | Use Reflexion? |
|---|---|
| Code generation with executable tests | Yes — clear pass/fail signal, big quality lift |
| Math reasoning (GSM8K, MATH) | Yes — wrong-answer catches via re-derivation |
| Multi-hop QA with verifiable facts | Yes — fact-check loop catches hallucinations |
| Open-ended creative writing | No — no clear pass/fail; subjective |
| Latency-critical (< 2s) | No — Reflexion adds 30%+ latency |
| Tasks where first-pass succeeds 90%+ | No — overhead not worth it |
#Pitfalls
#Self-Refine: Reflexion Without an External Evaluator
The pipeline is:
output_0 = actor(task)
for k in range(max_iters):
feedback_k = self_critic(task, output_k)
if feedback_k.contains("no changes needed"):
return output_k
output_{k+1} = actor(task, output_k, feedback_k)
Where Reflexion needs a pass/fail signal (tests, ground truth, downstream environment), Self-Refine works on any open-ended task — summarization, code style, story generation. Madaan et al. reported 5-40% gains across math, code, dialogue, and constrained generation tasks, without any external feedback channel.
Two practical differences to keep in mind:
- Risk of degradation. Without ground truth, the critic can push the output in the wrong direction. A safety check: keep
output_0and only commit tooutput_kif the critic explicitly approves it; otherwise return the earliest "good enough" version. - Domain mismatch. Self-Refine shines on tasks with clear stylistic or structural criteria (e.g., "make this code more readable"). It struggles when correctness is what you want — there, you still need Reflexion with executable checks.
Your team wants to improve the quality of LLM-generated meeting summaries. There's no ground-truth 'correct summary'. Which pattern fits?
#Retention of Reflections Across Trials
A subtle but important design choice: how long do reflections live? Three regimes:
| Memory regime | Behavior | When to use |
|---|---|---|
| Per-task | Reflections cleared at the end of each task. Each task starts fresh. | Default. Reflections from "factor a polynomial" don't help "translate to Spanish". |
| Cross-trial within episode | Reflections persist across trials of the SAME task, then clear. | Original Reflexion paper setup. Standard. |
| Persistent across tasks | Reflections from every task aggregated into a long-term memory. | Advanced — risks unbounded growth, requires summarization/retrieval. |
In practice, the cross-trial-within-episode regime is what most production code does. Persistent cross-task reflections sound powerful but introduce a new failure mode: reflections from unrelated tasks contaminating context ("I should remember the array was 0-indexed last time" applied to a task that has nothing to do with arrays). If you do go persistent, gate retrieval by task similarity or summarize aggressively.
#Reflexion vs o1/o3: Train-Time vs Inference-Time
The 2024-2026 frontier: train models to do Reflexion internally rather than wrapping them with an external loop. OpenAI o1 generates a long internal "reasoning chain" that includes self-critique steps — Reflexion built into the weights.
For most production: external Reflexion loops still win for tasks where you have explicit ground truth (tests, fact-check sources). For exploratory reasoning: o1/o3-class models with built-in critique are simpler.
#Key Takeaways
- Reflexion adds a self-critique step that lifts agent quality 10-30% on hard tasks — at 30% latency cost
- Loop: attempt → evaluate → reflect → retry with reflection in context. Hard cap at 3-5 attempts.
- Use Reflexion when: clear pass/fail signal exists (tests, ground truth) and one-shot fails 30%+ of the time
- Avoid Reflexion when: latency-critical, subjective tasks, first-pass succeeds 90%+
- o1/o3-style models bake Reflexion into the weights. For non-verifiable reasoning, prefer those over external loops
#Quick Check
Why does Reflexion improve LLM agent quality without any model retraining?
A team adds Reflexion to a customer-support chatbot. Quality on hard tickets goes up 12%, but the median p95 latency climbs from 2s to 6s. What's the most defensible production decision?