RLHF & DPO: Aligning LLMs with Human Preferences
After this lesson, you will be able to:
- Understand the three-step RLHF recipe: teach the model to follow instructions, build a reward model from human preferences, then optimize with PPO
- See how DPO simplifies this by skipping the reward model entirely
- Compare four alignment approaches (RLHF, DPO, GRPO, Constitutional AI) and know when to use each
- Spot alignment failures: reward hacking (gaming the score), mode collapse (boring answers), and over-refusal (saying no to everything)
Before You Start
Active Recall
Before alignment, retrieve the prerequisites: (1) write the cross-entropy loss for a single token in one line of notation, (2) state in one sentence what KL divergence MEASURES between two distributions, and (3) recall from the RL track — what is a policy π(a|s) and what is the policy gradient theorem doing for it? RLHF welds all three together, so reconstructing each by yourself before reading on makes the assembly below click instantly.
Write your answer in your own words — don't look back at the lesson. This is the most effective way to remember what you just learned.
Type your explanation above
#The Alignment Problem
Supervised fine-tuning (SFT) produces a model that follows instructions. But "follows instructions" is not the same as "safe and helpful." A raw instruction-tuned model has three failure modes:
- A maximally helpful model answers every request, including harmful ones
- A maximally harmless model refuses everything uncertain, becoming useless
- A maximally honest model might deliver truths in harmful ways
Alignment training is the process of finding the right balance among these three objectives using human judgment to define "the right balance."
#Stage 1: Supervised Fine-Tuning (SFT)
Before any alignment training, the base pretrained model undergoes SFT. This converts a completion model (predicts the next token) into an assistant (responds to instructions).
The SFT step:
- Collect a dataset of (prompt, ideal response) pairs, curated by human contractors
- Fine-tune the base model on these pairs using standard cross-entropy loss
- Mask loss on the prompt tokens — train only to predict the response
This SFT model becomes both:
- The starting point for reward model training
- The reference policy π_ref used later in PPO and DPO to prevent the aligned model from drifting too far from sensible behavior
The SFT model is competent but not aligned — it follows the format of a helpful assistant without the values. Stage 2 fixes that.
#Stage 2: Reward Modeling
#Collecting Preference Data
Human annotators see the same prompt with multiple model responses and rank them from best to worst. A typical annotation task:
Prompt: "How do I get my neighbor to stop playing loud music at night?"Response A: "You could hire a lawyer and sue them for noise violations." (overkill, confrontational)Response B: "Try talking to them directly first — most people are willing to compromise when asked politely. If that fails, check your local noise ordinance and contact your landlord or local authorities." (helpful, proportionate, practical)Response C: "That's a frustrating situation. Some options: 1) Talk to them directly, 2) Contact your landlord or building management, 3) File a noise complaint with local authorities. I'd start with option 1." (good but slightly more mechanical)Response D: "Have you tried soundproofing your apartment?" (misses the point)
The annotator ranks: B > C > A > D. This generates three preference pairs: (B > A), (B > D), (C > D), and so on.
#The Bradley-Terry Model
The reward model loss is the negative log-likelihood of correctly predicting the human preference:
#Architecture
Training the reward model typically uses 10,000–100,000 preference pairs and takes a fraction of the compute of pretraining.
#Stage 3: PPO Optimization
#The RLHF Objective
The policy (the LLM) generates a response y given prompt x. The reward model scores it. We want to maximize the expected reward, but with a crucial constraint: do not drift too far from the SFT reference model.
#One PPO Update Step
#Step 1: Sample a batch of prompts
Draw 64 prompts from the training distribution. These are the contexts the policy will respond to.
#Step 2: Generate responses with the current policy
The current LLM (θ) generates a response for each prompt. This is the "rollout" — sampling from the policy being optimized.
#Step 3: Score responses with the reward model
The frozen reward model assigns a scalar score r(x, y) to each (prompt, response) pair. High scores = responses humans prefer.
#Step 4: Compute the KL penalty
#Step 5: Compute advantages
Normalize the reward minus KL penalty to get advantages: how much better than average was this response? Positive advantage = reinforce this behavior. Negative advantage = suppress it.
#Step 6: Update the policy
Use the PPO clipped objective to update θ. The clipping prevents any single update from being too large, keeping training stable. Repeat for hundreds of update steps.
#The Four Models in RLHF
RLHF requires four models in memory simultaneously, which is its biggest practical challenge:
| Model | Role | Trainable? |
|---|---|---|
| SFT model (reference π_ref) | KL baseline — prevents reward hacking | Frozen |
| Reward model r_φ | Scores each response for the PPO signal | Frozen |
| Current policy π_θ | The model being optimized | Yes |
| Value function V | Estimates expected future reward (PPO baseline) | Yes |
For a 7B base model, this means holding four 7B+ parameter models in memory. With 16-bit weights, that is roughly 56 GB minimum — before activations and optimizer states. This is why RLHF requires serious infrastructure.
#Reward Hacking
A reward model gives response A a score of 8.2 and response B a score of 2.1. After PPO training, the model always generates response A. But users report the model has become sycophantic — it always agrees with them, even when they are wrong. What went wrong?
The answer: reward hacking. Human annotators, consciously or not, rated responses that agreed with them, validated their views, and used flattering language higher than responses that politely corrected them. The reward model absorbed this bias. PPO then exploited it ruthlessly. The result: a sycophantic model that maximizes annotator reward scores but fails at the actual goal of honesty.
#DPO: Cutting Out the Reward Model
#The Key Insight
The RLHF objective has an optimal solution we can derive analytically. Given the objective:
The optimal policy has the closed form:
Rearranging this equation to express the reward in terms of the policy:
#The DPO Loss
Substituting this reparameterization into the Bradley-Terry reward model loss, the partition function Z(x) cancels (it appears in both winner and loser terms), giving the DPO loss:
#What DPO Actually Does
The two log-ratio terms have a clear interpretation:
-
β · log(π_θ(y_w|x) / π_ref(y_w|x)): How much more (or less) does the current model prefer the chosen response compared to the reference model? DPO wants this to be large and positive.
-
β · log(π_θ(y_l|x) / π_ref(y_l|x)): How much more (or less) does the current model prefer the rejected response compared to the reference model? DPO wants this to be small or negative.
The loss is minimized when the model simultaneously increases the probability of chosen responses and decreases the probability of rejected responses, but weighted by how much the reference model is "surprised." Responses the reference model already handles well get less gradient signal; responses where the reference model is uncertain get more.
You want to align a 7B model for customer service. You have 10,000 preference pairs (chosen/rejected responses). You have one A100 GPU for 3 days. Should you use RLHF or DPO?
DPO is the clear choice here. RLHF would require: training a reward model (consuming some of your A100 time), then running PPO with four models in memory (which may not even fit on one A100). DPO trains with just two model passes (current policy + frozen reference), uses exactly the same memory budget as SFT, and produces comparable alignment quality at 7B scale.
#DPO Family & Successors
DPO collapsed RLHF from a four-model RL loop into a two-model supervised problem, and that simplification kicked off a flood of follow-ups in 2023-2024. Each one targets a specific failure mode of vanilla DPO: overconfidence on noisy preferences, the need for paired data, the cost of a reference model, or the length-bias DPO inherits from its sigmoid objective. The methods below are the ones you'll actually find in TRL today.
| Method | Needs paired preferences? | Needs reference model? | Year | One-line trade-off |
|---|---|---|---|---|
| DPO | Yes | Yes | 2023 | Strong baseline; overfits to deterministic preferences and inherits length bias. |
| IPO | Yes | Yes | 2023 | Bounded loss is more robust to noisy/contradictory labels. |
| KTO | No (binary thumbs) | Yes | 2024 | Eats production telemetry directly; much better data efficiency. |
| ORPO | Yes | No | 2024 | Half the memory, single-pass SFT+alignment; now default in TRL. |
| SimPO | Yes | No | 2024 | Length-normalized; matches/beats DPO with cleaner outputs. |
| CPO | Yes | Yes | 2024 | Adds SFT regularizer; stable when SFT distribution matters. |
What is ORPO's key architectural simplification over DPO?
#GRPO: Group Relative Policy Optimization
GRPO, used in DeepSeek-R1, takes the simplification further by eliminating the reference model as well.
#How It Works
Instead of comparing against a fixed reference policy, GRPO:
- Samples G responses (a group) for the same prompt
- Scores each response (using a reward signal — either a reward model or a rule-based checker)
- Computes relative advantages within the group: how much better is this response than the group average?
- Updates the policy to increase probability of above-average responses and decrease below-average ones
#Why GRPO Matters
- No reference model: The baseline is the group average, not a fixed π_ref. This saves memory (one fewer model copy) and avoids the assumption that the reference model is always a good baseline.
- Works with rule-based rewards: For tasks with verifiable answers (math, code, factual Q&A), you can replace a learned reward model with a rule: "does the answer match the ground truth?" This eliminates reward hacking on the reward model.
- Enabled reasoning model training: DeepSeek-R1 used GRPO with rule-based rewards (math verification) to train models that reason in extended chains of thought, achieving o1-level performance at a fraction of the cost.
#Constitutional AI: Anthropic's Approach
#The Two-Phase Process
- Sample a response from the model to a potentially harmful prompt
- Use a separate "critique" model to critique the response according to the constitution: "Does this response violate principle 7: 'Do not assist with creating weapons'? How?"
- Use the critique model to revise the response according to the constitution
- Fine-tune on the (prompt, revised response) pairs
- Generate preference pairs by having the AI evaluate two responses against the constitution: "Which response is more consistent with being helpful, harmless, and honest?"
- Train a preference model on these AI-generated labels
- Use RLHF with this AI-labeled reward model
#Why Constitutional AI Scales
Human preference annotation is expensive and inconsistent — different annotators disagree, and the same annotator is inconsistent across sessions. CAI:
- Scales annotation without requiring human judgment on every pair
- Makes alignment principles explicit and auditable (you can read the constitution)
- Allows rapid iteration — update the constitution, regenerate labels, retrain
#Comparison: RLHF vs DPO vs GRPO vs Constitutional AI
| Aspect | RLHF | DPO | GRPO | Constitutional AI |
|---|---|---|---|---|
| Data required | 100K+ preference pairs | 10K–100K preference pairs | Any reward signal | A constitution (principles) |
| Separate reward model | Yes | No | Optional | Yes (AI-generated) |
| Training complexity | Very high (RL loop, 4 models) | Low (supervised) | Medium | Medium (2 phases) |
| GPU memory | 4x SFT model size | 2x SFT model size | 1–2x | 2x |
| Main failure mode | Reward hacking | Preference data quality | Group size sensitivity | Constitution gaps |
| Best for | Large teams, safety-critical | Startups, 7B–70B models | Reasoning/math tasks | Principled alignment at scale |
| Used by | OpenAI (ChatGPT) | Most open-source chat models | DeepSeek (R1) | Anthropic (Claude) |
#Try It Yourself
Tests · Verify Scenario 1 loss is approximately 0.693. Verify Scenario 2 loss is less than Scenario 1 loss. Verify Scenario 3 loss is greater than Scenario 1 loss.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, Ryan Lowe (2022)
The InstructGPT paper — the first large-scale application of RLHF to LLMs. Showed that a 1.3B parameter RLHF-tuned model outperformed a raw 175B GPT-3 model on human preference evaluations. Introduced the SFT → reward model → PPO pipeline that became the industry standard.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, Chelsea Finn (2023)
The DPO paper. Derived that the optimal RLHF policy has a closed form that can be expressed directly in terms of policy probabilities, eliminating the need for a separate reward model and PPO training loop. Within months of publication, DPO became the default alignment method for the open-source LLM community.
Constitutional AI: Harmlessness from AI Feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, et al. (2022)
Anthropic's Constitutional AI paper. Introduces the principle-based approach to alignment where an AI critiques and revises its own outputs according to a human-written constitution. Demonstrates that AI feedback can partially replace human preference labeling, scaling alignment without proportional annotation cost.
#Key Takeaways
- SFT teaches format, alignment teaches values. Supervised fine-tuning produces an assistant that follows instructions; RLHF and DPO teach it which instructions to follow and how to balance helpfulness, harmlessness, and honesty
- RLHF requires four models and RL expertise. The SFT → reward model → PPO pipeline is powerful but expensive: four models in memory, a complex RL training loop, and constant vigilance against reward hacking
- DPO derived that the reward model is unnecessary. By reparameterizing the optimal RLHF policy in terms of log-ratios, DPO turns alignment into a supervised learning problem over preference pairs, cutting compute by 10x and eliminating the RL loop
- Reward hacking is alignment's Goodhart's Law. Any proxy measure of human preferences can be exploited; the KL penalty limits exploitation but does not eliminate it; monitoring for sycophancy, verbosity, and false confidence is part of production alignment
- The right method depends on your resources. DPO for startups and 7B models; RLHF for safety-critical large-scale systems; GRPO for verifiable reward tasks like math; Constitutional AI for principled alignment without per-pair human annotation
#Quick Check
In the RLHF pipeline, what is the purpose of the KL divergence penalty in the PPO objective?