After this lesson, you will be able to:
- Explain why the public web is no longer sufficient to scale frontier LLM pre-training.
- Describe the Phi recipe, Self-Instruct, Evol-Instruct and persona-driven synthesis at a recipe level.
- Diagnose the model-collapse failure mode and name the canonical defences (mix + re-base + diversity).
- Argue when a synthetic-data sample is safe to keep — i.e. only when you can VERIFY it.
- Reason about the unit economics: why $0.60 / 1M tokens beats human labelling by ~100x.
Before You Start
#The web ran out
In 2020 GPT-3 was trained on roughly 500B tokens scraped from the open web — Common Crawl, books, Wikipedia, code. That felt huge. Four years later it looks small.
Will we run out of data? Limits of LLM scaling based on human-generated data
Villalobos et al. (2024)
Epoch AI's projection: stock of high-quality public text gets exhausted between 2026 and 2032 at current consumption rates.
#The Phi recipe: textbook-style synthesis
Microsoft's Phi line (Phi-1 → Phi-1.5 → Phi-2 → Phi-3 → Phi-4) is the single most influential public demonstration that small models can punch above their weight if you feed them better data instead of more data. The recipe is straightforward.
- Curate a seed list of educational concepts — topics from textbooks, programming tutorials, mathematics curricula, reasoning patterns.
- Prompt a strong teacher model (originally GPT-3.5, later GPT-4) to generate "textbook-style" expositions of each concept: definitions, examples, worked problems, common pitfalls.
- Filter aggressively for quality. Phi-1 used an educational-content classifier; later iterations added perplexity gating, deduplication, and topic balance.
- Train a small model from scratch on the resulting synthetic corpus.
Textbooks Are All You Need
Gunasekar et al. (Microsoft) (2023)
The Phi-1 paper. A 1.3B model trained on 7B tokens of synthetic Python textbooks beat much larger models on HumanEval.
Textbooks Are All You Need II: phi-1.5 technical report
Li et al. (Microsoft) (2023)
Generalises the textbook approach beyond code into common-sense reasoning.
The headline result of Phi-2 (released late 2023) was that a 2.7B-parameter model trained mostly on synthetic textbook-quality data outperformed Llama-2-7B — a model with ~2.6x more parameters — on a broad slate of common-sense and reasoning benchmarks. Phi-4 (December 2024) pushed this further: 14B parameters, ~100% synthetic training mix, performance competitive with much larger open-weights baselines.
Phi-4 Technical Report
Abdin et al. (Microsoft) (2024)
14B parameters, training mix dominated by synthetic data generated and filtered with a careful curriculum.
Phi-4's training data is roughly what percentage synthetic?
#Self-Instruct and Alpaca: bootstrapping SFT from 175 seeds
Self-Instruct: Aligning Language Models with Self-Generated Instructions
Wang et al. (2022)
The seminal paper that showed you can bootstrap a high-quality instruction-tuning dataset from a tiny seed pool plus a strong teacher LLM.
The recipe is almost embarrassingly simple:
- Hand-write ~175 seed instructions covering a diverse set of tasks — summarise this, explain that, write code for X, classify Y.
- Repeatedly sample 8 of these as in-context examples and prompt a strong LLM ("write a new instruction in the style of the examples").
- Have the same LLM write a response to its own newly-generated instruction.
- Filter: drop near-duplicates, drop instructions that produce empty / refused / malformed answers, drop low-diversity outputs.
- Fine-tune a target model on the resulting (instruction, response) pairs.
Stanford Alpaca: An Instruction-following LLaMA Model
Taori, Gulrajani, Zhang et al. (2023)
Applied Self-Instruct to text-davinci-003, fine-tuned Llama-7B on 52K synthetic instructions. Cost: ~$600 of API calls.
This is the moment "training your own ChatGPT clone" became something a graduate student could do over a weekend. It also forced every major lab to think carefully about whether teacher-output licensing terms applied to distilled students.
#Evol-Instruct: turn easy problems into hard problems
WizardLM: Empowering Large Language Models to Follow Complex Instructions
Xu, Sun, Xu et al. (2023)
Introduced Evol-Instruct — four prompt-engineered operators that iteratively make seed instructions harder. Spawned WizardCoder and WizardMath.
The trick is to treat instruction-difficulty as something you can climb, not something you can only sample. Evol-Instruct uses four prompt-engineered "evolution operators":
| Operator | What it asks the teacher to do | Example |
|---|---|---|
| Add complexity | Add a new constraint or requirement | "Now require the answer to also handle negative inputs" |
| Deepen | Push the topic into more specialised territory | "Rephrase using domain-specific terminology" |
| Concretize | Replace abstract terms with concrete entities | "Use a specific real-world example instead of the abstract case" |
| Increase reasoning | Force multi-step rather than single-step thinking | "Require the answer to combine at least three independent facts" |
You apply one operator per iteration. A seed like "list three benefits of exercise" might evolve, after five rounds, into something like "for a sedentary office worker in their 40s with mild hypertension, compare three exercise regimes by VO2-max impact, joint-loading risk, and time cost, and justify which to start with assuming a 30-minute weekday window." Same shape, much harder.
The resulting WizardLM models (and the derivative WizardCoder for code, WizardMath for math) consistently outperformed Alpaca on benchmarks that test multi-step reasoning, which is exactly the dimension Evol-Instruct was designed to amplify.
Which of these is NOT one of the four canonical Evol-Instruct operators?
#Persona-driven synthesis: the 1B-persona trick
Scaling Synthetic Data Creation with 1,000,000,000 Personas
Chan, Wang, Yu, et al. (Tencent AI Lab) (2024)
Built a billion-persona library and used it to condition synthetic data generation. Demonstrated dramatic diversity gains on math, code and tool-use benchmarks.
The pipeline:
- Build a large persona library. Persona Hub describes ~1 billion personas as Cartesian products of occupation × interests × demographic features × situational context. A persona might be "a 52-year-old veterinary surgeon in rural Wales who is learning Python to automate clinic billing."
- Condition every synthetic prompt on a sampled persona: "Write an instruction this person would realistically send to an AI assistant."
- Generate the response as normal — possibly with Evol-Instruct on top.
- Filter and dedupe globally.
The diversity gain is dramatic. Without personas, "ask the model to generate an instruction" produces a heavy concentration of generic tasks. With personas, you cover the actual surface of human queries — including the long tail that benchmarks struggle to measure but real deployments live in.
What is the main benefit persona-conditioned synthesis adds over plain Self-Instruct?
#Distillation from strong teachers
A great deal of synthetic-data work is dressed-up distillation: have a strong model generate data, train a smaller model on it. This is how Alpaca was built; it is how a great deal of the open-weights chat ecosystem still works. The recipe is simple, but it has two failure modes you have to engineer around.
The standard mitigations: mix synthetic with a substantial real-data tail, regenerate from the base teacher periodically rather than chaining student-of-student-of-student, filter aggressively with a quality classifier, and use diverse personas / decoding temperatures to avoid mode collapse.
#Math: synthesise then verify
For mathematics, synthetic data is uniquely powerful because math has a cheap, exact verifier — you can check whether the final number is correct. MetaMath (2023) and the MathInstruct family pioneered the pattern.
MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
Yu, Jiang, Shi et al. (2023)
Augments MATH and GSM8K via rephrasing, backward reasoning, and self-verification. The verifier filter is the load-bearing piece.
The recipe:
- Start with a real corpus (MATH, GSM8K, NumGLUE).
- For each problem, ask the teacher LLM to generate multiple solution chains with explicit chain-of-thought.
- Use the same teacher as a verifier — give it the problem and a candidate solution, ask "is this correct?"
- Keep only the (problem, chain-of-thought, answer) triples where the verifier agrees and the final numeric answer matches a reference (when available).
- Round-trip filter: paraphrase the question; run the solver again; demand the same final answer. Inconsistency under paraphrase is a strong signal of luck rather than reasoning.
Magicoder: Empowering Code Generation with OSS-Instruct
Wei, Wang, Lin et al. (2023)
Conditions instruction generation on real open-source code seeds, then verifies via execution. The execution loop is what makes the synthetic data trustworthy.
#Self-Rewarding: closing the loop with no humans
Self-Rewarding Language Models
Yuan, Pang, Cho et al. (Meta) (2024)
Iterative DPO using the model itself as the preference judge. No human labellers needed beyond the seed instructions.
The loop:
- Start with seed instructions (Self-Instruct style).
- Sample multiple response candidates per instruction.
- Use the same model as an LLM-as-judge to rank the candidates against a rubric (helpfulness, correctness, conciseness).
- Construct preference pairs (chosen, rejected) from the rankings.
- Run DPO on these synthetic preferences.
- The post-DPO model is both a better responder AND a better judge — iterate.
The Meta paper showed that this iterative loop produces consistent gains over multiple rounds, with the model improving as both responder and grader. The human input is minimal — only the seed instructions; the rest is the model judging itself.
Self-Rewarding training (Yuan et al., 2024) requires what kind of human input?
#Model collapse: when the snake eats its tail
AI models collapse when trained on recursively generated data
Shumailov, Shumaylov, Zhao et al. (2024)
Published in Nature, July 2024. Showed empirically and theoretically that recursive training on synthetic outputs causes loss of tail-of-distribution coverage and eventual mode collapse.
What happens, and Shumailov's group showed this in language, image, and Gaussian-mixture settings — is that the distribution narrows. Tails disappear first. Rare events, rare vocabulary, rare reasoning paths drop below the model's representational threshold and never come back. After enough generations the model is fluent but pathologically repetitive, having forgotten the part of the data manifold that was never well-represented in the synthetic outputs.
The empirical demonstrations in the paper are striking. For a Wikipedia-trained OPT-125M model resampled across nine generations, the model's outputs on the prompt "some books that some recommend to others" devolve from coherent recommendations into surreal repetition about jackrabbits within five generations. The math case is even cleaner — the Gaussian-mixture variance progressively shrinks until the support is a single mode.
What is the canonical defence against model collapse?
#Constitutional / safety synthesis
Synthetic data is also how modern safety training works. The Constitutional AI line at Anthropic was the first widely-known instance.
Constitutional AI: Harmlessness from AI Feedback
Bai, Kadavath, Kundu et al. (Anthropic) (2022)
Replaced human harmfulness labels with model-generated critique-and-revise pairs guided by a written constitution. The synthetic preference data drove RLAIF training.
The pipeline turns safety alignment into a self-supervised synthetic-data exercise:
- Generate a response to a potentially harmful prompt — let the assistant answer freely.
- Critique the response against a written constitution (a list of principles — be helpful, avoid harm, respect autonomy, etc.). The critic is the same model, in a different prompt role.
- Generate a revised response that addresses the critique.
- Use (revised, original) as a synthetic preference pair: chosen = revised, rejected = original.
- Train via DPO / RLAIF on the resulting preference dataset.
This is synthetic-data generation pointed at alignment instead of capability. The human input shrinks to the constitution itself — a few hundred words. Anthropic's Claude line was the first commercial product to train its safety behaviour predominantly on data of this shape; the technique has since been widely adopted under various names (RLAIF, self-critique, principle-guided revision).
#A practical pipeline
If you are actually shipping a synthetic-data pipeline today, the canonical eight-step recipe looks like this:
- Define the skill surface. What does the student need to be good at? Enumerate domains, task types, reasoning patterns, formats.
- Curate seeds. Real examples per skill. Hundreds, not millions.
- Generate at scale using a strong teacher LLM. Use personas for breadth and Evol-Instruct for depth.
- Filter aggressively. Educational-content classifier, perplexity gate, format validator, near-duplicate removal (MinHash / SimHash), toxicity / PII screens, length caps.
- Verify. Code → execution. Math → calculator + round-trip paraphrase. Prose → LLM-as-judge with a rubric. Without a verifier, you are gambling.
- Mix with real data. Typical frontier mixes today are 50-90% synthetic, with the rest real for tail-of-distribution grounding. The Phi line is the high-end of synthetic share; most other labs sit lower.
- Train and evaluate. Eval on held-out benchmarks AND on red-team probes targeting known synthetic-data failure modes.
- Iterate on the gap. Identify skills where the student underperforms. Generate more data targeting those skills. Loop.
#Cost arithmetic: why synthetic wins
The economic argument is brutal. Suppose you need 1 trillion tokens of high-quality instruction-tuning data.
| Source | Unit cost | Total | Notes |
|---|---|---|---|
| GPT-4o-mini at $0.60 / 1M output tokens | $0.60 / M | ~$600,000 | Wall-clock: days, parallelised across many API keys. |
| Human labellers at $25 / hour, ~500 tokens / hour quality output | $50 / 1K tokens | ~$50,000,000 | Wall-clock: years of human-time. |
| Human experts (lawyers, doctors, PhDs) at $150 / hour | $300 / 1K tokens | ~$300,000,000 | Wall-clock: longer; expert recruitment dominates. |
That is ~100x in favour of synthetic vs general human labour and ~500x against domain experts, at comparable token counts. The catch, repeatedly, is verification. The $600K cost gets you raw tokens; getting them to ship-quality requires the verifier infrastructure (execution sandboxes, math checkers, judge-LLMs, classifiers) on top. Even with the verifier overhead, the cost gap to human labelling is generally one to two orders of magnitude.
#Key takeaways
Key Takeaways
- The public high-quality web is effectively exhausted; frontier models are now >50% synthetic data, and Phi-4 disclosed ~100% synthetic.
- The Phi recipe = curated seed concepts + GPT-4 textbook-style generation + aggressive quality filtering. It is the public proof that small models can punch up if fed better data.
- Self-Instruct (Wang 2022) bootstraps SFT data from ~175 seed instructions. Evol-Instruct (Xu 2023) iteratively raises difficulty via four operators: add-complexity, deepen, concretize, increase-reasoning.
- Persona-driven synthesis (Persona Hub, 2024) recovers the long tail of human queries by conditioning generation on diverse personas.
- Synthetic-math and synthetic-code rely on EXACT verifiers (execution, calculators, round-trip paraphrase). Without verification you amplify the teacher's biases.
- Self-Rewarding (Yuan 2024) and Constitutional AI (Anthropic 2022) close the RLHF loop with the model judging itself — humans only write seed instructions and principles.
- Model collapse (Shumailov 2024) is the canonical failure mode of recursive synthetic training. Defences: mix with real data, regenerate from base, use diverse personas, dedupe aggressively.
- Cost: ~$600K for 1T synthetic tokens via GPT-4o-mini vs ~$50M for the same volume of human labels. ~100x advantage, but only if your verifier infrastructure is sound.