LLM Data & Pretraining: Common Crawl to FineWeb
After this lesson, you will be able to:
- Trace the lineage of open pretraining corpora — Common Crawl → C4 → The Pile → RefinedWeb → RedPajama → Dolma → FineWeb / FineWeb-Edu / DCLM — and identify the curation move that distinguishes each
- Walk through a modern pretraining cleaning pipeline: language ID, boilerplate stripping, heuristic filters, KenLM perplexity filters, classifier filters, PII redaction, and the trade-offs of toxicity filtering
- Apply MinHash + LSH near-duplicate detection on a small document set and explain why dedup is the single biggest lever in data-quality work
- Reconcile Chinchilla (D ≈ 20N) with the modern over-training regime, and design a data-mixing schedule that weights code, math, and web text deliberately
- Make defensible decisions on vocabulary size, byte-level fallback, tokenizer training corpus, and document packing, and explain tokenizer fertility's equity implications
Before You Start
Don't be misled by the headlines about model architectures — the 2024-2025 frontier-lab consensus is that data work matters more than architectural innovation for the next 1-2 generations. The Chinchilla rule tells you how much data; this lesson is the missing piece on which data and why.
#Where Does the Data Come From?
The map of pretraining-data sources splits into three buckets:
| Source family | Examples | Why it matters |
|---|---|---|
| Web crawls | Common Crawl, OSCAR | The volume backbone. ~250 TB/month for CC, multilingual, messy. |
| Curated text | Wikipedia, Project Gutenberg, ArXiv, StackExchange, PubMed | High quality, low volume. Punches above its weight on benchmarks. |
| Code | The Stack, Stack v2 (BigCode), GitHub crawls | Roughly 5-15% of modern training mixes; improves reasoning. |
#Common Crawl: the firehose
#Curated derivatives (the lineage)
#Code corpora
#The legal landscape
You're training a 7B model. Chinchilla-optimal token count is roughly...
#Cleaning Pipeline: The Unsexy 80%
Below is the canonical 2026 cleaning order — roughly the order FineWeb, Dolma, and DCLM all use, with minor variations. Skip any stage and you pay for it twice: in training compute and in downstream eval.
#1. Language identification
Why this matters: a 5% leak of low-confidence French into your "English" corpus quietly degrades your tokenizer fertility, distorts your downstream evals, and confuses your data-mix ratios.
#2. Boilerplate removal
HTML pages are 30-70% boilerplate (nav bars, footers, cookie banners, sidebars, "related articles"). Three open extractors dominate:
- trafilatura. Slow but high-quality; the FineWeb default
- resiliparse. Faster Rust extractor; the DCLM default
- justext. Older, faster, lower quality
The CC default WET extraction is poor — re-extracting from WARC with trafilatura or resiliparse alone moves downstream eval by several points.
#3. Heuristic filters
Cheap rule-based filters that strip obvious garbage before anything expensive:
| Filter | Threshold | What it catches |
|---|---|---|
| Length | <50 words OR >100K words | Stubs and dumps |
| Mean word length | <3 or >10 chars | Garbled OCR, machine-generated nonsense |
| Punctuation ratio | <0.12 | List-only pages, code dumps |
| Symbol-to-word ratio | >0.1 | Math/code disguised as prose, currency dumps |
| Ellipsis lines | >30% of lines end in "..." | Search-result excerpts |
| Repetition | top-2gram >0.2 of doc | Repeating boilerplate |
| Bullet/numbered lines | >90% | Stack Overflow listings, recipe sites |
| Stopword ratio | <2 of {the, of, and, a, in, to} | Not actually English |
#4. KenLM perplexity filter (CCNet)
Train a 5-gram language model (KenLM is the standard) on a high-quality reference corpus — typically Wikipedia. Score every web document by its perplexity under that model. Keep documents in the middle perplexity band: very low perplexity is template/boilerplate, very high is gibberish, the middle is real prose.
#5. Classifier-based filter (FineWeb-Edu)
The 2024 quality-filter breakthrough: train a small classifier (DistilBERT-scale, sometimes a 1B Llama distillation) to predict "is this educational" on a few hundred thousand documents rated by a strong model. Then run it across the full corpus and keep the top X%.
FineWeb-Edu used Llama 3 70B to rate 500K documents on a 0-5 educational-value scale, trained a small classifier on those ratings, and kept the top ~30% of FineWeb. The result: a 1.3T-token subset that beats raw 15T FineWeb at small training budgets by 5-10 MMLU points. This is the single most impactful 2024 data-curation result.
DCLM-baseline-1.0 used a similar move with a slightly different rating axis. The pattern is now standard.
#6. PII redaction
#7. Toxicity / safety filter
#Deduplication: The Biggest Single Lever
Why dedup matters: a duplicate document trained 10 times is mostly wasted compute (the gradient updates collapse), inflates memorization risk (Carlini et al. 2022 showed memorization scales with duplication frequency), and skews downstream eval (because eval sets leak into training data via duplicates). Dedup is the most-replicated "free win" in pretraining data work.
#Exact deduplication
- URL-level dedup. The easiest 30% win. Two crawls of the same URL produce two copies. Hash the URL, keep the latest fetch.
- Exact-line dedup. Drop lines that appear in many documents (cookie banners, license boilerplate, footer text). The Pile and C4 both did this.
- Exact-document hash dedup. SHA-256 the document, drop exact matches. Cheap but only catches byte-identical duplicates.
#Near-duplicate detection
The real problem: two news articles republished with a one-line attribution change, scraped versions of Wikipedia with different cookie banners stripped, paraphrased clickbait. Three families of solutions:
- MinHash + LSH. The field standard. Hash document n-grams ("shingles"), compute a MinHash signature, bucket signatures via Locality-Sensitive Hashing, only do exact Jaccard inside buckets. Sublinear cost.
- SimHash (Charikar 2002) — a single 64-bit fingerprint per document; near-duplicates have low Hamming distance. Used by Google for web dedup. Less common for LLM corpora than MinHash but cheaper at very large scale.
- Suffix array methods (Lee et al. 2022, "Deduplicating Training Data Makes Language Models Better") — find all substrings of length k that appear more than once across the corpus, delete one copy. Strictly more thorough than MinHash, but harder to scale.
#Cross-document vs intra-document dedup
- Cross-document. Different documents that are mostly the same. The big category. Catches near-duplicates of news articles, mirrored content, scraped copies.
- Intra-document. Repeating phrases within one document. The Gopher repetition filters catch this; suffix array methods catch both at once.
#Empirical impact
The FineWeb release report documents a clean 20% downstream-score lift from a strong dedup pass (MinHash with 128 hashes, 14-band LSH at ~0.7 Jaccard threshold) over a URL-only dedup baseline. SlimPajama observed similar lifts over RedPajama. The 2025 consensus: a serious dedup pass is the highest-ROI data-work intervention available.
You have two 1000-word documents that share 950 words verbatim but the 50 unique words are different (e.g., a news article republished with a different byline and intro paragraph). Will exact-hash dedup catch them as duplicates?
#Hands-On: MinHash + LSH Deduplication
Time to build MinHash + LSH from scratch on a small set of near-duplicate documents. This is the canonical "interview question that's actually how the field works" exercise.
A 100M-document corpus. URL-level dedup removes 30% of duplicates. Why bother with MinHash + LSH after that?
#Hands-On: KenLM-Style Perplexity Filtering
KenLM is the production tool, but you can build a tiny n-gram language model in pure numpy and demonstrate the CCNet filter behavior on a handful of documents.
#Quality vs Quantity: Chinchilla and After
You read the Chinchilla rule in the scaling-laws lesson: D ≈ 20N is compute-optimal at a fixed training budget. The 2024-2025 evidence layered on top:
-
Llama 3 broke Chinchilla deliberately. Llama 3 8B was trained on 15T tokens — about 100x past Chinchilla-optimal. Llama 3 70B saw ~6T tokens, ~4x past. The reason: inference economics. Over-trained small models serve cheaper.
-
DCLM showed data quality can substitute for data quantity. DCLM-baseline at 4T high-quality tokens matches or beats RedPajama at 1.2T raw tokens on downstream eval — even at the same model size. Quality filtering can buy you 2-3x in effective tokens.
-
The "10x token rule" is the modern over-training heuristic. For a model that will be heavily served at inference, train it on ~10x the Chinchilla-optimal token count. A 7B model: ~1.4T optimal, ~14T in over-trained practice. Llama 3 went further still on 8B because the team had compute headroom and the curve kept moving.
#Data mixing schedules
The mix matters as much as the volume. The modern recipe is roughly:
| Source | Share (typical) | Why |
|---|---|---|
| Web (CC-derived) | 50-70% | Volume backbone |
| Code | 5-20% | Reasoning gains, code capability |
| Math / academic (ArXiv, books) | 5-10% | Reasoning, technical capability |
| Wikipedia / encyclopedic | 3-5% | Factual grounding |
| Books / long-form | 5-15% | Long-context coherence |
| StackExchange / Q&A | 1-5% | Instruction-following baseline |
| Multilingual web | 5-20% | Per-language eval |
Curriculum scheduling is a 2024-2025 area of active work: train on noisier web data first, switch to higher-quality and code-heavy data in the final ~10% of tokens. Llama 3 disclosed a curriculum-style annealing phase; DeepSeek-V3 documented an explicit phase-shift in the data mix. The signal: late-stage data quality matters disproportionately because the late gradients shape downstream eval most heavily.
You're filtering a 50T-token web crawl down to 5T tokens via FineWeb-Edu-style classifier filtering. Same 7B model trained on the 5T filtered set vs the 50T unfiltered set. Which wins on MMLU?
#Tokenizer Training: A Data Decision in Disguise
#Vocab size
Modern practice has converged on 32K-128K:
| Model | Vocab | Algorithm |
|---|---|---|
| Llama 3 | 128K | Tiktoken-style BPE |
| Mistral | 32K | SentencePiece BPE |
| GPT-4 (cl100k) | ~100K | Tiktoken BPE |
| Claude | ~100K | Custom BPE |
| Gemma | 256K | SentencePiece |
| Qwen 2.5 | 152K | BPE |
vocab_size × d_model, so a 256K vocab at d=8192 is a 2.1B-parameter embedding table — significant for small models, negligible for 70B+.#Byte-level fallback
<UNK> for any input. The trade-off: rare Unicode characters (emoji, CJK ideographs) require multiple byte tokens.#Tokenizer training corpus
This is the under-appreciated decision: train your tokenizer on a representative slice of your eventual training data. If your training data is 15% code, train your tokenizer on a 15%-code slice — otherwise you get bad code tokenization that costs you across the whole training run. Llama 3 specifically retrained its tokenizer with more code than Llama 2 to fix this.
#Tiktoken's regex pre-tokenization
OpenAI's tiktoken uses an explicit regex to split text into chunks before applying BPE. The cl100k_base regex looks like:
(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\r\n\p{L}\p{N}]?\p{L}+|\p{N}{1,3}| ?[^\s\p{L}\p{N}]+[\r\n]*|\s*[\r\n]+|\s+(?!\S)|\s+
This is opinionated: numbers get split into 1-3 digit chunks (which is why GPT-4 fails on long-number arithmetic), contractions get split off, and whitespace is handled explicitly. The choice of pre-tokenization regex propagates directly into model behavior — it is why "12345" and "1234567" tokenize very differently.
#Tokenizer fertility
| Language | Fertility (tokens/word) |
|---|---|
| English | ~1.0 |
| French | ~1.2 |
| Spanish | ~1.2 |
| Chinese | ~1.5 (per character) |
| Japanese | ~1.6 (per character) |
| Hindi | ~2.5-3.0 |
| Arabic | ~2.5 |
| Burmese / Khmer | ~5-8 |
That fertility number directly translates into: cost per API call, effective context window, and number of forward passes the model must do for the same semantic content. Building a tokenizer that minimizes mean fertility across your intended user base is an equity decision dressed up as a hyperparameter.
Why is FineWeb-Edu's classifier-based filter so effective compared to heuristic + perplexity filters?
#Document Packing and Sequence Assembly
The final stage before bytes hit the GPU. Your training data is millions of documents of wildly varying length; your model trains on fixed-length sequences (2K, 4K, 8K, 32K depending on the model). Two choices:
#EOS-separated packing
#Document-attention-mask packing
Pack documents into a fixed-length sequence but mask attention so each document only attends to its own tokens. Plus optionally reset position embeddings at document boundaries. Pros: cleaner training signal. Cons: more bookkeeping in the data loader, and many open-source training scripts don't implement it correctly.
The 2024-2025 consensus drift: document-attention masks are worth the implementation effort for serious training runs. Llama 3, DeepSeek-V3, and Qwen 2.5 all use them. Early open-source models (Llama 1/2, Mistral) used EOS-separated packing and accept the noise.
#Position resetting
Within a packed sequence, do you reset the position index to 0 at each document boundary? RoPE positional encodings make this cheap and clean. Most modern implementations reset. Some still don't, and the resulting "wraparound" position artifacts can show up as weird in-context-learning behavior near the start of the second document.
#Putting It Together
A 2026 frontier-lab pretraining pipeline, end-to-end:
- Acquire — Common Crawl WARCs (250-400 TB/month), GitHub mirror, ArXiv dump, Wikipedia dump, licensed sources (news, books), opt-out filtered with C2PA / ai.txt.
- Extract — trafilatura or resiliparse on WARC → plaintext.
- Language ID — fastText lid.176, threshold ~0.65 per target language.
- Heuristic filters — Gopher + RedPajama-V2 quality signals.
- Perplexity filter — KenLM 5-gram trained on Wikipedia.
- Classifier filter — DistilBERT-scale "educational quality" classifier trained on a strong model's ratings.
- PII redaction — Presidio.
- Deduplication — URL dedup → MinHash + LSH (128 hashes, ~14 bands) at 0.7-0.85 Jaccard threshold → exact-line dedup pass.
- Mix — weighted blend of web/code/math/books/wiki/multilingual.
- Tokenizer train — 32K-128K SentencePiece or BPE on a representative slice with byte fallback.
- Pack — 4K-8K sequences with document-attention masks and position resets.
- Shuffle — global shuffle, then training-time micro-batch shuffle.
The actual training run, the part that draws the headlines, is roughly steps 13-15 of a pipeline that's mostly data engineering.
#Key Takeaways
- The open-data lineage runs C4 → The Pile → RefinedWeb → RedPajama → Dolma → FineWeb → DCLM. Each step is a curation move, and FineWeb's 15T-token release (April 2024) reset the open-source bar against Llama-3-scale closed data
- Dedup is the highest-ROI data work. URL dedup is the easy 30% win, MinHash + LSH catches the long tail of near-duplicates, and FineWeb's report documents roughly 20% downstream-score lift from a strong dedup pass alone
- Classifier-based quality filtering beats heuristic filtering by a wide margin. FineWeb-Edu's small classifier trained on Llama-3-rated educational value produces a 1.3T-token subset that beats the raw 15T set by 5-10 MMLU points at small budgets
- Chinchilla (D ≈ 20N) is the floor; over-trained small models are the new default. Llama 3 8B at 15T tokens (~100x past Chinchilla) reflects inference economics, not a contradiction of the scaling law
- Tokenizer choices are equity choices in disguise. A tokenizer trained on English-heavy data charges Hindi, Burmese, and Arabic users 2-5x more in tokens, API cost, and effective context; mixing the tokenizer training corpus to match your target users is the single most impactful equity intervention in pretraining
#Quick Check
Why is MinHash combined with LSH (Locality-Sensitive Hashing) instead of just computing exact Jaccard for every document pair?