Evaluating LLMs in 2026: MMLU-Pro, GPQA, SWE-Bench & Agentic Evals
After this lesson, you will be able to:
- Read a frontier-model benchmark table critically — know which evals are saturated, which leak, and which still produce signal at the 2026 frontier (MMLU-Pro, GPQA Diamond, FrontierMath, SWE-Bench Verified, LiveCodeBench, RULER, GAIA, MMMU, Chatbot Arena Elo)
- Compute pass@k correctly with the unbiased Chen-et-al estimator, and understand why pass@1 systematically understates a model's true capability when verification is cheap
- Apply real statistical tests to benchmark results — bootstrap confidence intervals and the Wilcoxon signed-rank test — so you can tell whether 'Model A beats Model B by 1.2 points on a 50-question subset' is signal or noise
- Spot benchmark contamination, judge-model bias, and decoding-config tricks in published leaderboards, and design a defensible evaluation for your own production task that survives 6 months
Before You Start
#Part 1: Why the Old Benchmarks Are Dying
A short tour of the casualties:
- MMLU (Hendrycks 2020): 57 subjects, four-choice MCQ. Frontier models score 85–91%. The differences are smaller than the noise floor of prompt-template choice. Still useful as a floor — any model below 70% is unfit for general use — but not as a ranking signal.
- HumanEval (Chen 2021): 164 Python problems with hidden unit tests. GPT-5, Claude Opus 4.7, and o3 all sit in the 95–98% band. Multiple analyses have shown direct contamination signals in pretrained models. Saturated and suspect.
- HellaSwag, ARC, TruthfulQA: still cited, mostly as quick sanity checks. Saturation at the top makes them poor discriminators for frontier models, though they remain useful for small/open models below the frontier.
- GSM8K (Cobbe 2021): 8.5k grade-school math word problems. Non-reasoning frontier models hit 92–95%; reasoning models hit 97%+. Saturated for the frontier; still relevant for smaller models.
#Part 2: Knowledge and Reasoning Evals
The benchmarks below are the current discrimination set for general knowledge and reasoning at the frontier.
#MMLU-Pro
#GPQA Diamond
#Humanity's Last Exam
Anthropic + Center for AI Safety (2025). Roughly 3,000 questions in math, science, humanities, and reasoning, written by domain experts and explicitly calibrated against frontier models. Designed to remain difficult even as MMLU-Pro and GPQA saturate. By early 2026 frontier scores sit in the single-to-low-double digits — this is currently the hardest broad academic benchmark in the public canon.
#FrontierMath
#AIME, MATH, AlphaProof / AlphaGeometry 2
#GSM8K and the Saturation Floor
GSM8K is now what MMLU was three years ago — a floor, not a discriminator. Use it to verify a model can do basic arithmetic word problems; don't try to rank with it.
GPT-5 scores 78% on MMLU-Pro. Claude Sonnet 5 scores 80%. You evaluate both on a 50-question random subset of MMLU-Pro from your own machine. Which statistical test should you use to claim Sonnet beats GPT-5 on this subset?
#Part 3: Code Evals
Code is the capability that closed the loop fastest. HumanEval was the gold standard in 2021–2023 and is fully saturated. The new canon:
#LiveCodeBench
#SWE-Bench, SWE-Bench Verified, SWE-Bench Lite
- SWE-Bench Verified (500 issues): an expert-curated subset where OpenAI human contractors confirmed each task is well-specified and the hidden tests are correct. This is the version every lab now reports.
- SWE-Bench Lite (300 issues): a smaller, easier subset for faster iteration. Lower discrimination at the top but cheaper to run.
#HumanEval+ / MBPP+ (EvalPlus)
Liu et al. (2023). Same prompts as HumanEval and MBPP, but with 80x more aggressive test cases. Edge cases that the original tests missed reveal another 5–15 points of weakness in models that scored 95% on HumanEval. A useful reality check on the "we got 97%" headlines.
#Other Worth Knowing
- BigCodeBench. Wider API usage (1,140 tasks across 139 libraries); harder than HumanEval, tests integration not just algorithms.
- CodeContests (DeepMind) — competitive programming problems with rich test suites; harder, lower scores, useful for reasoning models.
- APPS (Hendrycks) — older, broader; still cited as a baseline.
You're picking between SWE-Bench Lite and SWE-Bench Verified for the eval of your new code agent. Which choice makes most sense?
#Part 4: Long-Context Evals
Frontier APIs in 2026 advertise 200K (Claude Haiku 4.5), 1M (Claude Sonnet 5 and Claude Opus 4.7, both at standard pricing with no long-context premium; Gemini 2.5 Pro / 3), and up to 10M (Llama 4) context windows. But advertised context length is not the same as effective context length. Real-world recall degrades long before you hit the documented maximum.
#Needle in a Haystack (NIAH)
#RULER
NVIDIA (Hsieh 2024). The serious long-context benchmark. 13 tasks across four categories: retrieval (multi-needle, multi-key, multi-value), multi-hop tracing, aggregation, and question-answering — all at controllable context lengths. RULER exposes the real cliff: even by 2026, Claude Sonnet 5, GPT-5, and Gemini 2.5 Pro still degrade measurably between 128K and 1M context (less than they did in 2024, but the cliff is real), and most open-weights models degrade much earlier. The headline 10M-token claims (Llama 4) collapse under RULER-style stress.
#LongBench, LOFT, InfiniBench
- LongBench (Bai 2023): diverse tasks (summarization, multi-document QA, code completion) at long context. Practical and well-cited.
- LOFT (Long-Context Frontiers): multi-hop tasks that require integrating retrieval results inside a long context window. Tests whether your "1M context" model can actually beat a smaller model + RAG pipeline.
- InfiniBench: 1M+ token tasks; useful for the Gemini 2.5 Pro / Gemini 3 / Claude Opus 4.7 1M regime.
#Part 5: Agentic Evals: the Frontier of "Frontier"
#SWE-Bench Verified (revisited as agentic)
#TheAgentCompany
#OSWorld
#WebArena and VisualWebArena
Autonomous web browsing on fully self-hosted clones of GitLab, Reddit, an e-commerce site, OpenStreetMap, and a CMS. The agent reads HTML (or screenshots, in VisualWebArena) and executes actions. Multi-step shopping, social-media browsing, and content-management tasks. Frontier scores: 20–35% across the suite.
#GAIA
#AgentBench
8 distinct agentic scenarios (operating system, database, knowledge graph, card games, etc.). Earlier and broader than the others; useful for comparing scaffolding choices, not for ranking frontier models.
#ARC-AGI-2
Chollet et al. (2025). The successor to ARC-AGI-1, designed to stay hard for o3-class reasoning models. It still tests visual abstract reasoning from few examples, but with a fresh private set and harder compositional puzzles. Frontier scores entered 2026 in the 5-15% range with extreme test-time compute budgets — currently the most-watched single-number measure of "general fluid intelligence" in LLMs.
#Part 6: Multimodal Evals
- MMMU (Yue 2023): 11,550 college-level multimodal questions across 30 disciplines. The current frontier multimodal benchmark; scores 65–75% for top closed models.
- MathVista: visual math problems — geometry, charts, function plots. Tests whether the model can read a math diagram, not just describe it.
- MMStar: vision-saturation tests; many existing multimodal benchmarks have visual leakage where the model can answer without looking at the image. MMStar filters out those items.
- VideoMME: video understanding across short, medium, and long clips; tests temporal reasoning.
- EgoSchema: long-form first-person video with 3-minute clips; tests sustained attention.
- VQAv2: older, still cited for compatibility.
#Part 7: Open-Ended Generation and Preference Evals
Static benchmarks measure capabilities you can grade automatically. Real-world chat is open-ended; you need preference signal.
#LMSYS Chatbot Arena (now LMArena)
#MT-Bench
80 multi-turn prompts, scored by GPT-5 / Claude Opus 4.7 as judge on a 1–10 scale. Cheap and reproducible. Useful for fast internal iteration; less trustworthy than Arena for cross-lab claims because of judge-model bias.
#AlpacaEval and Arena-Hard
#LLM-as-Judge Pitfalls
LLM judges are cheap and scalable but biased in known ways:
- Position bias: the model in position A tends to win more often than the model in position B, regardless of content. Mitigated by running each pair in both orders and averaging.
- Verbosity bias: longer responses are rated higher even when content is equivalent. Mitigated by length-controlled scoring (AlpacaEval 2.0 LC).
- Self-preference bias: GPT-4 judges think GPT-4 outputs are better; Claude judges prefer Claude. Mitigated by using a different judge family than the candidates, or by averaging over multiple judges.
- Format bias: judges over-reward Markdown formatting, bullet points, and headers. Watch this in production: a model that learns to format pretty can win MT-Bench without actually being better.
You're using GPT-5 as a judge to score outputs from GPT-5-mini, Claude Haiku 4.5, and Llama-4-Scout. Which result should make you suspicious?
#Pairwise vs Pointwise Judging
- Pairwise: judge sees two responses, picks the better one. More reliable; calibrates the judge's threshold against itself.
- Pointwise: judge sees one response, scores it 1–10. Cheaper (N evaluations vs N choose 2), but ratings drift over time and across runs.
Default to pairwise for cross-model comparisons; use pointwise for tracking a single model's improvement over training checkpoints.
#Part 8: Safety, Reliability, and Truthfulness
The frontier-eval suite for safety in 2026:
- HarmBench (Mazeika 2024): 510 harmful behaviors across 7 categories with automated classifier-based scoring. The mainstream red-team eval.
- JailbreakBench: standardized jailbreak evaluation with a frozen attack/defense protocol.
- ALERT: categorical safety eval across 6 macro-categories and 32 micro-categories.
- TruthfulQA (Lin 2021): older but still relevant; tests whether the model gives false-but-popular answers ("Does swallowed gum stay in your stomach for 7 years?").
- HaluEval: hallucination detection across QA, summarization, and dialogue.
- CRASH: common-sense robustness — adversarial perturbations of CommonsenseQA-style inputs.
For any product that touches medical, legal, or financial advice, run at least HarmBench plus your domain-specific guardrails before shipping.
#Part 9: Reasoning Evals in the Post-o1 Era
The reasoning-model evaluation suite is overlapping but distinct from general LLM eval:
- AIME 2024 / 2025: 30-question math olympiad qualifiers; AIME 2025 is the fresh contamination-free split. Reasoning models live in the 85–98% range here; non-reasoning models 15–50%.
- MATH-500: 500-problem subset of MATH used as a fast reasoning benchmark.
- GPQA Diamond with chain-of-thought scoring: report both the final answer accuracy and the reasoning quality (judged separately).
- FrontierMath: Tier 1 in the mid-40s to 50% by 2026 for the best reasoning models; Tier 2 / Tier 3 (added 2025) still in the single digits — the new headroom benchmark.
- Humanity's Last Exam: 3000 expert-curated questions (Anthropic + CAIS 2025); frontier scores low double-digits — currently the hardest broad academic benchmark.
- ARC-AGI-2 (Chollet 2025): the successor to ARC-AGI-1, designed to stay hard for o3-class reasoning models. Frontier scores 5–15% with extreme test-time compute.
- Codeforces / LiveCodeBench: competition programming, where reasoning models meaningfully outperform standard models.
#pass@k and the Sampling-Budget Trade-off
You have a code-generation model that solves 30% of problems on a single try (pass@1 = 0.30). You can afford to generate 100 candidates per problem and run the unit tests on each. What is your approximate pass@10 — i.e., the chance that at least one of 10 randomly chosen candidates passes?
#Test-Time Compute Trade-offs
#Part 10: Methodology: How to Actually Evaluate
Knowing benchmark names is useless if your statistics are wrong. The methodology section below is what separates a defensible eval from a screenshot from Twitter.
#Reporting Hygiene
Always specify:
- Decoding config (greedy vs temperature=T, top_p, max_tokens). Greedy decoding for benchmark reporting; temperature sampling for pass@k.
- Prompt template (zero-shot, few-shot with N examples, chain-of-thought yes/no, system prompt). Same template across all models being compared.
- Aggregation (mean across questions, macro-average across categories, micro-average). Be explicit.
- N (number of questions in the subset). A claim on 50 questions has a different confidence interval than the same claim on 5,000 questions.
#Statistical Significance
- McNemar's test (paired binary outcomes): tests whether two models disagree in a biased direction. Simple, well-suited for accuracy comparisons.
- Wilcoxon signed-rank test (paired ordinal/continuous): more general; works for binary, integer scores, or continuous metrics.
- Bootstrap confidence intervals: sample with replacement from your N questions B=10,000 times; compute the metric on each bootstrap; the 2.5th / 97.5th percentiles give a 95% CI. Works for any metric — accuracy, F1, BLEU, win-rate, even composite scores.
Avoid the independent two-sample t-test for paired benchmark results. It ignores the pairing and inflates the standard error.
#Eval Cost is Real
Running MMLU-Pro (12k MCQ) on a 70B-class model is several dollars; running SWE-Bench Verified end-to-end with agent scaffolding is hundreds. Reasoning models burn 10–100x more tokens. Plan the eval budget like you plan the training budget. A common mistake: design a 50,000-example eval and then never run it because it's too expensive — better to run a 500-example eval thoroughly with statistical hygiene than a 50,000-example eval once with no error bars.
#Contamination Detection
Tools and signals to watch for:
- Substring search in pretraining corpora (when available). HuggingFace has open-llm-contamination tools for many open datasets.
- Perplexity asymmetry: a contaminated model has much lower perplexity on memorized benchmark text than on similar held-out text.
- Rephrased-question performance gap: if a model scores 90% on the original MMLU and 65% on syntactically rephrased versions, that's a memorization signal.
- Use rotating benchmarks (LiveCodeBench, contamination-resistant sets like FrontierMath, your own held-out internal set).
Claude Sonnet 5 beats GPT-5 by 1.5 points on MMLU (90.3 vs 88.8) but loses to GPT-5 by 30 Elo on LMArena. Why does this happen? Pick the BEST explanation.
#A Working 2026 Eval Stack
For a new product launch in 2026, the defensible eval stack is roughly:
- Floor: MMLU-Pro + GSM8K (sanity — any model scoring below the floor is unfit).
- Knowledge / reasoning: GPQA Diamond + AIME (for reasoning models) + MATH-500.
- Code: SWE-Bench Verified + LiveCodeBench (rotating split postdating training cutoff).
- Long context: RULER at your application's actual context length.
- Agentic: GAIA + the relevant slice of AgentBench / WebArena / OSWorld / TheAgentCompany for your domain; add ARC-AGI-2 and Humanity's Last Exam as frontier-reasoning stress tests.
- Multimodal (if applicable): MMMU + MathVista.
- Preference: LMSYS Arena Elo (consume) + your own pairwise comparisons (build).
- Safety: HarmBench + your domain-specific guardrails.
- Internal: 200–500 held-out examples from your real task, refreshed quarterly, scored with a strong LLM judge plus regex/exact-match checks, with bootstrap CIs and paired Wilcoxon tests on every model comparison.
The last item, your internal eval, is the one that actually matters for product decisions. Public benchmarks are how you compare to other labs. Internal evals are how you decide whether to ship.
You see a paper claim: 'Our model achieves 94.3% on MMLU, +1.1 over GPT-5.' What's the FIRST question you should ask?
#Key Takeaways
- The old canon is dead at the frontier. MMLU, HumanEval, GSM8K, HellaSwag are saturated and contamination-suspect; they remain useful as floors but not as ranking signals between frontier models.
- The current discriminators are MMLU-Pro, GPQA Diamond, FrontierMath, SWE-Bench Verified, LiveCodeBench, RULER, GAIA, MMMU, and LMSYS Arena Elo — each measures a distinct capability slice; no single number is sufficient.
- Pass@k is not pass@1 with a multiplier. Use the unbiased Chen-et-al estimator. When verification is cheap, pass@10 tells a very different (and more deployment-relevant) story than pass@1.
- Use paired statistics. When two models answer the same questions, the data is paired; bootstrap confidence intervals and Wilcoxon signed-rank tests are the minimum hygiene for any "A beats B" claim.
- LLM-as-judge is biased in known ways. Position bias, verbosity bias, self-preference bias, format bias. Mitigate with both-order pairwise scoring, length-controlled metrics, and cross-family judges.
- Your internal eval is what matters for shipping. 200–500 held-out examples from your real task, refreshed quarterly, with statistical hygiene, will be more trustworthy than any public leaderboard.
#Quick Check
Why is GPQA Diamond more discriminating at the frontier than MMLU?