Skip to content
An Estimated 6.49% of MMLU Questions Contain Errors, a Fresh Maths Test Cut Some Models' Scores by up to 8 Points, and the Llama 4 Model Meta Entered on LMArena Was Not the One It Released
Artificial Intelligence

An Estimated 6.49% of MMLU Questions Contain Errors, a Fresh Maths Test Cut Some Models' Scores by up to 8 Points, and the Llama 4 Model Meta Entered on LMArena Was Not the One It Released

Illustration by tuput

English

Researchers who re-checked MMLU estimate 6.49% of its questions contain errors. Annotators hired through Scale AI wrote 1,205 new school maths problems, and some open models scored up to 8 points lower on them than on the public GSM8K test. LMArena said Meta should have made clearer that its Llama 4 entry was customised for human preference.

· · 7 min read

MMLU, a 57-subject multiple-choice exam used to score language models, contains errors in an estimated 6.49% of its questions, according to a June 2024 re-check by 16 researchers. In the virology section, 57 of the 100 questions they examined had a problem. What follows are the documented cases of benchmarks wearing out, leaking into training data or being gamed, who measured each figure, and what the benchmarks’ own authors say about their limits.

When a test runs out of room

By 3 June 2024 the authors of MMLU-Pro, a harder successor to MMLU, wrote that performance on such benchmarks “has begun to plateau.” MMLU-Pro has 12,032 questions with ten answer options, three times as many wrong choices as MMLU. Its builders dropped 5,886 MMLU questions that more than four of eight open models answered correctly, so only 6,810 of its questions come from MMLU and the rest come from other sources. The authors report accuracy falling 16% to 33% against MMLU on the models they tested, and sensitivity to prompt wording falling from 4 to 5% to 2%. For GPT-4o they report that step-by-step reasoning lifted MMLU-Pro from 53.5% to 72.6% but MMLU only from 87.2% to 88.7%.

Stanford’s 2026 AI Index says “Evaluations intended to be challenging for years are saturated in months.” It adds that frontier models gained 30 percentage points in one year on Humanity’s Last Exam, a test designed to be hard for AI.

Errors in the answer key

A wrong answer key caps the best possible score below 100%. The MMLU-Redux team, with Aryo Pradipta Gema as first author, checked 100 random questions from each of the 57 subjects, 5,700 in all out of 14,042. They sorted problems into unclear questions, unclear options, no correct answer, several correct answers and wrong keys. Virology had the worst record at 57%, with 33 wrong keys. Logical fallacies had 26%, college chemistry 25% and professional law 18%. The 6.49% figure is their estimate for the whole test, extrapolated from the sample. The authors note that 8,342 questions went unchecked and that their protocol “might still be prone to the personal biases of the annotators.”

A new set of maths problems

GSM8K is a set of school-level maths word problems. On 1 May 2024 Hugh Zhang and colleagues released a paper built around GSM1k, 1,205 new problems written by human annotators hired through Scale AI, with no language model used to write them. On GSM1k, accuracy fell by up to 8 points on some models. The largest gap was Yi-6B-Chat, from 43.7% to 35.7%. The Phi and Mistral families scored lower on GSM1k across almost every release, while Gemini, GPT and Claude models showed little sign of overfitting. The authors also found a correlation between how likely a model was to produce GSM8K examples verbatim and its score gap (a Spearman r-squared of 0.36). They caution that the two sets are “only highly similar, but not identically distributed.” The authors withheld GSM1k at first to keep it from leaking into training data.

How contamination gets measured

Contamination means test material leaked into a model’s training data. Three papers attack it from different angles.

Chunyuan Deng and colleagues asked models to guess a wrong answer option that had been masked out of a multiple-choice question. On MMLU, ChatGPT and GPT-4 matched it exactly 52% and 57% of the time. The authors read this as a sign of exposure but call the method less reliable than direct retrieval. Yonatan Oren and colleagues tested whether a model prefers a benchmark’s original question order over shuffled versions. Applied to five public models, the test found little evidence of widespread contamination. Ruijie Xu and colleagues checked 31 models on maths tests using perplexity and n-gram scores and reported substantial instances of test-set misuse. For a case where the model’s maker ran nearly every test, see our look at on-device models.

From HumanEval to problems with dates on them

LiveCodeBench, posted on 12 March 2024 as a replacement for older coding tests such as HumanEval, takes problems from LeetCode, AtCoder and Codeforces and records when each was published, so a model can be tested only on problems released after its training cutoff. Its authors report that DeepSeek’s code models dropped sharply on LeetCode problems released after August 2023. They write that HumanEval is “an easier benchmark with small and isolated programming problems and thus easier to overfit on.” They add that the benchmark covers only Python and that problem sampling alone moves scores by 1 to 1.5%, so small gaps between models mean little.

SWE-bench Verified, a human-filtered set of software tasks, drew the same scrutiny. Epoch AI rated it “Flawed” on 3 September 2026. It cites OpenAI’s audit of 23 February 2026, which covered 27.6% of the tasks and found tests that rejected correct fixes in 59.4% of those, a floor of 16.4% of all tasks. Epoch says it has reasonable confidence that more than 20% of questions have scoring problems, and that all tasks and solutions are now public. Our report on coding agents covers the maintainer review of passing patches and Scale AI’s SWE-Bench Pro board.

The Llama 4 entry on LMArena

Meta’s Llama 4 announcement of 5 April 2025 said Llama 4 Maverick came with “an experimental chat version” that scored an Elo of 1417 on LMArena, a ranking built on human preference votes. On 7 April Meta executive Ahmad Al-Dahle called the claim that Meta trained on test sets “simply not true,” according to TechCrunch.

LMArena posted its own statement on 8 April. It said Meta’s interpretation of its policy “did not match what we expect from model providers,” and that Meta should have made clearer that Llama-4-Maverick-03-26-Experimental was a customised model built to optimise for human preference. It released 2,000 or more battle results for review and said it was updating its policies.

On 29 April 2025 Shivalika Singh and colleagues published “The Leaderboard Illusion.” It counts as many as 27 private Meta variants on the arena before the Llama 4 release, and estimates that Google and OpenAI received 19.2% and 20.4% of test prompts against 29.7% for 83 open-weight models combined. In a fine-tuning test, a 7-billion-parameter model trained on a mix with 70% Arena data raised its win rate on the Arena-Hard prompt set from 23.5% to 49.9%, a 112.3% relative gain, while its MMLU score slipped from 66.5% to 65.9%. The authors call their Meta count conservative, say they cannot see private variants’ scores, and disclose that several authors work at Cohere. They recommend banning score retraction and setting a stated cap on private variants per provider.

Arena replied on 9 May 2025. It disputes the paper’s open-model share, putting open models at 40.9% in its 27 April statistics. It says the paper’s chart of pre-release testing gains was a simulation unrelated to arena data, and that the fine-tuning test used a static 500-prompt benchmark judged by a model. It also said its pre-release testing policy has been public since 1 March 2024, and it promised to allow multiple pre-release variants and to label early scores provisional.

Exams built to resist, and their own audits

Humanity’s Last Exam (HLE) was posted on 24 January 2025 with 2,500 questions. Each was tested against frontier models and rejected if they answered it correctly. The paper’s table lists o3-mini (high) at 13.4% on the text-only questions, the best of eight models, and a private held-out set checks for overfitting. On 23 July 2025, FutureHouse checked 321 text-only chemistry and biology questions with a literature-search agent, had experts review 150 of its outputs, and estimated that 29% (plus or minus 3.7 points) had answers in direct conflict with peer-reviewed literature. FutureHouse’s update says the HLE team’s own follow-up found about 18% of a similar subset problematic, and that at least one of three expert reviewers disagreed on 25%. The team plans a rolling revision.

ARC-AGI, introduced by François Chollet in 2019, is a set of puzzles meant to be accessible to people and hard for AI. The ARC-AGI-2 paper of 17 May 2025 reports human testing with 407 participants. Its table, as of 14 May 2025, lists 3.0% as the best score on the semi-private set. The authors say accuracies below 5% are generally not treated as meaningful. In a report of 1 May 2026, NIST’s testing centre CAISI, which calls that set “held-out and uncontaminated,” measured 79% for GPT-5.5 and 63% for Anthropic’s Opus 4.6 on it. Our piece on Chinese models covers the rest of that report. CAISI says its figures are a mean across tasks, which differs from the benchmark’s official scoring.

Share
Copied!

Sources & further reading

  1. Wang and others, MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark (arXiv:2406.01574, 3 June 2024)
  2. Gema and others, Are We Done with MMLU? (arXiv:2406.04127, 6 June 2024)
  3. Zhang and others, A Careful Examination of Large Language Model Performance on Grade School Arithmetic (arXiv:2405.00332, 1 May 2024)
  4. Deng and others, Investigating Data Contamination in Modern Benchmarks for Large Language Models (arXiv:2311.09783, NAACL 2024)
  5. Oren and others, Proving Test Set Contamination in Black Box Language Models (arXiv:2310.17623, 26 October 2023)
  6. Xu and others, Benchmarking Benchmark Leakage in Large Language Models (arXiv:2404.18824, 29 April 2024)
  7. Jain and others, LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code (arXiv:2403.07974, 12 March 2024)
  8. Epoch AI: Benchmark review, SWE-bench Verified (3 September 2026)
  9. Meta AI: The Llama 4 herd (5 April 2025)
  10. TechCrunch: Meta exec denies the company artificially boosted Llama 4's benchmark scores (7 April 2025)
  11. LMArena statement on Llama 4, posted on X and reproduced by Simon Willison (8 April 2025)
  12. Singh and others, The Leaderboard Illusion (arXiv:2504.20879, 29 April 2025)
  13. Arena (formerly LMArena): Our Response to 'The Leaderboard Illusion' Writeup (9 May 2025, updated 13 June 2025)
  14. Stanford HAI: 2026 AI Index Report, Technical Performance
  15. Phan and others, Humanity's Last Exam (arXiv:2501.14249, 24 January 2025; version 11, 28 July 2026)
  16. FutureHouse: About 30% of Humanity's Last Exam chemistry/biology answers are likely wrong (23 July 2025)
  17. Chollet and others, ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems (arXiv:2505.11831, 17 May 2025)
  18. NIST CAISI: Evaluation of DeepSeek V4 Pro (1 May 2026)

Researched and written with the help of AI tools and edited for accuracy. Provided for general information and discussion only, not professional advice. See our editorial standards and disclaimer. Spotted an error? Tell us.

#ai benchmarks#benchmark contamination#mmlu#gsm8k#swe-bench#lmarena#humanity's last exam#arc-agi

Enjoyed this? Get the next one.

One good read at a time, straight to your inbox. No spam, unsubscribe anytime.

More in Artificial Intelligence
Scale's SWE-Bench Pro Board Puts the Best Coding Agent at 61.5%, Maintainers Reject Many Fixes That Pass the Tests, and METR Says Its Latest Trial Cannot Show How Much Time Agents Save
Benchmarks say agents are strong. Maintainers, a randomised trial and a pile of incident reports say the picture is narrower. Here is what each number actually measures.
Mistral's 1-Trillion-Parameter Le Chonk Scored 38 on an Independent AI Index, Eighth Among Open Models, Before Its Weights Have Even Shipped
France's Mistral has put a trillion-parameter model online and says the downloadable weights follow this month. Outside testers rank it the best open model outside China and eighth among open models, and Mistral's own figures for its size do not all agree.
Two Supreme Court Rulings in 2026 Set Aside Orders Built on Fake AI Citations, While Indian Courts Use AI for Translation and Transcription and the Court's Draft AI Rules Are Still Not Final
Parliamentary replies and court records show which AI tools are live, which are still pilots, and what judges are barred from doing with them. Three Supreme Court orders show what happens when invented case law gets through.
← all articles