How to check if a benchmark score can be trusted
Before you trust a benchmark score, check that its grader accepts correct answers and rejects wrong ones, that its tasks are clear and that the model shows no sign of having seen them during training. Review a sample of tasks by hand, try to pass the grader without solving them and look for leaked answers. Then recalculate the score on the tasks that pass every check.
What a benchmark score counts
The benchmarks in this guide are fixed sets of tasks, each scored by an automatic grader such as a set of tests, an exact answer or another AI model acting as judge. The grader decides whether each answer passes or how well it does, and the score is the share of tasks that pass or the average grade. A training environment for reinforcement learning has the same parts (a setting such as a code repository, tasks and a grader), and the model learns from the grader's verdicts while it trains.
The score therefore depends on three things besides the model. The grader has to accept every correct answer and reject every wrong one, each task has to be clear enough to solve and the model must not have seen the tasks or their answers during training.
What audits of one benchmark found
SWE-bench Verified is a set of 500 coding tasks that OpenAI published in August 2024. OpenAI calls it "a standard metric reported in frontier model releases", meaning announcements of the newest large models. In February 2026 OpenAI said it had stopped reporting the score. It had audited 138 tasks that its o3 model did not solve consistently, and it found that "59.4% of the 138 problems contained material issues in test design and/or problem description". Every frontier model it tested could also reproduce the original fix or exact details of the problem for some tasks, which OpenAI reads as a sign that all of them had seen some of the problems and solutions during training.
A June 2026 preprint found the opposite problem in the same benchmark. In a sample of 49 tasks, 28.5% had tests weak enough that a code change confirmed to be incorrect still passed. Across 134 model submissions, models scored about 14 percentage points higher on the tasks the author flagged as easy to pass this way than on other tasks of similar difficulty.
The same weakness is a risk in training, because a model rewarded for passing a weak grader can learn to pass the grader instead of solving the task. Epoch AI interviewed people who build and buy training environments, and they named resistance to this kind of gaming as the most important quality of an environment.
Four outcomes for each task
| Outcome | What it does to the score |
|---|---|
| The grader passed it, and the task was solved | Counts correctly |
| The grader passed it, but the task was not solved | Raises the score without any real progress |
| The grader failed it, but the task was solved | Lowers the score and punishes a correct answer |
| The grader failed it, and the task was not solved | Counts correctly |
How to check a score yourself
You can run this check with your own team if the people reviewing the tasks did not build them.
- Write down how each task is graded. Note what the grader checks, such as tests, an exact answer or a second AI model acting as judge, and what it ignores.
- Take a random sample of tasks and have someone who knows the subject solve or review each one. Mark whether the task is clear and whether the grader accepts a correct answer.
- Try to pass the grader without solving the task. Submit empty, partial and deliberately wrong answers, and count how many pass.
- Look for leaked answers. Give the model only the task's name or the start of the problem, and see whether it reproduces the reference answer or details it could not know otherwise.
- Recalculate the score on the tasks that passed every check, and report it next to the published score with the number of tasks behind each.
You end with the published score and the score on the tasks you checked side by side, plus a list of tasks to rewrite or remove.
What to ask whoever reports the score
- Which tasks and which version of the benchmark does the score use?
- How does the grader decide that an answer passes, and has anyone checked its verdicts by hand?
- Were the tasks or their answers public before the model was trained?
- How many runs is the score based on, and how much does it change from one run to the next?
- Can we see the tasks the model failed and the ones it passed?
Sources
- OpenAI, "Why SWE-bench Verified no longer measures frontier coding capabilities", 23 February 2026, checked 1 October 2026
- Shreshth Rajan, "Auditing reward hackability in code RL training environments", arXiv preprint 2606.16062, 14 June 2026, checked 1 October 2026
- JS Denain and Chris Barber, Epoch AI, "An FAQ on reinforcement learning environments", 12 January 2026, checked 1 October 2026
Before you rely on a score
If you build, buy or sell a model or a training environment and a decision depends on its score, tell us the decision and when you need to make it. We reply within 24 hours.
Contact us