Executive Summary
This article goes back over where a received view came from, the view that frontier models cannot do physics. A team based mainly at Yale opened up six widely used physics benchmarks and had physicists re-grade, one at a time, the responses those benchmarks had recorded as wrong. Of the 250 responses filtered out as wrong, twelve were cases where the model had actually got the physics wrong. In the rest, either the item itself was defective or the grader failed to recognize a correct answer.
The team fixed the answer keys, removed the items it could not fix, and measured again. The scores rose sharply. Not all of the rise is owed to the re-grading, however. One benchmark lost nearly half its items, and to that extent the gain includes the effect of clearing away hard problems. The firmest piece of evidence in the study sits elsewhere: a benchmark that kept its items almost intact and had only its answer key repaired also gained a great deal. And the design of the audit itself was not flawless either, which the paper writes down in its own appendix.
No model's measured error rate can fall below the defect rate of the test paper. The better a model gets, the more the errors left on the scoreboard belong to the paper rather than to the model. The same question carries over to the score tables an organization studies when it weighs whether to adopt AI. Who reviewed those answers, what is the defect rate, and has anyone written that number next to the score?
12
model errors among the audited wrong answers
Out of 250 responses recorded as wrong. The other 238 came from the test paper
143
cases where the item itself was defective
Out of the same 250. The problem statement or the answer key was wrong
61.0 → 87.2%
score on the benchmark where only the answer key was fixed
CMT-Benchmark. Its item count fell only from 50 to 49
0
open theoretical physics problems solved by an agent setup alone
Saturation is a claim confined to closed-form, problem-set items
Recounting the 250 wrong answers
The question put by a paper posted to arXiv on 11 September 2026 is a short one. Are frontier models really unable to do physics? The place the question comes from is just as plain. Recent models scored low on widely cited physics benchmarks, and those scores circulated as an impression that physics is still hard. The paper, which carries 51 authors, does not set out to overturn that impression. It opens up the place where the scores were made. Responses the benchmarks had recorded as wrong were pulled out one at a time and graded again by people.
The team had its reasons for doubting those scores. The paper's introduction first sets down what had been happening in mathematics over the same stretch of time. Report after report described frontier models, and the agents built on top of them, contributing new proofs and solutions to open mathematical problems, and the introduction lists those cases. Even so, the paper draws a line: none of it accounts for the physics scores. Physics adds the work of building a model, abstracting, and settling on which assumptions to hold. The mismatch the introduction treats as decisive lies elsewhere. Physicists who have used these models in their own research report capabilities that do not square with the benchmark scores. That is the spot the paper's central question came from.
The data auditors who did the reviewing number 39, and most of them are physics doctoral students. Twelve more people are listed separately as physics advisors, whose role Appendix G describes as discussing the physics of their own subfields and vouching for the graduate students as auditors. The two lists are not exclusive, however. Cross-check the 39 auditors against the 12 advisors and five people turn up on both. The body of the paper describes the group only as "a team of physicists, primarily based at Yale University." Some of the authors are not physicists, so rendering the study as "51 physicists reviewed it" gets the head count and the roles both wrong.
1.1The denominator comes first
The figure from this paper most likely to be misquoted is 250. It is not the total item count. The audit run across four benchmarks covered 502 responses in all, and 252 of those were scored correct and never reached an auditor. What people looked at again were the remaining 250 recorded as wrong. Every proportion that follows should therefore be read with "of the 250 wrong answers" attached to it. A sentence saying that 95% of physics problems are defective is one this study never wrote.
How the sample was chosen carries a caveat as well. Across the four benchmarks, the audit was narrowed to items on which every attempt by one model had been scored wrong. That choice cut the burden of hands-on review, and the paper records it in §2.3. So these 250 are not a random sample but a sample tilted toward what looks hard. The defect rates below, then, are not the defect rates of the benchmarks as a whole.
The number of attempts allowed varied by benchmark too. Appendix B.3 shows that HLE-Physics items were attempted up to five times each, with tools permitted only on the fifth attempt. PHYBench allowed up to five attempts and no tools. PRISM-Physics and UGPhysics gave each item a single attempt. Items missed five times running and items missed once thus sit together inside the same 250. What an auditor saw on screen was also not every response to an item, but one stored response and the final binary verdict the original grader had issued. On top of that, this audit run is a separate run from the one that produced the pre-audit scores in Table 1. The selection rule and the response budget both differ, and the paper insists that the attribution counts in Appendix C not be read as a reconstruction of the Table 1 scores.
1.2Three kinds of wrong
The part of this method that actually earns its keep is not the arrows but this classification. For each wrong answer, the auditors sorted the responsibility for the error into one of three kinds. The paper writes the definitions out itself, and letting the names drift would collapse the argument, so this article holds to the paper's own three labels.
- Model error: the problem is well posed and the reference solution is correct, but the model under test gives an incorrect answer. In the paper's words, "only these count as genuine model errors."
- Grader error: the problem and the answer key are both sound and the model gives a correct answer, but the evaluator marks it incorrect.
- Benchmark error: the problem statement or the reference solution is itself defective. A wrong answer key, inconsistent conditions, ambiguous wording and a missing assumption all land here.
Sorted and counted, the 250 came out as 143 benchmark errors and 95 grader errors. Twelve were cases where the model had actually got the physics wrong. Add the first two together and 238, which is 95.20% of the re-graded wrong answers, sat on the test paper's side rather than the model's. Broken out by benchmark, the shape of the breakage also turns out to differ from one to the next.
Attribution of the 250 responses recorded as wrong and sent to audit. Source: arXiv:2609.13009, Appendix C, Table 2. Bar length is proportional to the count. CritPt and CMT-Benchmark, which were audited in full, are absent from this tally because their audit scope differs, and the paper states separately that the tally is not a reconstruction of the Table 1 scores.
Gather up only the three benchmarks drawn from public problem banks and the slope gets steeper. Of 152 cases, 148, or 97.37%, sat on the test paper's side, and four were model errors. In the PRISM-Physics sample, all 74 audited wrong answers belonged to the test paper. Not a single point in that sample was docked because the model got the physics wrong. These 74, however, came out of a run that gave each item one attempt. That run offers no assurance that the count would hold with more attempts.
Turning the figure of twelve into "models now hardly ever get physics wrong" throws the denominator away. Twelve is a count inside the 250 wrong answers that went to audit, and the 252 responses scored correct in the same run, together with every item outside the sample, have nothing to do with it. This article can claim one thing only. When people look again at responses recorded as wrong, most of them turn out to have come from the test paper rather than from the model's physics reasoning.
This is not the first time. Back in July the Pebblous blog covered a study that re-graded the failures of a data-engineering benchmark and found answer-key errors. There, 82.7% of the tasks recorded as failures had a grading or answer-key problem, and the score rose by a little under ten percentage points. This case adds three things to that piece. The effect size has grown into the thirty-point range; the count taken was the absolute number of the model's own mistakes rather than a defect rate; and the paper writes down in an appendix what the auditors themselves were looking at when they judged.
This is how the test paper was broken
Proportions alone tell you the size of the breakage and nothing about its shape. The paper's appendices carry the actual scenes along with the notes the auditors attached, and reading them puts within reach what "the test paper was wrong" really means.
2.1Scenes where a right answer was recorded as wrong
The first scene is a multiple-choice item asking for an electron's time of flight. The options were 330 nanoseconds, 66 nanoseconds and 33 nanoseconds. Instead of picking one of the three, the model answered: "Cannot be determined from the given information; no option can be uniquely selected." The grading came back wrong. Yet the note left by the reviewer who looked at the item again runs to a single line. "Not enough information is given in the problem." The model had put its finger squarely on the defect in the question, and it lost the point for being right about it.
The second and third scenes come from the same benchmark and concern its grader. The model's answer was T = P/(3√6) and the answer key carried T = (√6/18)P. Rationalize the denominator and the two are exactly the same expression. The score the grader assigned was zero. Even on a scale built to award partial credit it came out at 0.0. The reviewer note is one sentence: "Expressions are algebraically the same." In another item the model wrote Fmax = 5mRω² where the key had Fmax = 5mω²R. The same scalar factors in a different order, and zero again.
The fourth scene is the answer key itself being wrong. In one item the reference solution set up the algebra correctly and then botched the final arithmetic. The reviewer note reads, "The final arithmetic is performed incorrectly, despite having the correct algebraic expression. They should obtain 2−√2 instead." The model's answer was exactly 2−√2, and the value printed in the key was −√2. Other types show up as well. One condensed-matter theory item never wrote down the Hamiltonian, leaving it unclear whether the coupling constant attached to the Pauli operators or to the spin operators. The two conventions split the energy by a factor of four, so, in the paper's words, "the numerical ground-state energy is not uniquely defined." A label that left out its conventions had pulled the answer four ways.
Repairing the answer key also produced cases where the model's original answer turned out to be the closer one. One item's reference answer listed two entries before the repair and three after it. The model had originally answered with two entries, and on the repaired problem it got all three on the first attempt. A single missing assumption had been changing the correct answer itself.
The appendix carries defects of another texture too. One item asks for the minimum speed of a flowing medium given the speeds of two spacecraft. The appendix notes that the expression printed in the answer key simplifies to zero, and the reviewer note says only, "Reference answer equals 0, which is incorrect." An expression that always returns zero was sitting in the answer slot of a question about a minimum. Another item asks whether, in damped oscillation, a particle can return to equilibrium faster when the resistance falls below the critical value, and its key says no. The reviewer wrote that critical damping gives the fastest decay without overshoot rather than the fastest possible route to equilibrium, and added that the question may have meant "without overshoot" but never said so. The first is arithmetic gone wrong. The second is a sentence two competent people can read two ways.
The scene where the grader travels furthest is in the appendix as well. In a problem with four sub-parts, covering a diver's motion from the board and down through the water, the model's final answer and the key's final answer pointed at the same value in all four. They parted over the use of an equals sign on a rounded value, the order in which the factors of a fraction were written, and whether the speed was written V or dx/dt. The reviewer note is one sentence, "The two responses are identical," and the binary score the grader awarded was zero.
2.2All six makers wrote that they had verified it
Here this report steps outside the paper for a moment. We opened the original papers of the six audited benchmarks to see how each set of makers described the verification of its own answer key. All six say they verified it. The wording varies, but the direction is single: experts built it, the answers are unambiguous, and a machine grades them exactly.
| Benchmark | What its makers wrote | What this audit found |
|---|---|---|
| HLE | "Each question has a known solution that is unambiguous and easily verifiable" | 86 benchmark errors among 98 wrong answers (87.76%) |
| CMT-Benchmark | "The dataset was designed and verified by a panel of expert researchers from around the world" | Benchmark errors in 30 of all 50 items (60.00%) |
| CritPt | "Every problem is hand-curated to admit a guess-resistant and machine-verifiable answer" | Benchmark errors in 21 of the 56 audited challenges (37.50%) |
| PHYBench | "a systematic curation pipeline to eliminate flawed items": expert review plus model-assisted detection and a large-scale evaluation by 81 students, compressed to a 66.1% acceptance rate | 40 grader errors among 56 wrong answers (71.43%) |
| PRISM-Physics | A "fully rule-based" symbolic equivalence method by which "we ensure consistent validation across diverse formulations without heuristic judgments" | 48 grader errors among 74 wrong answers (64.86%) |
| UGPhysics | Items "all rigorously screened for data leakage", with a judgment pipeline "ensuring accurate evaluation" | 18 benchmark errors among 22 wrong answers (81.82%) |
The self-descriptions are taken from the abstract or methods section of each benchmark's original paper (HLE arXiv:2501.14249 · CMT-Benchmark arXiv:2510.05228 · CritPt arXiv:2509.26574 · PHYBench arXiv:2504.16074 · PRISM-Physics arXiv:2510.03185 · UGPhysics arXiv:2502.00334, checked 15 September 2026). The denominators in the right-hand column differ by benchmark. For the top four the denominator is the responses recorded as wrong; for CMT-Benchmark and CritPt it is the full item set or the whole audited set. The column cannot be summed or averaged down its length.
HLE stands out most in this table. Its abstract says the solutions are "unambiguous and easily verifiable," while the review procedure described in the same paper reads: submissions were so advanced and specialized that "reviewers were not expected to verify the full accuracy of each provided solution rationale if it would take more than five minutes, instead focusing on whether the question aligns with guidelines." And the same passage closes like this. "Given this limitation in the review process, we welcome community feedback."
The defects were not hidden. The makers wrote them into their methods sections, and nobody read them as a warning. This paper is less an indictment than the first reply to that invitation. The audit put HLE-Physics at the highest benchmark-error rate of the four, which also means the published account of the limitation had been pointing at the real defect rate all along.
Holding a self-description up against the material was not always possible, either. For three of the six, the audit team never got hold of the answer key or the grading code. PHYBench had reference solutions and worked solutions public for only 100 of its 500 items, and the paper notes in a footnote that it asked the authors for the rest and did not receive them. CMT-Benchmark had reference solutions, but its authors' automatic grading code was unavailable, so the team used the HLE-adapted evaluator instead. CritPt keeps its official reference solutions private, so the auditors worked the problems themselves and built a new answer key. The scope of the audit was set less by the team's design than by how much material had been published.
2.3The loudest claims about grading came with the most grader errors
Among the graders alone, the order inverts. The two benchmarks that made the largest claims for the rigor of their grading rank first and second in grader-error rate. PHYBench introduced its own expression edit distance metric and wrote that it lifted sample efficiency by 204% over binary grading; in the audit its grader-error rate came out at 71.43%. PRISM-Physics wrote that its rule-based matching secures consistent validation across diverse formulations, and came out at 64.86%. The HLE-adapted evaluator, which makes no separate claim about grading accuracy, came out at 4.08%, the lowest of any grader this paper audited.
The claim that rules deliver consistency collapses at exactly that point. The two scenes above are the evidence. One answer differed from its key by a rationalized denominator, the other by the order of two factors. A person sees one answer where the rule reads two. A rule that cannot recognize equivalence is wrong steadily and without wavering. Consistency is no warrant of accuracy, and it can just as well mean that the error reproduces.
One choice on the auditors' side is folded into that 71.43%, and leaving it out of the sentence inflates the number. This paper set aside PHYBench's partial-credit design and counted an answer correct only when its edit-distance score was full marks. Honor the partial credit and some of those cases would not have been classed as wrong. Both of the scenes above, even so, scored 0.0 under partial credit. Defects remain that the decision to binarize does not explain.
UGPhysics's 13.64% carries a caveat as well. That benchmark layers an auxiliary model judge on top of a rule-based evaluator, and before running the pre-audit the paper swapped the original, older judge model for a frontier model. So 13.64% is not the value of the pipeline as it stood, but the value after the judge had been raised a generation. Nor is the best grader in this audit sitting at 4.08% any ground for comfort. It works out to one error in every twenty-five. Scores have been circulating while nobody measured the error rate of the grader.
That is also why the paper audited the grader and the answer key separately. Pipelines that hand off to a model judge when rule-based matching fails are common, and a model judge brings failures of its own: position bias, verbosity bias, and low accuracy on objective reasoning tasks. Nor does the choice of judge settle whether the reference solution is correct. The two layers are independent of each other, and so the paper kept the audit of items and reference solutions apart from the audit of response grading. One more thing belongs on the record. All six corrected scores were graded by a single evaluator, the HLE-adapted pipeline that came out best at 4.08%. Even the repaired scores were not measured with a defect-free ruler.
Why the scores went up
This study will travel mostly as an arrow. On one benchmark a score went from 47.3% to 78.7%, and on another it went from 61.0% to 87.2%. Inside a single arrow, though, three different operations are stacked on top of one another. The answer keys were repaired, the responses were graded again, and the items that could not be repaired were dropped outright. The paper does not break out how much of the rise each of the three delivered. So the reader has to do the breaking out.
A word on the metrics first. mean@4 is the average accuracy across four attempts at the same problem, and pass@4 counts an item correct when any one of the four attempts gets it. The pass figure is therefore always equal or higher. Every number in the table below is a mean@4, and all of them come from one model.
3.1The item count belongs beside the score
The paper attaches the caveat itself, in its abstract. "Corrected scores are computed on the retained evaluation subsets following expert review." §3 of the body is blunter: "Corrected evaluations use the retained or repaired question sets (Appendix B), so pre-audit and corrected scores are not always computed on the same questions." The item count therefore has to stand next to the score.
| Benchmark | Items | mean@4 | Change |
|---|---|---|---|
| PHYBench | 100 → 87 | 26.50 → 90.23 | +63.73pp |
| PRISM-Physics | 100 → 74 | 13.00 → 94.59 | +81.59pp |
| UGPhysics | 100 → 82 | 83.00 → 92.07 | +9.07pp |
| HLE-Physics | 202 → 116 | 47.28 → 78.66 | +31.38pp |
| CMT-Benchmark | 50 → 49 | 61.00 → 87.24 | +26.24pp |
| CritPt | 70 → 54 | 32.29 → 87.50 | conditions differ |
Source: arXiv:2609.13009, Table 1. The figures are GPT-5.6-Sol's mean@4, except for CritPt's pre-audit value, which is a mean@5 measured by an outside aggregator. The three models' scores cannot be set side by side to read off a ranking. On the three public problem banks every model answered without tools, and on the three expert-authored sets only two models used tools.
The item-count column on its own shows what happened. HLE-Physics started at 202 items and finished at 116. Eighty-six items came out, and the ones that came out were not chosen at random but judged defective in the audit. Since the audited set was itself made of items the model had missed on every attempt, the natural reading is that the hard items were concentrated among the ones that left.
Item counts before and after correction. Source: arXiv:2609.13009, the item-count column of Table 1 and the per-benchmark audit procedures in Appendix B. CMT-Benchmark repaired 29 of its 50 items and excluded only one; CritPt repaired 19 challenges and excluded two.
3.2Only one comparison is clean
So the paper's argument rests on the benchmark that lost a single item, not on the biggest arrow. CMT-Benchmark was audited across all 50 items, defects turned up in 30, 29 of those were repaired, and exactly one was excluded. Of this benchmark the paper writes, "The same HLE-adapted evaluator is used before and after correction," and adds, "The corrections are to the benchmark materials only." With the item set and the grader both left essentially in place, the score went from 61.00% to 87.24%.
The explanation that scores rose because items were removed does not survive this benchmark. One item was removed. A rise of more than 26 percentage points came out of repairing answer keys and problem statements alone. If only one arrow is to be quoted, it should be this one rather than 47.3% to 78.7%.
The figures from the two benchmarks audited in full point the same way. CMT-Benchmark had defects in 30 of its 50 items, and CritPt in 21 of the 56 challenges audited. With every item opened rather than only the sample of wrong answers, the defect rate still sits in the 30 to 60 percent range. The objection that a hard sample produced a high defect rate is one the paper had already absorbed inside its own data.
3.3CritPt has to sit apart
At the other end sits CritPt. Its arrow is of a different character from the other five. The pre-audit score was not measured by the authors but is a mean@5 taken by an outside aggregator on 70 challenges, while the corrected score is a mean@4 on the 54 challenges the authors' pipeline retained. And since the official reference solutions were unavailable, in the paper's words, the "corrected evaluation uses reference solutions derived independently by our auditors." The items changed, the metric changed, the party doing the grading changed, and the answer key changed.
The corrected pass@4 of 94.4% on CritPt is probably the most provocative figure in the paper. This article does not use it as a headline. It is a metric that counts an item correct when one of four attempts lands, and the pre-audit pass@4 it would be compared against does not exist in the first place. That cell of the paper's table is empty. A figure with no arrow available should not be read as the tip of one.
To put it together: calling the whole rise an effect of re-grading overstates the case, and calling it an effect of item removal understates it. The blend of the two effects differs by benchmark, and the one case where the blend can be taken apart is CMT-Benchmark, which lost a single item. When quoting any other benchmark's arrow, the safer course is to write the item count into the same sentence.
The floor a defect rate sets
The scores in this paper will age. One line in its §5 discussion will not. Every benchmark carries some share of defective items, and on those items a model is marked wrong whatever it answers. And so "no model's measured error rate can fall below the defect rate." The defect rate is the floor of the measurement, and the floor stays there no matter what the model does.
That much is arithmetic. The interesting part comes next. While a model is weak, the floor stays out of sight. In the paper's words, "While a model still gets many valid items wrong, the defects add only a little to its error." Then the model's true error rate drops below the defect rate and the picture inverts. "Once its true error rate drops below the defect rate, most of the errors on its scorecard are the benchmark's, not the model's." The paper takes the view that the true error rate of frontier models has already fallen far below the defect rate of the benchmarks available.
Conceptual diagram, rendering the argument of the §5 discussion in arXiv:2609.13009 as a picture. The axes carry no ticks because the paper never measured defect rate and true error rate against an axis of time.
The paper points, inside its own data, to places where this inversion has already happened. In the PRISM-Physics sample not one audited wrong answer was a model error, and in UGPhysics there was one. In those two benchmarks the recorded wrong answers are, for practical purposes, all defects. HLE-Physics and PHYBench have not reached that point, but the paper notes that even there the raw error rate inflates the model's own share by roughly an order of magnitude, which is to say around tenfold.
The prescription the paper offers in the same section is modest. Small repairs, filling in context an item left out or giving the model more attempts, would fix most of the reported failures. The second half of that catches the eye. Raising the attempt budget does not repair the test paper; it changes the way the measuring is done. How much of the reported failure belongs to that share is something the paper never measured separately.
4.1Why it surfaced now
The defects were not created yesterday. The conditions under which they show have changed. Graders go wrong mostly by marking a correct answer wrong, and a weak model does not often produce a correct answer. Which leaves the grader few chances to be wrong. As capability climbs and correct answers turn common, the grader's mistakes accumulate and the gap between the measurement and the reality opens up. The paper carries this account over from prior work, and it connects to an older complaint: a fixed benchmark loses statistical power in the high-accuracy range. Moving from 98% accuracy to 98.1% removes the same proportion of error as moving from 80% to 81%, but detecting that difference costs roughly an order of magnitude more evaluation data.
Statistical power is not all the paper takes from that prior work. As summarized there, the earlier study laid out four conditions a usable benchmark should meet, and two of them land squarely on the ground this audit covers. One is the power just described; the other is reliable annotation. The definition attached to the second interlocks with this audit immediately. Items that merely carry a wrong label have to be held apart from items that have no single determinate answer to begin with. The earlier study named two routes by which the latter arise: the problem is underspecified, or competent people read its sentences differently. The three-way classification this audit used to sort wrong answers, and the disagreement among auditors that section 5 gets to, were already inside that definition.
The Pebblous blog has covered the same structure once before, from the opposite end: a study that computed the highest attainable score on a visual benchmark without running a model even once. The value computed there pinned down where the model's share begins. The defect rate in this article works the other way around. It hides where the model's share ends. One hands responsibility back to the model, and the other blends the model's responsibility into the test paper's.
4.2Evidence that this is not a physics problem only
That answer-key errors can overturn model rankings has already been quantified in another field. A 2021 study of label errors found them in ten representative test sets across computer vision, natural language and audio. The label error rate in the ImageNet validation set was 5.83%. The damage showed up in the ranking, not in the rate. Re-ranking some sixty pretrained models on the corrected subset sent the model that had been first down to 29th, and lifted a much smaller model from 34th to first. Higher-capacity models had memorized the original's wrong labels along with everything else.
Setting 5.83% beside this article's 57.20%, the 143 benchmark errors among the 250, to conclude that physics is in far worse shape than image recognition would be a misreading. The first has the whole test set as its denominator; the second has the 250 responses filtered out as wrong. The two numbers do not sit on the same ruler. The two studies share a structure rather than a proportion. Wrong labels make scores wrong, and wrong scores make rankings wrong.
The precedent from coding is one the paper cites directly in §4. The Pebblous blog has separately covered how that benchmark was retired over defects and saturation, so the figures stay there. Two sentences the paper adds are worth carrying over. "Defects survived two rounds of curation in a domain where every task ships with an executable test. Physics benchmarks are graded against written reference solutions and have no comparable check, so there is no reason to expect them to be cleaner, and no way to find their defects short of re-deriving each answer."
One more distinction belongs here. The trust problem in agent benchmarks, which the Pebblous blog covered earlier, was a story about parties digging into the gaps in the rules to score well. This case is a different animal. Nobody here deceived anyone. The people who wrote the items, the people who coded the grader and the people running the leaderboard were each diligent, and the scores came out wrong all the same. Gaming goes away once the incentive is removed. A mistake stays until somebody builds a verification procedure.
Who audits the audit?
Everything in this section comes from Appendix F of the paper. That fact goes down first. The weak links in the audit design were not dug up from outside; the authors wrote them down themselves, and that is why we get to read them. This section is a reading that follows the record rather than a criticism. The question of who reviewed the answers comes back around to the paper that did the reviewing.
5.1What the auditors had in front of them
The people who looked at the 250 again are physics doctoral students. Four things were on their screens: "the problem statement, the reference solution, the model response, and an AI-generated preliminary review." Reviewers chose an area of expertise and were assigned items they had not reviewed before, and the interface "requested 30 reviews per contributor and allowed up to 15 skips for questions outside their expertise."
One precedence rule hangs over the classification. When the problem statement or the reference solution is defective, benchmark error takes precedence over the other verdicts. Part of the reason 143 is the largest of the three categories sits right there. The grader may well have been wrong on the same item, but once the test paper is wrong first, that item is written into the benchmark-error column and nowhere else.
The two benchmarks audited in full had a tighter procedure and fewer people. "Each problem was assigned to a single reviewer," and for CritPt that one person was an expert in the relevant subfield, either faculty or a researcher whose advisor vouched for their participation. Reviewers had to set out their own approach before reading the reference solution. The rule at the final stage reads: "AI tools could be used for assistance, but reviewers were responsible for independently verifying the submitted results."
5.2Of the items two reviewers saw, 28.57% split
The verdicts attached to the 250 come to 446 in all. 196 items were seen by two reviewers and 54 by one, and the paper gives the reason for the latter as "limited human resources." Of the 196 seen twice, 140 pairs of labels matched and 56 differed. All 56 disagreements were settled by a third review.
The proportion that split matters less than where it split. Counted by type, the disagreements before resolution break down like this.
| Conflicting label pair | Count | What rides on it |
|---|---|---|
| Benchmark error ↔ grader error | 34 | Both sit on the test paper's side, so the model-error count does not move |
| Benchmark error ↔ model error | 19 | The boundary around the headline figure of 12 model errors runs right here |
| Grader error ↔ model error | 3 | The other side of the same boundary |
| Total | 56 | 28.57% of the 196 items seen by two reviewers |
Source: arXiv:2609.13009, Appendix F, Tables 3 and 4. The counts in Table 4 are the values before resolution by a third review. These two tables rest on a transcription of the paper's full text, and the internal sums were verified (196 + 54 = 250, 140 + 56 = 196, 34 + 19 + 3 = 56).
Original Pebblous diagram. Source: arXiv:2609.13009, Appendix F, Tables 3-4. The 22 is the sum of 19 (benchmark error ↔ model error) and 3 (grader error ↔ model error).
Of the 56 split verdicts, 22 sit on the boundary between the test paper's fault and the model's fault. That boundary is exactly where the figure of twelve, which this article has carried since its first paragraph, gets made. So twelve reads more accurately as a value that emerged after a three-way judgment on which people differed to some degree and which was then resolved, rather than as a settled figure with no rounding in it.
5.3Is that much disagreement common?
The 140 matching verdicts are 71.43% of the 196 items seen twice. How to read that agreement rate cannot be settled from this paper alone. On a three-way task, guessing at random lands about one third of the time, so 71.43% is plainly above chance. The band that the annotation literature calls adequate, on the other hand, usually starts above 80% raw agreement. So the figure lands somewhere in between. A chance-corrected statistic goes unreported here, and with the three categories distributed as lopsidedly as 143 to 95 to 12, correcting for chance could bring the number out lower than the impression it leaves.
This blog holds one record taken through the same lens: a case where three labelers' verdicts diverged. Its agreement rate cannot be lined up against this paper's figure, however. That task had three people judging across thirteen categories; this one had two people judging across three. The more categories there are, the lower the chance of agreeing by accident, so the two figures stand on different baselines from the start. The two cases share an attitude rather than a number. Both kept the disagreement on the record instead of erasing it as noise.
5.4On having an AI-written preliminary review alongside
What happens in annotation work when a model's suggestion is shown to a person first has been measured several times over the past few years, and the results do not converge. In one large pre-registered experiment the label distribution shifted significantly toward the model's suggestion. In other studies, agreement came out about the same whether or not the model's hint was supplied. The strong form of the claim, that a first impression governs the whole of the subsequent judgment, has been rejected on testing in at least one case.
There is thus no ground to say the design of this paper is wrong. Something else can be said. This audit used a design whose bias risk is known, and the size of that risk cannot be measured from inside the paper. Of the 250 items, 54 were seen by one person, and of those seen by two, 28.57% split. The layer that reviews the answers has quality problems of its own, and nobody has audited that layer yet.
Which version is the score on the leaderboard?
The paper's title carries the phrase "near-saturation." How far that phrase reaches is something the paper narrows down in the same section, and stripping the caveat off it sends the whole piece rolling the wrong way. So the range goes first.
6.1Saturation is a story inside the problem set
The saturation claim is about closed-form problems, the problem-set items whose answers come out to one thing. In the same section the paper writes that while models may have reached a tipping point in physics, "it by no means follows that frontier models are capable of end-to-end physics research." The grounds are an experiment the authors ran with their own hands. They took agentic harnesses of the kind that had succeeded on open mathematics conjectures and set them on several open problems in theoretical physics, and the result is short. "As of now, we have not been able to fully solve a single one of these open problems this way."
Reading that zero as "models added nothing to physics research" would clash with another list the paper sets down on the same pages. §4 gives three pieces of physics work carried out with frontier models outside the standard benchmarks. One report describes a model reconstructing a hidden symmetry algebra after a simpler warm-up problem; one researcher describes steering a model through an extended theoretical-physics project that produced a new factorization theorem and a resummed C-parameter calculation; another team paired a model with tree search and numerical feedback to derive analytic results for cosmic-string radiation. The caveat the paper attaches to that list right afterward is the part that matters. "None of these is a score on held-out questions. Each puts a human expert or an external checker in the loop, over hours or weeks, on a calculation with no reference solution to grade against." The zero is the figure for the setting with the person taken out, and the record from the setting with the person left in sits elsewhere.
Saturation on closed problems and a zero on open ones are not a contradiction. A problem set gathers problems whose answers already exist, and research begins where the answer does not yet. This paper says the difficulty of the former has run out, not that the latter has been solved. Which is why its conclusion closes on a call for a new style of benchmark, "built from new hard physics tasks and adequately verified," rather than on any declaration that the job is done.
6.2Two scores under one name
The scene a practitioner runs into first sits outside the paper. It is on the scoreboard as it stands. Of the six benchmarks audited, CritPt is the only physics benchmark inside a major aggregate index, and as of 15 September 2026 the Artificial Analysis CritPt leaderboard still shows GPT-5.6-Sol at 32.3%. That is effectively the same value as the mean@5 of 32.29% the paper cites as the pre-audit figure. The physics score a reader sees on that page today is the value from before the answer keys were repaired.
The people running the leaderboard are not the ones at fault. Artificial Analysis did not build the benchmarks; it takes evaluation sets built by others and runs them, and it puts real effort into consistency of execution. On how it verifies the correctness of reference answers or the reliability of graders, though, no public methodology exists. The interesting part is that the same outfit has already repaired an answer key on a different evaluation set. The notes accompanying a revised edition of its own long-reasoning evaluation set contain a sentence about having corrected 16 answer keys. In practice they already know that answer keys can be wrong, and that procedure has not been written up as a public document.
On the other side a different version circulates. Anthropic's system card lists an item called CritPt-Corrected, which is an internal corrected edition rather than the public one. As the paper reports it, expert revisions were obtained to 31 of the 71 problem statements; each problem allowed sixteen max-effort, tool-enabled attempts; the judging was done by the vendor's own model; and the value that came out under those conditions is 88.4%. The metric name written there is average pass@1, which the paper separately notes corresponds to mean@16 in its own notation. The system card carries the corrected score alone, and no pre-correction score is placed next to it. This account, however, travels by way of the paper's citation and a secondary summary. We did not open the system card file itself.
Original Pebblous diagram. Source: arXiv:2609.13009, Appendix B.2.3 (via citation, system card not opened directly), Artificial Analysis CritPt leaderboard (checked 15 September 2026).
Two scores circulate under nearly the same name. The public leaderboard shows the pre-correction figure only, and the vendor's system card shows the post-correction figure only. Neither writes the other one alongside. Subtracting one from the other, or reading them as a multiple, does not work, because the model, the attempt budget, the item set and the judge all differ. The paper itself states that the differing "corrected question sets, models, attempt budgets, and judges" are what "make the two evaluations not directly comparable." The fact left standing here is not the size of either score but a single asymmetry in disclosure.
This is the fourth day since the paper went public. Whether the parties maintaining the audited benchmarks have replied publicly or issued corrected editions is something we could not confirm as of this writing. That means we could not confirm it, not that there has been no response. The history of each benchmark's repository was not investigated separately.
6.3What to ask when a score table arrives
An organization choosing a model or weighing adoption usually has a score table in hand. This paper adds one axis to that table. What is the defect rate of this score? The same axis carries over intact to an in-house evaluation set. Boiled down to five lines, the questions run like this.
- Who reviewed the answer key behind this score, and how many people spent how much time on that review?
- Has anyone looked again at the responses recorded as wrong? If so, were the results sorted into model error, grader error and benchmark error?
- Was that proportion written next to the score? An evaluation report with no defect rate is a number read without knowing its order of magnitude.
- Which version produced the figure in front of us? The original, or a corrected edition, and where is that recorded?
- Does the documentation for the evaluation set state what its makers did not verify?
The last item is the cheapest prescription in this affair and the most effective. HLE's five-minute rule is the evidence. Because the makers wrote down what they had not checked, the audit could tell in advance where to look and what to look for. Writing down what went unverified takes no extra budget.
The side that does take budget is the next benchmark. The paper writes that a new evaluation "will likely require substantial financial resources to achieve high-quality curation," without saying how much. Outside the paper the situation is much the same. One exam put up a $500,000 prize pool, one mathematics benchmark published rewards in the range of several hundred to a thousand dollars per item, and one benchmark audited here records an average of more than 40 hours of labor per item. Yet not a single case has disclosed its total production cost. In Korea the gap has the same shape. Training-data quality sits inside a certification scheme, while a provision governing the quality of answer keys and graders for evaluation data cannot be found.
The paper puts an actor rather than an amount in that slot. It points to nonprofit organizations taking the work on through a more elaborate process that ensures rigor, or to open competitions in the style recently adopted in mathematics. The part worth noticing is that the paper does not write this up as a problem of evaluation design. It writes it up as the question of how to set up data collection that is "fair, fast, and of high quality." The question hanging over the next benchmark comes back, in the end, to how the data is gathered and reviewed.
To close, one neighboring piece marks where this article stands. A study that audited the redundancy of an evaluation suite statistically took up the question of what gets measured how many times, which is the weighting inside evaluation design. This article took up the layer beneath it. Is the item being measured once written correctly to begin with? Skewed weights shake the ranking, and a wrong item makes the absolute level itself wrong.
Why this matters to Pebblous
Pebblous diagnoses data and issues quality reports on it. So this paper does not read as news from some other industry. The subject of the audit happened to be evaluation data rather than training data, and the shape of the breakage is the shape we look at every day.
7.1The breakage has the same shape
Translate the defects this paper found into the language of data quality and three familiar things come out. The labels are wrong, the labeling convention is missing, and the machine that reads the labels cannot recognize equivalence. An arithmetic error in a reference solution is the first. An item that split its answer by a factor of four because it never said which operator the coupling constant multiplied is the second. A rule-based grader that handed out a zero over one rationalization and one reordering of factors is the third. Our work on training data deals with exactly those three: locating the defective entries, stating the conventions, and defining the equivalence test.
The difference lies in institutions. On the training-data side, procedures for those three exist with names and forms attached. On the evaluation-data side, the same procedures have never been institutionalized, and in that state they have been propping up the industry's standard scoreboards. That 143 of the 250 were benchmark errors means the subject of diagnosis has widened by one layer.
7.2Label quality sets the floor of the measurement
In the language of data quality, that one line about the defect rate reads like this. The label error rate sets the floor of the measurement, and the floor looms larger the better the thing being measured gets. While the subject is poor, defects in the instrument get buried in the result; once the subject improves, the instrument's defects are all that is left. Few sentences explain more precisely why the practice of comparing models without measuring the label error rate is coming apart now.
7.3How to carry this into an in-house evaluation set
The method in this paper runs just the same at a smaller scale. There are four steps. Sample the responses recorded as wrong and have people look at them again. Sort the results into model error, grader error and benchmark error. Report that proportion alongside the score. And state, in the documentation for the evaluation set, what was left unverified.
The third step is the one most often dropped in practice. A report carrying a score and no defect rate is a graph with no error bars. The fourth step is what this affair newly taught. Write down what went unverified, and later somebody opens that spot first.
7.4The uncomfortable fact this article leaves
This article does not close on the sentence that evaluation data is, in the end, data too. That sentence has been written many times over, and the work now lies after it. This paper leaves us with a more uncomfortable fact, exactly as section 5 had it. The judgments that picked out the defects varied from person to person, part of those judgments passed through a single pair of eyes, and the judges had a review written in advance by a model sitting next to them. Which is to say the layer that reviews the answers has quality problems of its own, and that layer is not yet the subject of any audit.
Pebblous has nothing to sell at this spot. We can, though, draw the shape of the gap precisely. There is a layer that makes the scores, a layer that reviews those scores, and a layer that reviews the reviewing; all three carry data quality problems, and the further down you go, the fewer people are measuring anything. The reports we issue cannot dodge the same question. What ruler produced this score, what is the defect rate of that ruler, and can the customer see that number?
The figures and verbatim quotations in the body were checked directly against the full text and appendices of the public arXiv version. The self-descriptions of the six audited benchmarks were confirmed in each original paper, and the current leaderboard value was checked on the page itself. One Anthropic system card was the only source whose original file we could not open, so we went by way of the paper's citation. Please read the paper's results and Pebblous's interpretation as separate things. Sections 1 through 6 carry what the researchers measured and what we were able to confirm, and section 7 is work this paper did not do. Thank you for reading this far.
References
The figures in the body come from three streams. The table and appendix values from paper 1 were carried over after direct comparison with the arXiv full text. Items 2 through 7 are the original papers of the audited benchmarks, and the self-description quotes in the comparison table in 2.2 were confirmed directly in each paper's abstract or methods section. Item 9, a system card, was reached by way of paper 1's citation and a secondary summary because the original file could not be opened, and the body says as much.
The spine of this report (checked against the primary text)
- 1.Ali Ansari, Haoran Sun, Andy Zeyi Liu, Mark Jabbour and 47 others. "How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks." arXiv:2609.13009v1, submitted 11 September 2026, CC BY 4.0. arXiv: 2609.13009 — a preprint that has not been through peer review. Of the 51 authors, 39 are listed as data auditors, and five of the 12 physics advisors also appear among those 39 (confirmed by cross-checking the two lists in Appendix G). The corresponding author is John Sous of Yale, and the work was funded by the Yale Office of the Provost AI Initiatives and a gift from Jump Trading Group. Every citation of Tables 1 and 2 and Appendices B, C, D, E and F in this article comes from this version.
Original papers of the audited benchmarks (self-description check)
- 2.Center for AI Safety, Scale AI. "Humanity's Last Exam." arXiv:2501.14249. arXiv: 2501.14249 — the "unambiguous and easily verifiable" quote in 2.2 comes from the abstract; the five-minute rule and "we welcome community feedback" come from the review section, §3.2.
- 3.Zhu et al. "Probing the Critical Point (CritPt) of AI Reasoning." arXiv:2509.26574. arXiv: 2509.26574 — the figure of more than 40 hours of labor per item on average is also this paper's own account.
- 4.Pan et al. "CMT-Benchmark." arXiv:2510.05228. arXiv: 2510.05228
- 5.Qiu et al. "PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models." arXiv:2504.16074. arXiv: 2504.16074 — the source for expression edit distance (EED) and for the claim of 204% sample efficiency over binary grading.
- 6.Zhao et al. "PRISM-Physics: Causal DAG-Based Process Evaluation for Physics Reasoning." arXiv:2510.03185. arXiv: 2510.03185 — the similarly named arXiv:2512.05930 "PRiSM" is a different paper.
- 7.Xu et al. "UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models." arXiv:2502.00334. arXiv: 2502.00334
Adjacent audits and the label-quality literature
- 8.Curtis G. Northcutt, Anish Athalye, Jonas Mueller. "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks." NeurIPS 2021 Datasets and Benchmarks, arXiv:2103.14749. arXiv: 2103.14749 — the source for the 5.83% label error rate in the ImageNet validation set and for the re-ranking on the corrected subset in which the top model fell to 29th and the 34th-place model rose to first. Its denominator is the whole test set, and section 4.2 of the body records that difference.
- 9.Anthropic. "Claude Fable 5.1 & Claude Mythos 5.1 System Card", 1 September 2026, §8.9 CritPt-Corrected — we could not open the original file. The account in section 6.2 travels by way of the Appendix B.2.3 citation in paper 1 and published secondary summaries.
- 10.Artificial Analysis. "Intelligence Index methodology". artificialanalysis.ai — the current CritPt leaderboard value (GPT-5.6-Sol Max, 32.3%) was checked on 15 September 2026. No methodology for verifying the correctness of reference answers or the reliability of graders was found in this document.
- 11.Bowman, Dahl (2021). "What Will it Take to Fix Benchmarking in Natural Language Understanding?" — the statistical-power argument in section 4.1 and the account of the four conditions for an adequate benchmark both travel by way of the summary in §4 of paper 1; the original was not separately checked.
- 12.Tan et al. "JudgeBench: A Benchmark for Evaluating LLM-Based Judges." arXiv:2410.12784. arXiv: 2410.12784 — a rare case of benchmarking a judge's own accuracy against objective ground truth.
- 13.Zheng et al. "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena." arXiv:2306.05685. arXiv: 2306.05685 — the basis for the statement in 5.3 that the judge literature mostly reports judge-human agreement and not absolute error rates.
- 14.Stanford HAI. "AI Index Report 2026". hai.stanford.edu — the source for the background claim that benchmark lifespans are shortening to a matter of months.
Neighboring pieces on the Pebblous blog
- 15.The benchmark that underrated AI agents because of answer-key errors — the predecessor to this article.
- 16.The attainable ceiling of a visual benchmark, computed without running a model — the concept on the far side of the floor a defect rate makes.
- 17.The coding benchmark retired over defects and saturation — referenced in one line in section 4.2.
- 18.The trust problem in agent benchmarks — referenced to separate gaming from mistakes.
- 19.A case where labelers' verdicts diverged — referenced in 5.3 along with the reason direct comparison is impossible.
- 20.A study auditing the redundancy of an evaluation suite statistically — section 6.3 records how it differs from this article.