Executive Summary
As LLMs increasingly grade the answers of other LLMs, audits that check whether the judge is biased have settled on a standard design. The one that has passed as the strictest holds the item fixed, changes a single condition, and differences twice. Compare two candidate responses inside the same item, then compare that comparison across conditions, and whatever loads equally onto both responses cancels. A paper posted to arXiv on August 27 shows that the cancellation argument does not hold on a scale that stops at 1 and 5, and it shows it inside the authors' own audit data.
The audited system was a judge that scores the next tutor turn in a tutoring dialogue. The pre-registered primary endpoint barely moved at +0.085 points, and only one of four rubric sub-scores reached significance: +0.378 on productive struggle. The authors then built a construction that forces both responses to lose exactly the same amount, so differential preference is zero by hand, and that construction reproduced 79 to 85% of the observed value.
The reconstruction does not prove that no bias exists. The authors are explicit that it is a counterexample, not a decomposition. Still, anyone receiving an audit report now has a question worth asking. Can the reported interaction be traced to responses parked at the end of the scale, using nothing but the ratings the audit already holds?
Key figures
Source: Fan et al., arXiv:2608.27309 (2026-08-27)
79–85%
Reproduced with zero differential preference
The share of the one significant +0.378 that clipping alone rebuilds. Rounding to the integers the judge actually emits brings it to 79%
+0.085
Pre-registered primary endpoint
95% BCa interval from −0.167 to +0.353, p=0.684. Any mark the profile left on the judge's scaffolding preference went undetected at this resolution
17 / 30
Stimuli pinned at the floor
On productive struggle the low-scaffolding response sat at exactly 1.000 in all three arms. On those items the double difference equals one response's shift
6.62 vs 1.38
Headroom the scale left behind
Averaged over the 30 weak-stratum stimuli the two responses sit at 4.58 and 1.96. Room to move positive runs close to five times the room to move negative
Differencing twice was trusted as the strictest design
Scores from LLM judges already move leaderboards and benchmark rankings, and in education they have reached the point of grading the instructional quality of tutoring systems against a rubric. Whether a judge should hold that role is itself a measurement question, and the audits that answer it mostly share one design. Fix the item, vary a single attribute of how it is presented or of the identity attached to it, and read the bias off the contrast between the two matched conditions.
The stronger design subtracts once more. Take the comparison between two candidate responses inside the same item, and subtract that comparison across the manipulated attribute. Anything that pushes one response's score up or down by the same amount in every condition disappears along the way. What survives is a difference in differences, and in judge audits it is read off a rating scale that runs from 1 to 5.
The cancellation argument itself is not wrong. Additive artifacts really do cancel. The trouble is the quieter assumption that followed it. Once an effect survived the double difference, the bounded scale was treated as an annoying but conservative nuisance: clipping at the scale ends can attenuate a real effect, the reasoning went, but it cannot create one. For a single contrast that is true, because clipping is non-expansive.
These authors wrote that assumption into their own pre-registration for the double difference, and there it is false. Econometrics has long separated a latent interaction from the cross difference of observed means. Tobin's censored regression is where that line starts, and along it sits Puhani's result that the observed cross difference in a nonlinear difference-in-differences is not the treatment effect. On ordinal outcomes there is the further objection that the estimand is not even defined without latent-scale assumptions. LLM judge audits nonetheless kept reporting interaction effects from short bounded rubrics with no identifiability check. The example the authors name in their introduction is a factorial interaction study read off a 1 to 10 rubric with most ratings within two points of the floor.
A floor on the scale turns a common fall into an interaction
The mechanism itself is simple. Each response carries an unobserved latent quality, and the score the judge emits is that value clipped into the interval between 1 and 5. Now suppose the manipulated condition takes the same amount away from both responses. Differential preference is exactly zero. At the latent level the difference in differences is zero too.
The observed value is another matter. Because of the clipping, each response shows only part of that common fall. How much gets swallowed depends on how close that response sits to a bound, and two responses at different distances swallow different shares. What the observed difference in differences ends up measuring is how differently the two responses hid the same fall. That is a difference in attenuation, not in preference.
The extreme case is the easy one to see. If a response is flush against the end of the scale and cannot move under any condition, its shift is zero. The difference in differences then equals the shift of the remaining response exactly, and a one-response contrast gets reported under the name of an interaction between two conditions. The statistic is zero only when both responses clip the common fall identically.
An asymmetry of direction rides on top of this. When the two responses are spread out on the scale, the headroom available to the endpoint in the positive direction exceeds the headroom in the negative direction by twice the separation between them. Averaged over the 30 weak-stratum stimuli in this audit, the two responses sat at 4.58 and 1.96 under the novice profile, leaving 6.62 scale points of headroom upward against 1.38 downward. An audit hunting for bias is structurally far more likely to meet a positive number.
No amount of design polish escapes this, and better stimuli make it worse. A good contrast sharpens the quality gap between the two candidate responses, which is another way of saying it pushes them toward opposite ends of the scale. The very property that lets a manipulation check pass is what pins the responses to the bounds.
Across 990 ratings the primary endpoint did not move
What the authors audited is a judge released by earlier work. It takes a tutoring dialogue and scores the pedagogical quality of the next tutor turn, and its system prompt went unchanged through all 990 calls. The material is 55 stimuli drawn from tutoring dialogues over six acid-mixture word problems. Each stimulus is a dialogue that breaks off at a student turn, paired with two candidate replies. One is the high-scaffolding response, which leaves the next reasoning step to the student from wherever the student stands; the other is the low-scaffolding response, which takes that step for them.
The manipulation is a single block of learner profile text. There are three arms: dialogue and responses with no profile, a stated novice profile prepended, and a stated advanced profile. The two profile texts are sentence-by-sentence parallel at 59 words each, and neither hints at the dialogue nor instructs the judge how to weigh it. Fifty-five stimuli times three arms times two responses times three repeated ratings gives 990 calls. The judge returns one integer between 1 and 5 for each of four pedagogical principles, plus an overall rating. The identifier the provider returned on all 990 calls was Claude Opus 4.8.
The weak and strong strata split the stimuli by the competence the student actually demonstrated in the dialogue. Three independent blind passes labeled the dialogue contexts alone, never the candidate responses or the tutor metadata, and the majority vote decided, with pairwise agreement above 93%. That produced 30 weak-stratum stimuli and 25 strong. The unit of inference is not the stimulus. Stimuli drawn from the same tutoring session are not independent, so tests and intervals run on means taken first within each source run, of which the weak stratum has 18 and the strong 10. The clusters referred to throughout this article are those source runs.
Start with the result. The judge preferred high-scaffolding responses by 2.58 to 3.19 points under every arm and in both strata. The baseline itself is enormous. Against it, the amount the learner profile moved that preference, which is the pre-registered primary endpoint, came to +0.085 points on the weak stratum. The 95% BCa interval runs from −0.167 to +0.353, and the exact signed-rank p-value is 0.684. The strong stratum points the same way at +0.106 with p=0.914.
The authors state plainly that this is not to be read as proof of absence. No equivalence margin, no smallest effect size of interest, and no power statement were pre-registered, so no claim of the form "the bias is below such and such" is available. The upper end of the interval is +0.353 on its own, and in simulation this test crosses 0.80 power somewhere between +0.40 and +0.42. This is a failure to detect at a stated resolution, not an absence.
Nor can the null alone be read as the judge grounding its scores in the dialogue evidence. The frozen system prompt instructs the judge to rate on the dialogue alone, and the profile block sits outside what that instruction points at. The reading that the judge grounded its scores in the evidence and the reading that it simply followed the instruction predict the same null. The pre-registration therefore built in a device to separate the two readings.
For the null to mean anything, the profile had to reach the rating at all. It did. On absolute scores with the response held fixed, the advanced profile scored the same low-scaffolding response 0.153 points lower than the novice profile did (p=0.00781) and the same high-scaffolding response 0.238 points lower (p=0.0625). That both numbers are negative is the important part. The profile did not favor one response; it pulled both down together. That is the common severity shift analyzed in the previous section, in exactly that configuration.
The movement is concentrated on the advanced side. Measured against the no-profile arm, the novice profile touched the high-scaffolding response by +0.009 and the low-scaffolding response by −0.128. The advanced profile pulled the same two down by −0.228 and −0.281. One caveat attaches to the significant term: its p-value of 0.00781 is the floor attainable with eight nonzero clusters, so it says only that every cluster that moved moved the same way.
The authors also head off the next move: reading that fall as direct evidence that a label swayed the judge. The two profile texts differ in more than ability: they state a help-seeking preference as well, and help-seeking is the construct one of the four sub-scores measures directly. A judge that scores the same response lower for a student introduced as advanced may be anchoring on the label, or may be measuring what it was told to measure, correctly calibrated. The registered evidence-gradient check that would have separated those two readings, testing whether the profile effect shrinks where the dialogue shows clearer competence evidence, came back null. Spearman ρ was −0.179 with p=0.415. Neither reading is supported.
A construction with zero preference rebuilt the one significant +0.378
Of the four pre-registered per-field gaps, one was nominally significant: +0.378 on productive struggle. The other three did not reach the threshold. Break each field down by response, though, and what the ranking is measuring becomes visible.
| Rubric sub-score | Difference in differences | p | Shift of the high-scaffolding response |
|---|---|---|---|
| Productive struggle | +0.378 | 0.002 | −0.395 |
| Assistance calibration | +0.309 | 0.102 | −0.465 |
| Elicitation | +0.134 | 0.227 | Not reported |
| Contingent scaffolding | +0.130 | 0.502 | Not reported |
Assistance calibration moved its high-scaffolding response further, by −0.465, and yet its gap is smaller. What made the difference is how much room the low-scaffolding response had left to move. The field ranking is measuring the geometry of clipping rather than the judge's preferences. Across all four fields the high-scaffolding response accounts for 72 to 96% of the total movement of both responses and 104 to 164% of the gap itself.
The ranking is not explained by the four fields measuring different things either. The sub-scores the judge emits are highly correlated, and two of them agreed in 815 of the 990 ratings. Fields that are close in content diverge far more in how exposed they are to the scale bounds. That is the direct reason not to line up per-field gaps and read the largest as the strongest effect.
Nor does the response have to be flush against the floor for this to happen. On contingent scaffolding, not one stimulus has its low-scaffolding response at 1.000 in all three arms, and the same one-sided pattern appears anyway, with the high-scaffolding response carrying nearly all of the movement. Two responses sitting at different distances from a bound is enough. Being pinned is just the extreme where that distance is largest.
On productive struggle the structure is at its most blatant. On 17 of the 30 weak-stratum stimuli, the low-scaffolding response scored exactly 1.000 on this field in all three arms. That is the floor. On those stimuli the difference in differences becomes algebraically identical to the shift of the high-scaffolding response alone, an identity the data satisfy exactly. Splitting on that condition separates the two: +0.472 (p=0.00391) where the response is floored, +0.149 where it is free. The second figure rests on only five nonzero clusters, so no test could have rejected there.
The reconstruction procedure is simple. For each stimulus, transport the observed shift of the high-scaffolding response onto the low-scaffolding response, add it to the three integer ratings from the novice arm, clip into the interval between 1 and 5, and average again. Differential preference is pinned to zero by hand, so nothing but attenuation is left between the two responses. That construction reproduces +0.321, which is 85% of the observed +0.378. Rounding to the integers the judge actually emits brings it down to 79%. A full clip of the low-scaffolding response would have manufactured 104.5% of the observed gap.
It is tempting to book the remaining +0.057 as the share belonging to real bias, and the authors close that road. The residual is the construction's prediction error on the low-scaffolding response, computed on five nonzero clusters, where even the smallest attainable p-value exceeds the significance threshold. It bounds nothing. The construction also assumes the very common fall that is in question, which is why the authors call it a counterexample, not a decomposition. That the productive-struggle result does not support an anchoring interpretation is the whole of what this data can say.
Censoring did not run in only one direction
The pre-registration contained a sentence saying that censoring only attenuates, so the observed value is a lower bound. That holds if a bound hides only latent movement past it. It does not. A latent value above the cap produces an observed fall only once the latent fall exceeds the headroom, so a ceiling masks falls as well as rises.
The data show it. Among weak stimuli whose high-scaffolding response was pinned at 5.000 under the novice profile, 1 of 16 fell under the advanced profile. Among the unpinned ones, 8 of 14 did. Fisher's exact test gives p=0.0043. Which way de-censoring would move the endpoint is therefore not even self-evident. In this audit the high-scaffolding response moved more, which means the swallowed share was larger on the low-scaffolding side.
The primary endpoint is itself a mixture rather than a single value: −0.205 on the 18 stimuli whose no-profile high-scaffolding response already sat at 5.000, and +0.352 on the 12 where it did not. Neither is significant.
Resolution has to be read alongside this. A per-unit score is the mean of three integers, and 284 of the 330 units returned the same integer in all three repetitions. The estimand lives on a sparse lattice. Five optimally chosen changes of a single integer rating carry the primary from +0.085 past zero, and six land it exactly there. An interior zero here is a threshold that was not crossed, not an observed indifference.
Multiplicity is also left unresolved. The pre-registered analyses report 21 hypothesis tests and 26 interval estimates, and no correction was registered or is claimed. The authors write that the sign test most directly supporting the productive-struggle result survives no correction at a family larger than four. The counterexample this article is about needs no correction to hold, but how much weight to put on the claim that the field was significant in the first place does depend on it.
None of this means a false assumption inside the pre-registration brings the whole audit down. The authors' accounting is narrow and clear. Because the primary endpoint is null, nothing in the headline result is corrupted by the error. What changes is what a non-null primary would have entitled them to claim, and that is precisely what the pre-registration was meant to fix in advance. The registered protection covered only the weak reading that the profile influenced the rating somehow, and it did not cover the per-field secondaries at all. Binding an analysis in advance is not validating it.
Better stimuli make the endpoint harder to identify
What carries from this paper to other audits is the mechanism, not the null. Matched-condition contrasts have been the standard defense against artifacts in LLM judge audits, and differencing once more inside the same item was the natural way to strengthen that defense. On a bounded scale, though, the design turns into a one-response severity contrast as the stimuli get better, because the separation between responses that makes a manipulation check pass is the same separation that pins them to the bounds. As the systems under evaluation improve and short rubrics pile up at their top, more of the interaction effects that get reported will be attenuation differences under the wrong name.
6.1Three things you can check inside the audit's own data
The saving grace is that this is diagnosable from ratings an audit already holds. The authors set out a three-step procedure.
- • Count how many responses sit at a scale bound in every arm. Their shift is zero by definition, and on those items the difference in differences is the other response's shift.
- • Split the sample on that condition and compute the endpoint separately. This shows how much of the reported effect sits where the endpoint has become a one-response contrast.
- • Transport the free response's observed shift onto the pinned one and clip. Whatever this zero-preference construction reproduces is a magnitude that censoring alone can generate.
The authors draw a line between diagnosis and correction, and they are offering the first. They publish no de-censored point estimate, and note only that recovering the latent endpoint would take a response format that avoids the relevant bounds, or a latent-variable model that handles censoring and ordinal thresholds explicitly. Whatever such a model returned would then lean on its own assumptions.
The paper's own scope is narrow. A single judge model, one rubric, six acid-mixture word problems, and 55 stimuli clustered in 23 source tutoring runs. The two candidate responses differ in more than scaffolding. Question marks appear in 37 of the 55 high-scaffolding responses against 1 on the low-scaffolding side, boxed final answers appear 18 times on the low side and never on the high, and 26 of the 29 corpus-drawn low-scaffolding responses come from a single tutor model.
The judge's identity is not fully fixed either. The provider returned one alias and no dated snapshot appears in any response or in the released logs, so a later call under the same name carries no guarantee of reaching the same model. What the authors hold is one agreement between two points in time. Re-scoring each stimulus's original tutor turn with no profile gave a mean of 3.891 against 3.879 in the published corpus, with Pearson r of 0.979 and exact agreement on 40 of the 55 stimuli. That comparison covers 165 of the 990 evaluations.
Editor's Note: There is an illusion here that Pebblous keeps running into on data quality work. Evaluation data gets treated as the ruler that measures the model, and what the markings on that ruler were built to measure rarely gets examined. What this paper shows is that the spacing of the markings and the ends of the scale help produce the verdict. No label is wrong and no sample is skewed, yet the conclusion wobbles. Any organization that has managed evaluation data quality through label accuracy alone might check whether scale design and response distribution are on that list too. Anyone receiving an audit from a vendor has something to ask for: show that the reported interaction can be rebuilt from the audit's own ratings. A number that cannot be rebuilt is not yet a verdict.
The paper is available at arXiv:2608.27309. The audited judge and the tutoring corpus were released by the same group at arXiv:2607.28128.