Executive Summary
A research team anonymized 120 papers submitted to ICLR 2026, held the methods, the experimental settings and the reported numbers fixed, and rewrote only the prose to produce 4,080 manuscripts. Several LLM judges then read those manuscripts under two scoring protocols. The scores moved. But the finding is not that AI judges fall for flashy writing. The sensitivity turned out to be structured.
How the evidence was framed and how novelty was claimed shifted the scores most, while making the vocabulary and syntax harder drew almost no reaction at all. Change the judge and the same edit could push a score down instead of up. Tightening the rubric pulled the whole scale down by 1.36 points, yet sensitivity to the prose did not shrink with it, and in some combinations it grew. The most uncomfortable part is where the change was recorded. Only the sentences had changed, and the model wrote it up not as better presentation but as a difference in contribution and soundness.
Detection cannot close this gap. Hidden prompts and artificial perturbations can be caught, but a rewrite that states the evidence more clearly and puts the contribution up front is indistinguishable from what an advisor asks a student to do. And every norm that conferences and publishers have put in place over the past year, within what we were able to verify, concerns confidentiality, human accountability and disclosure of use, in other words who wrote it. We found no clause asking what the judgment responded to. For any organization handing evaluation to a model, one question remains. When the same content arrives in different words, does your pipeline return the same verdict, and is there a record of anyone checking?
Key numbers
Source: measured values from Li et al. (2026), main text and appendices. All four are taken up again in the body.
Score contrast of the axis that moved most
(evidence framing) and least (lexical complexity)
Swing in weak-accept probability
from evidence framing alone
Drop in mean overall score under the strict rubric.
Sensitivity did not drop with it
Evaluation cells left entirely unscored
on one paper (of 84 missing records)
One paper, 4,200 manuscripts
"How Can Rhetoric Reward-Hack AI Reviewers?", posted to arXiv on 10 August 2026 by researchers at the University of Maryland, Virginia Tech, MBZUAI and the University of Waterloo, does not observe AI reviewers in the wild. It intervenes on the manuscript side under control. The design is simple: hold the scientific content of a paper fixed, change only the way it is written, and measure across a grid how far that variation moves the score. In a study like this, what determines how much you can trust the result is not the conclusion but the scope of the control. What was locked and what was left open is very nearly the whole study.
This is not the first time the Pebblous blog has covered AI peer review, but the axis is different. When 21% of ICLR 2026 reviews were flagged as AI-written, the question was who wrote the review, and the remedy was detection. What this report covers is what that judge responds to. It also differs from bias differences between judge models. What wobbles here is the verdict of a single judge when the same content arrives written a different way.
1.1What was locked, what was left open
The team pulled ICLR 2026 submissions through the OpenReview API, dropped withdrawn and desk-rejected entries, and sampled 20 papers from each of six bands of mean human rating to reach 120. Each paper was matched by metadata to its public arXiv LaTeX source, and where several versions existed, the best one was chosen by similarity against the normalized text. An agentic anonymization harness then stripped author names, affiliations, acknowledgements, submission-status phrasing and identity-revealing links. Source diffs and human inspection confirmed that the edits stayed inside the identity region.
Control during the rewrite came in two layers. One layer was enforced programmatically: citation and cross-reference keys, labels, URLs, figure paths, math and code environments, and bibliography files were all locked. Violations were returned to the agent for repair, and manuscripts that could not be fixed or failed to compile were discarded. The other layer was deliberately open: wording, organization, emphasis, captions, tables, and the narration and interpretation of reported results. The requirement to preserve methods, experimental settings, reported values, comparisons and core findings lived at the prompt level only. Sentence-level semantic identity was never demanded.
So it is wrong to describe this experiment as "manuscripts with the numbers untouched and barely a word changed." The accurate description is that the scientific content was preserved while the prose was rewritten wholesale. That boundary matters because it changes how the results read. If captions and result interpretations could also be rewritten, then the score changes that follow are not "nothing but a formula stayed untouched and it still moved," but "the same finding told in different sentences moved the score."
The scale sits on top of that. Six dimensions applied in both a positive and a negative direction gave twelve conditions, and two rewriter models each produced their own version of every manuscript. Single-dimension rewrites alone account for 2,880 manuscripts; adding the simultaneous rewrite that applies all six axes at once, three-round iterative rewrites, and a condition that revises once more after reading a judge's review brings the total to 4,080.
| Component | Specification | Scale |
|---|---|---|
| Corpus | ICLR 2026 submissions, stratified across six rating bands | 120 papers |
| Intervention space | 6 rhetorical dimensions × 2 directions | 12 conditions |
| Rewriters | GPT-5.5 (Codex CLI), Opus 4.8 (Claude Code) | 2 models |
| Single-dimension rewrites | One dimension at a time | 2,880 manuscripts |
| Simultaneous, iterative, feedback rewrites | All six axes 240 / three rounds 480 / review feedback 480 | 1,200 manuscripts |
| Judges | Gemini 3.5 Flash-Lite, Qwen 3.5 Flash, GPT-5 mini, GPT-5.5, Claude Sonnet 5 | 5 models |
| Review protocols | Standard (conference rubric), strict (demand evidence, forbid rewarding presentation) | 2 protocols |
| Total manuscripts / valid reviews | 120 originals + 4,080 rewrites | 4,200 / 42,396 |
| Direct API cost | Rewriting $20,583.34 + reviewing $8,582.55 | $29,165.89 |
From Table 1 and Appendix B.3 of the paper. Judges were told neither which condition a manuscript had been rewritten under nor which model had rewritten it.
The diagram below traces how a single manuscript multiplies into a stack of scorecards. Counts grow from left to right, and the two rows underneath separate what was locked from what the agent was free to touch.
Original Pebblous diagram, reconstructed from §3.1 of the paper. The boundary drawn in the lower two rows is what sets the range of valid interpretations.
1.2The procedure is public; the raw data is unverified
The team released a project repository under the MIT license. It contains the rewriting and review prompts, pipeline code, the anonymization harness and container definitions, so the barrier to walking the procedure again is low. At 36KB, however, the repository plainly does not hold the raw artifacts: 4,200 manuscripts and 42,396 review records. We were unable to confirm the scope of any data release. And at nearly $30,000 in API cost, rerunning the whole thing is not a realistic check either. Reproducibility here holds at the level of procedure and remains unverified at the level of raw data.
Only three of six axes moved scores
The six dimensions were fixed in advance: novelty stance, scope framing, evidence framing, contribution salience, technical register, and linguistic complexity. Each axis was applied once in the positive direction and once in the negative, the mean change in the overall score was measured, and the gap between the two directions was reported as a contrast. A large contrast means the judge's score swung sharply depending on which way that axis was pushed.
Three of the six axes moved appreciably; the other three barely moved at all. Among those three, evidence framing and novelty stance led. At the other end, the contrast for making vocabulary and syntax harder is 0.011, which rounds to zero by the second decimal place.
| Rhetorical dimension | Positive direction | Negative direction | Contrast |
|---|---|---|---|
| Evidence framing | +.289 | -.303 | .592 |
| Novelty stance | +.187 | -.353 | .540 |
| Scope framing | +.136 | -.289 | .425 |
| Contribution salience | +.192 | +.049 | .143 |
| Technical register | +.142 | +.034 | .108 |
| Linguistic complexity | +.017 | +.006 | .011 |
| Mean across six axes | +.161 | -.142 | — |
Table 3 of the paper, averaged over rewriters, judges and protocols. Values are changes in the overall score on a 1–10 scale.
Placing the two-direction spreads side by side makes the split between the top three axes and the bottom three visible at a glance. The bars below draw each contrast on the same scale.
Original Pebblous diagram, redrawn from Table 3 of the paper. The top three axes are properties that require judgment; the bottom three are properties you can check on the surface.
The asymmetry between directions deserves attention too. For novelty stance and scope framing, the drop in the negative direction is far larger than the gain in the positive one. Hedge your claims and the score reliably falls; push them harder and it does not rise nearly as much. Evidence framing is the only axis that moved substantially both ways. None of this is noise: on all three of evidence, novelty and scope, more than 100 of the 120 papers came out in the intended order, with the positive version scoring above the negative one.
A metric closer to practice than the raw score is the acceptance threshold. The paper defines an overall score of 6 or higher as weak accept. Six is the rubric's "marginally above acceptance" grade, and it sits above the 5.39 average review rating of accepted ICLR 2026 papers. On that measure, evidence framing alone swung the weak-accept probability by 13.0 percentage points, novelty stance by 12.0 points and scope framing by 9.0. Under standard scoring novelty is slightly stronger; under strict scoring evidence takes the lead.
2.10.93 is not an average
The abstract states that positive evidence framing raises the overall score by up to 0.93 points and negative novelty stance lowers it by up to 0.73. The most common error when those figures get quoted is to repeat them as averages. The 0.93 and the 0.73 come from a single cell where the Opus 4.8 rewriter, the Gemini 3.5 Flash-Lite judge and the strict protocol coincide. Averaged across all combinations, the same effects are +0.289 and -0.353, exactly as the table above shows. The distance between those two places is close to the paper's actual point. This sensitivity does not summarize into one average.
2.2Not another verbosity bias
Surface-form sensitivity in LLM judges is not new. The order in which candidates are presented has been shown to flip pairwise preferences, and the finding that response length substantially shifts automatic win rates produced an entire length-controlled benchmark. So it is tempting to file this study under "longer and fancier wins." Yet in this paper the manipulations that correspond to length and difficulty barely registered: linguistic complexity has a contrast of 0.011 and technical register 0.108. What moved was not the style but the framing of the argument.
The closest predecessor lies elsewhere. A 2025 study showed that when hedging language is added while the answer's accuracy stays the same, the judge's evaluation goes down. This paper takes that one marker, expands it into six axes, applies them to full papers rather than abstracts, and pushes in both directions. It also differs from the many prior efforts framed as optimization, searching for the revision that lifts a score. Instead of keeping the winning candidates and reporting the gain, this work fixes the axes in advance, preserves every result including the ones that went down, and reads them across a grid of two rewriters, five judges and two protocols. What comes out is not "how easily can it be gamed" but which axis moves, at which score band, under which judge. Not an attack success rate, a map of sensitivity.
The first score set the direction
On average, evidence framing lifts the overall score by 0.289 points. Split that average by the score the original received and you get a slope rather than a single value. The lower a judge scored the original, the more the rewrite raised it; where the judge had scored the original highly, the rewrite pushed it down. Under standard scoring, pooling the two rewriters within each paper, papers whose original fell in the 1–3 band moved by +1.05 on average, while the 8–10 band moved by -0.19.
Applying all six axes at once makes that slope steeper: +1.42 at the low end and -0.88 at the high end. Put the four values on one axis and it becomes clear that this sensitivity changes not just in size but in direction depending on where the paper started.
Original Pebblous diagram, redrawn from §4.3, §5 and Appendix Table 25 of the paper. Middle bands are omitted because the paper does not tabulate them.
Run the same slope against human ratings and it disappears. Whether the human reviewers had rated a paper high or low, the change in the AI judge's score after rewriting showed no consistent direction. What creates the slope is not the paper's actual quality but the value the judge assigned first. The intuition that better papers are more robust does not hold here.
In practice this property bites hardest around the acceptance line. Straddling the weak-accept cut, papers in the 4–5 band crossed upward while papers in the 6–7 band crossed downward, both in the same experiment. The directional contrast is sharpest in exactly that middle range. Verdicts are shakiest exactly where accept and reject get decided.
3.1The caveat the authors added themselves
The authors qualify this pattern twice. The observation is descriptive, they write, and may partly reflect scale-boundary effects and regression to the mean. A low score has room to move only upward and a high score only downward, so the scale itself can manufacture a slope. Regression to the mean is common in repeated-measurement data and grows more pronounced the larger the measurement error and the more you analyze subgroups defined by a baseline value. As the epidemiology literature settled long ago, it is a candidate that must always be considered as a possible cause of observed change.
Read alongside another table in the same paper, though, that caveat looks unlikely to account for the whole spread. In the reliability audit of Appendix F, the team compared estimates from scoring the same PDF three independent times and averaging against estimates from a single scoring pass. Across twelve baseline-matched effects the two correlated at 0.995, with a mean absolute difference of 0.040 points. If scoring noise is that small, regression to the mean, which grows with error, has a hard time explaining shifts on the order of a full point. We state this conditionally rather than as a conclusion. The authors' caveat is legitimate, and it may also not be the whole story.
Is human review any steadier? No. In the NeurIPS consistency experiments, two independent committees disagreed on 23–25% of accept and reject decisions, and roughly half of the accepted list was estimated to change on a rerun. But using that fact as absolution loses the point. Human disagreement comes from differing judgments about content; the instability this paper captures comes from differences in how the same content is written. The latter has no counterpart on the human side. The absence of any slope across human rating bands, noted in §4.2 of the paper, shows precisely that difference.
Stricter rubrics did not cure it
When bias shows up in an evaluation pipeline, the first response is almost always to tighten the prompt. This paper built that response into the experiment. The standard protocol reproduces a conventional conference rubric; the strict protocol demands concrete evidence before awarding a high rating and explicitly instructs the judge not to reward wording that does not change the scientific judgment. It is the prompt someone who already knows about this problem would write.
The outcome split three ways. The scale came down decisively, the ordering between papers was largely preserved, and sensitivity to the prose did not shrink. That last item is the point of this section.
Original Pebblous diagram, reconstructed from §6.2, Table 8 and Table 3 of the paper. In the authors' own words, the strict protocol neither consistently strengthens nor consistently weakens the rewriting effects.
The drop in the scale is large and consistent. The mean overall score fell from 6.296 to 4.934, and 95% of papers received a lower score. Even so, the paper-level rank correlation between standard and strict scoring was 0.862. The absolute numbers moved as a block while the relative ordering mostly survived. Up to this point, tightening the prompt looks like it worked.
The trouble is in the third panel. Under the protocol that explicitly forbids rewarding presentation, the evidence framing effect on manuscripts rewritten by Opus 4.8 grew from 0.320 to 0.560. The paper's third headline finding says the same thing: strict prompting neither consistently strengthens nor consistently weakens the rewriting effects. It shifted the whole scale downward without producing robustness.
4.1The scale was already out of step with humans
Looking at where the absolute scores sit makes the motive for tightening clearer. Under the standard protocol all five judges scored above the human mean, by margins ranging from +0.418 to +2.934 points. Under the strict protocol, GPT-5 mini, GPT-5.5 and Sonnet 5 dropped below the human mean, Qwen 3.5 Flash landed almost exactly on it, and only Gemini 3.5 Flash-Lite remained 1.676 points high. Paper-level absolute error ran from 1.029 to 2.934 points, and rank correlation with human ratings from 0.395 to 0.621.
This score inflation is not unique to the paper. A separate study of 1,441 ICLR and NeurIPS papers reported a human mean of 5.70 against an LLM mean of 6.86, a gap of 1.16 points in the same direction. But turning that into "AI judges disagree with humans" misses the point of this report. Across prior work on peer review, correlation between one human and another sits between 0.14 and 0.41. The absolute level of agreement is not the issue. The issue is that the verdict moves when the same content is written differently.
4.2A protocol condition is becoming policy
Inside the lab, the protocol was a variable. The same thing is now appearing as conference rules. ICML 2026's policy on LLM use in reviewing abandons the allow-or-forbid binary and lets authors choose a review regime per paper. Policy A bans LLM use entirely in the pre-review stage; Policy B permits using a model to understand a paper and polish review prose while forbidding quality assessment, strength-and-weakness identification and delegated review writing. Papers that request A are assigned only to reviewers who commit to A, and authors who request A must themselves follow A when they review.
The last condition in that policy is where this chapter lands. Authors cannot see which policy the reviewers of their own paper actually followed. Within one conference, different papers are measured with different rulers, and which ruler was used is not disclosed to the person being measured. Set that next to the experimental result where changing the protocol moved the scale by 1.36 points, and it becomes hard to treat this setting as harmless to the outcome.
ICLR 2026 approached the same problem from another direction. Its reviewer guide permits LLMs as writing aids while adding a disclosure field to the review form, and states that failing to disclose puts the reviewer's own submissions at risk of desk rejection. Responsibility for the content of a review published under your name remains entirely yours. The norms clearly got stronger. But look at their direction and both policies address who wrote it. Neither contains a clause asking what the judgment responded to.
Swap the judge, flip the sign
This is where the grid design pays off. Same manuscript, same dimension, but you can watch the result change as the rewriter and the judge change. The paper describes it as a division of labor: the rewriter sets the size of the contrast between the positive and negative directions, and the judge sets the magnitude and the sign of the effect. The word "sign" belongs in that sentence. Change the judge and the same edit could raise a score or lower it.
| Judge | Response to rewriting |
|---|---|
| Qwen 3.5 Flash | Widest spread of responses. Score shifts vary greatly by condition |
| GPT-5 mini | Similar profile, also a wide spread |
| Gemini 3.5 Flash-Lite | Where the largest effects were observed. Still 1.676 points above humans under strict scoring |
| GPT-5.5 | Moves least. The same edits produce small score changes |
| Claude Sonnet 5 | Pulls effects downward. Many cells show scores falling even under positive-direction rewrites |
Qualitative summary from §6.1 and Table 3 of the paper. Per-judge figures are scattered across conditions in the original tables, so only direction and spread are carried over here.
The Sonnet 5 row is the counterexample this report must not delete. The summary "AI judges fall for polished writing" collapses on that single row. The paper's own conclusion is that rhetoric is not a universal reward hack. The risk is not that everyone is fooled in the same direction. It is that the sign of the same intervention depends on which judge you use, and the teams running these pipelines usually swap judges without knowing that.
A more uncomfortable observation sits next to it. The five judges agree reasonably well on how papers rank. Yet they do not agree on which papers benefit from rewriting. Agreement about quality does not imply agreement about intervention effects. That is exactly the blind spot in any pipeline that believes averaging several judges has bought it stability.
5.1Where the change got recorded
Besides the overall score, the judges also returned soundness, presentation and contribution scores. Since only the prose changed, common sense says any movement should land in presentation. The actual result was the opposite. Changes in the overall score co-moved weakly with presentation and more strongly with contribution and soundness. Qwen 3.5 Flash tracked mainly with contribution; Gemini 3.5 Flash-Lite under the strict protocol tracked strongly with both contribution and soundness. Sonnet 5's coupling was weak and unstable.
The paper is explicit that this is co-movement, not causal mediation and not a real change in quality. This design cannot establish which component is cause and which is effect. Even so, co-movement alone carries the practical implication. Breaking a rubric into components does not help if those components absorb the change in prose, because reading the scorecard will not tell you what the verdict responded to.
The premise that a fine-grained rubric explains a score has been one of the main arguments for adopting LLM judges. This result shakes it. The sentences changed, and the model wrote that up not as "the presentation improved" but as "the contribution and the soundness are different." That is the sentence that ends up in the audit trail.
5.2Fancier loops did not pay off
If you apply all six axes at once, iterate several times, and revise again after reading the judge's review, does the score keep climbing? The paper ran all three conditions, and the answer was mostly no.
- All axes at once. Manuscripts from the Opus 4.8 rewriter gained +0.289 under standard and +0.463 under strict scoring, beating that rewriter's single-axis average by 0.160. Those from the GPT-5.5 rewriter managed +0.021 and +0.045, essentially nothing, and came in 0.074 below its own single-axis average. Doing everything at once does not reliably produce more.
- Three iterative rounds. The Opus 4.8 rewriter captured most of its gain by round two (0.410 standard, 0.624 strict), and round three added almost nothing. For some judges the trajectory flattened or gave back gains from the earlier round.
- Revising on review feedback. Taking the judge's own review and revising once more lifted the overall average from 0.204 to 0.250. Compared against simply revising twice without any feedback, however, all four combinations came out worse. Listening to the judge did not beat one more blind pass.
It is tempting to read those diminishing returns as reassurance. Before doing so, separate the levels. What this paper measured is a shallow loop at the prompt level. What happens when the same sensitivity enters a training signal is answered by other work. In a study that used rubrics as a reward for reinforcement learning, scores from the judge used in training kept rising while scores from a stronger held-out judge peaked and then declined. The gap reached 3 points on one benchmark and 22 on another. A fixed bias would have moved the two curves in parallel; instead they diverged.
There is an intermediate case that stops short of training. Another team used bandit search to hunt for meaning-preserving stylistic edits that raise a specific judge's score, and the search did find them. That diverges from this paper's iterative loop, which fixed the axes in advance, applied them in order, and ran into diminishing returns. Faced with the same sensitivity, pushing on predetermined axes and searching against the score produce different answers. Measure the size of the sensitivity but leave out how an adversary searches, and you will underestimate the risk.
Do not equate the two results. Three prompt rounds and thousands of gradient steps differ in both the scale of pressure and the mode of search. What emerges from placing them side by side is this: a static sensitivity may look small, but the moment it becomes an optimization objective, search will find it. "We ran it a few times and nothing much happened" is an observation about the shallow loop, not a guarantee about the deep one.
5.3The empty cells are data too
Of 42,480 planned reviews, 42,396 produced valid records. The team left the missing 84 as missing rather than filling them with imputed values. Where those gaps clustered is the most telling detail in the whole experiment. On a single paper about jailbreaking, "JULI: Jailbreak Large Language Models by Self-Introspection," Sonnet 5 repeatedly failed to produce a valid record across 26 evaluation cells. One Qwen strict-simultaneous evaluation was explicitly refused by the provider's content check. The authors treat the pattern as consistent with safeguards firing while declining to assert the cause.
The practical implication is plain. Missingness in an evaluation pipeline is not random. A safety filter can systematically empty out the verdicts on one topic, and anyone reading only the aggregate mean will never see the hole. In a pipeline where refused items quietly disappear, the absence of a verdict becomes the verdict on that topic.
Six tests for whether the verdict holds
The previous five chapters all point one way. The axes that moved a lot were evidence and novelty, properties that require judgment from the reader. The axes that barely moved were vocabulary and register, properties you can check on the surface. Even the place where score changes were recorded was contribution and soundness rather than presentation. Work that perturbs the judge side rather than the manuscript reads the same. In an experiment that rewrote the instructions given to an evaluator while preserving their meaning, what determined the stability of the verdict was not model size but whether the property being evaluated was verifiable. Shake the manuscript or shake the judge, and the instability concentrates on properties that are hard to verify. The starting point of this chapter is that those are precisely the judgments organizations most want to hand to a model.
6.1What this result does not say
The authors set the limits of their own result. First, the sample is restricted to ICLR 2026 submissions for which full-text sources and public review metadata could be recovered; nothing guarantees that it generalizes to other venues or fields. Second, the missingness noted above may not be random. Third, the six dimensions are not orthogonal. Each rewrite is an intervention on the full text, so the value on each axis cannot be read as an independent coefficient on an isolated linguistic feature. Fourth, most manuscripts were scored only once. The results characterize the models and prompts that were tested; they do not represent human review or the full space of possible variations.
The authors' ethics statement belongs here too. They note the dual use openly: controlled rewriting can be used to diagnose rhetorical sensitivity, and it can also be used to tune a paper to an AI judge's taste without improving the research. They explicitly warn against reading a score increase as a gain in scientific merit, and against applying the framework to reviews in progress. This report keeps the same line. What is at issue here is a defect in a judgment system, not a technique for getting through it.
6.2Where detection cannot reach
Most prior attack research on AI review has dealt with detectable interventions: hiding instructions invisible to human eyes inside a manuscript to inflate ratings, or perturbing characters, words and sentences to unsettle the judgment. Those can be caught, and once caught they are handled as misconduct. The intervention in this paper has no such property. Stating the evidence more clearly and putting the contribution up front is what an advisor tells a student to do. When a variation indistinguishable from ordinary academic revision moves the verdict, detection-based defense cannot reach it in principle. What remains is robustness on the judge's side.
One experiment aimed at the same judgment system from the opposite direction. A separate team deliberately corrupted the relationships among a paper's results, interpretations and claims, using meaning-preserving edits as a control condition. The automatic reviewers failed to catch the corruption. The designs, the models and the tasks differ, so these cannot be called two sides of one experiment. Placed side by side, though, a direction becomes legible. Manuscripts whose content is broken pass through, and manuscripts whose content is intact see their scores move because the writing changed. That is why an organization handing review to a model has more to check than misconduct detection alone.
Yet the norms established over the past year stand almost entirely on the detection-and-disclosure axis. Conference policy was covered in the previous chapter, and publishers point the same way. Within what we verified, the guidance from Elsevier, Springer Nature and Wiley converges on three points: do not upload unpublished manuscripts to generative AI, human accountability does not transfer, and disclose your use. One analysis found that 83% of high-impact journals now have AI guidelines. Policy density is rising, and its direction is singular. Within the scope we checked, not a single clause requires that an AI-assisted judgment be stable under content-preserving variation.
This is a claim about absence, so the scope is stated narrowly. We verified two conference policies directly on their official pages and compared publisher guidance at the level of summaries. We do not rule out that such a clause exists in regulations we did not examine.
6.3Six requirements
The six items below follow directly from the paper's results. Each one states what you lose by skipping that check. None of them requires inventing anything new.
① Regression tests against content-preserving variation
Rescore the same content written a different way and check whether the verdict holds. This is not a new invention. NLP model testing formalized it in 2020 as invariance testing: apply a perturbation that does not change the label and require the prediction to stay the same. Its roots lie in metamorphic testing from software engineering. What is missing is not the method but the practice of applying that check to the LLM judge itself. Without it, your pipeline selects for the more forcefully written entry rather than the better one.
② Judge ensembles, for checking signs rather than averaging
Add judges from different model families and see whether the sign of the conclusion flips. A panel of small judges has been reported to correlate better with humans than a single large judge at a fraction of the cost, but the same work left panel selection as an open problem. Overlay this paper's observation and the purpose of an ensemble is redefined. Judges agree on ranking yet not on intervention effects, so the real value of an ensemble is not in averaging but in surfacing disagreement. Without it, one judge's taste becomes your organization's quality standard.
③ Rankings and relative comparisons instead of absolute scores
Between standard and strict scoring the absolute numbers moved by 1.36 points while the paper-level rank correlation stayed at 0.862. Hang your cutoff on an absolute value and a single line of prompt wording shifts everything; hang it on rank and it moves less. Note that this is an observation from inside this paper, and we did not verify an external source establishing that rank-based metrics are generally more robust. Without it, your acceptance line moves every time the scoring prompt is revised.
④ Double-checking borderline cases
Around the weak-accept cut, upward and downward crossings happened at the same time. The band with the largest directional contrast is the band that separates pass from fail. Without it, the most consequential decisions get made automatically at the point where the verdict is least stable.
⑤ Logging judgment failures and refusals
The event where 26 of 84 missing records clustered on a single paper is invisible in an aggregate mean. Article 12 of the EU AI Act requires automatic lifecycle logging for high-risk systems, and Article 15 asks that feedback-loop risk, where biased outputs become future inputs, be reduced at design time. These provisions target systems classified as high-risk, so reading them as an immediate legal obligation for an internal quality gate overstates the case. What is clear is the language regulators have already started using. Without it, a safety filter quietly empties the verdicts on one topic and the aggregate hides it.
⑥ Measuring agreement with human judgment on a schedule
A one-off correlation varies enormously by domain. In medical settings with clear scoring criteria a median of 0.69 has been reported, while studies specific to peer review cluster between 0.12 and 0.49. On top of that, human-to-human agreement is itself low. Without it, a correlation measured once gets treated as a permanent license.
6.4The tooling exists in papers, not yet in products
These six lines are more than exhortation because the materials already exist. Invariance testing was formalized in a 2020 ACL best paper, and dedicated benchmarks for measuring a judge's prompt sensitivity exist on the research side. Among the evaluation platforms we examined in this study, commercial and open-source alike, we found none offering whether a verdict holds when the same content is written differently as a first-class feature. Some platforms do offer alignment features that tune judge prompts to human labels, but matching humans and being unmoved by input variation are different properties.
The platform comparison rests on secondary summaries rather than individually verified official documentation. It asserts nothing about any specific product lacking such a feature, and should be read only as a statement that we did not find one within the scope we checked.
Demand-side adoption is already substantial. In a late-2025 survey of 1,340 respondents including engineers, product managers and executives, 53.3% said they use LLM judges to evaluate quality, factual accuracy and guideline compliance. Another 59.8% still run human review alongside. More than half of these organizations have handed judgment to a model, and we found no survey that asks whether that judgment is stable under variations in wording. The gap itself is the argument of this chapter.
The EU AI Act, the NIST AI RMF and ISO/IEC 42001 all push their requirements as far as documenting evaluation procedures and making them reproducible. Yet none of these frameworks asks you to measure whether the evaluator itself is stable under content-preserving variation. Even the closest provision, the feedback-loop language in Article 15 of the AI Act, targets biased outputs cycling back into training rather than an evaluator's sensitivity to its inputs. So it is more accurate to set these six lines up not as compliance items but as voluntary requirements that run ahead of regulation.
Why Pebblous follows this
Editor's Note. What follows is not part of the analysis above but the editorial team's perspective on the topic. The preceding six chapters rest on the original paper, external research and published policy documents; this chapter explains where Pebblous places them.
7.1When the judge becomes a model
Pebblous judges data quality in DataClinic and uses those judgments as gates in AI-Ready Data pipelines. The moment the judging party shifts from a person to a model, the nature of a quality metric changes. What used to be a function of "what state is this data in" becomes a function of "how did the model read this data." By holding content fixed and changing only the writing, this paper measured the instability of that function head-on. That is why we regard it as a test that anyone selling or buying evaluation automation has to pass.
7.2Evaluation output is itself new data
A model's evaluation output does not stop at a report. The moment it is reused as ranking, filtering, labels or reward, score differences that responded to writing style harden into the distribution of the training data. This is not speculation but a measured path. In reinforcement learning that used rubrics as reward, scores from the training judge rose while scores from a stronger judge peaked and fell, with a gap of 22 points on one benchmark. Article 15 of the EU AI Act, asking that feedback loops where biased outputs become future inputs be reduced at design time, expresses the same worry.
Placed on the axes this blog has covered, the position is clear. Benchmark contamination was training data leaking into evaluation. Bias differences between judges was a question of which model you asked. Silent model swaps asked whether the model you evaluated is the model being served. This report is the fourth axis: the same judge, the same content, a different verdict because the input was written differently.
7.3Five lines to settle on day one
For an organization adopting model judgments as a quality gate, the day-one question is not which judge model is better. Translated into working language, the requirements in chapter 6 come to five lines. Do you have a regression test that rescores content-preserving rewrites? Have you checked whether swapping the judge model flips the sign of the conclusion? Are you using absolute scores or rankings? Do you double-check cases near the threshold? Do judgment failures and refusals leave a record? All five are cheaper to settle before adoption than to bolt on afterward.
7.4Measure the wobble first
The market sells better judge models. What Pebblous can speak to sits one step earlier: how to measure whether the verdict wobbles at all. Just as data quality diagnostics measure the data, evaluation pipeline diagnostics have to measure the judgment. This is not blocked for lack of a method. Invariance testing was formalized back in 2020. What is missing is the practice of applying that check to the LLM judge itself, and filling that gap is where the next installment of our benchmark reliability series goes.
References
The figures in this report come primarily from the body, tables and appendices of the original paper. External corroboration was limited to peer-reviewed work, official statistics and surveys with a stated sample. Two conference policies were verified directly on their official pages; publisher guidance and evaluation platform features were compared only at the level of summaries, a limit noted in the body as well. As of 16 August 2026 we found no secondary press coverage of this paper, so none is cited.
Primary sources
- 1.Li, M., Wang, C., Li, X., Zeng, X., Li, D., Shi, P., Zhou, D., & Zhou, T. (2026). "How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review." arXiv:2608.08975v1 [cs.CL], 2026-08-10. The source of record for the figures in this report.
- 2.Project repository (MIT). MingLiiii/Dissecting_AI_Reviews. Rewriting and review prompts, pipeline code, anonymization harness and container definitions.
Academic · Bias and robustness in LLM judges
- 3.Zheng, L., et al. (2023). "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena." NeurIPS 36, 46595–46623. The canonical framing of position, verbosity and self-preference bias.
- 4.Wang, P., et al. (2024). "Large Language Models are not Fair Evaluators." ACL 2024, 9440–9450. Candidate presentation order flips pairwise preferences.
- 5.Dubois, Y., et al. (2024). "Length-Controlled AlpacaEval." arXiv:2404.04475. Length bias and length-controlled debiasing.
- 6.Stureborg, R., et al. (2024). "Large Language Models are Inconsistent and Biased Evaluators." arXiv:2405.01724.
- 7.Panickssery, A., et al. (2024). "LLM Evaluators Recognize and Favor Their Own Generations." NeurIPS 37, 68772–68802.
- 8.Lee, D., et al. (2025). "Are LLM-Judges Robust to Expressions of Uncertainty?" NAACL 2025. Epistemic markers lower the evaluation even when accuracy is unchanged. The closest predecessor to this paper.
- 9.Bhat, S., & Varma, V. (2026). "All Prompts Are Created Equal? Evaluating Robustness of LLM Judges Against Non-Adversarial Prompt Variations." Findings of ACL 2026, 38730–38745. Stability depends on whether the evaluated property is verifiable, not on model scale.
- 10.Verga, P., et al. (2024). "Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models." arXiv:2404.18796. The case for a panel of small judges (PoLL), and panel selection as an open problem.
- 11.Ribeiro, M. T., Wu, T., Guestrin, C., & Singh, S. (2020). "Beyond Accuracy: Behavioral Testing of NLP Models with CheckList." ACL 2020, 4902–4912 (Best Paper). Formalization of invariance testing (INV).
- 12.Yang, S., et al. (2026). "Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges." arXiv:2605.26156. Search-driven meaning-preserving stylistic edits.
Academic · AI peer review and manuscript-side intervention
- 13.Yang, Z., et al. (2026). "No Hidden Prompts Needed! You Can Game AI Peer Review with Presentation-Only Revisions." arXiv:2606.13044.
- 14.Baumann, J., et al. (2026). "Stop Automating Peer Review Without Rigorous Evaluation." arXiv:2605.03202. 86.2% agreement between 58 author complaints and detection labels.
- 15.Kaneko, M. (2026). "Paraphrasing Adversarial Attack on LLM-as-a-Reviewer." arXiv:2601.06884.
- 16.Li, X., et al. (2026). "Gaming AI-Assisted Peer Reviews Poses New Risks to the Scientific Community." arXiv:2606.10159.
- 17.Du, J. (2025). "TRAP: Probing Presentation Bias in LLM-Based Scientific Reviewing." Proc. 5th Workshop on Evaluation and Comparison of NLP Systems, 119–125.
- 18.Dycke, N., & Gurevych, I. (2026). "Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers." TACL 14, 465–488. Automatic reviewers miss deliberately corrupted content.
- 19.Ye, R., et al. (2024). "Are We There Yet? Revealing the Risks of Utilizing LLMs in Scholarly Peer Review." arXiv:2412.01708. Hidden prompt injection.
- 20.Lin, R., et al. (2025). "Breaking the Reviewer." EMNLP Findings 2025, 4819–4839. Character-, word- and sentence-level perturbations.
Academic · When evaluation scores become training signal
- 21.Gao, L., Schulman, J., & Hilton, J. (2023). "Scaling Laws for Reward Model Overoptimization." ICML 2023, PMLR 202:10835–10866.
- 22.Coste, T., et al. (2023). "Reward Model Ensembles Help Mitigate Overoptimization." arXiv:2310.02743. Ensembles mitigate overoptimization without eliminating it.
- 23.(2026). "Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning." arXiv:2606.04923. Divergence between the training judge and a gold judge: a 3-point gap on HealthBench-Hard and 22 points on ResearchQA.
Statistics · Audits
- 24.Barnett, A. G., van der Pols, J. C., & Dobson, A. J. (2005). "Regression to the Mean: What It Is and How to Deal with It." International Journal of Epidemiology 34(1):215–220. (Erratum: IJE 44(5):1748, 2015)
- 25.(2021). "Inconsistency in Conference Peer Review: Revisiting the 2014 NeurIPS Experiment." arXiv:2109.09774. 25.0% disagreement between committees.
- 26.(2023). "Has the Machine Learning Review Process Become More Arbitrary as the Field Has Grown? The NeurIPS 2021 Consistency Experiment." arXiv:2306.03262. 23% disagreement, with roughly half the accept list estimated not to reproduce.
- 27.(2026). "Impact of Large Language Models on Peer Review Opinions." arXiv:2604.19578. Across 1,441 ICLR and NeurIPS papers, humans 5.70 versus LLMs 6.86.
- 28.Paper Copilot. "ICLR 2026 Statistics" (compiled 2026-01-28). Mean review rating of accepted papers, 5.39. We re-verified the paper's footnoted value against this source.
- 29.ICLR. "A Retrospective on the ICLR 2026 Review Process" (2026-03-31). 19,525 valid submissions, 76,139 reviews, 27.4% acceptance rate.
Policy · Institutions · Surveys
- 30.ICLR. "ICLR 2026 Reviewer Guide." New LLM-use disclosure field and sanctions for non-disclosure. Verified directly on the official page.
- 31.ICML. "ICML 2026 Policy for LLM Use in Reviewing." The dual Policy A / Policy B regime. Verified directly on the official page.
- 32.Wang, Y., et al. (2026). "A Cross-Disciplinary Analysis of AI Policies in Academic Peer Review." Learned Publishing, doi:10.1002/leap.2035. 83% of high-impact journals have AI guidelines.
- 33.Elsevier, "Generative AI policies for journals." Cited in the body with the caveat that this was compared at the level of a summary.
- 34.Springer Nature, "AI editorial policies." Cited alongside Wiley's author and reviewer guidelines with the same summary-level caveat.
- 35.EU AI Act, Art. 15 (accuracy, robustness, cybersecurity) and Art. 12 (record-keeping). These provisions apply to high-risk systems.
- 36.NIST. "AI RMF to ISO/IEC 42001 Crosswalk." TEVV documentation requirements under MEASURE 2.1.
- 37.LangChain. "State of Agent Engineering" (fielded 2025-11-18 to 12-02, n=1,340). LLM judges in use by 53.3%, human review run alongside by 59.8%.