Executive Summary

A fairness audit usually runs like this. You swap in a demographic attribute or something correlated with it, then watch whether the model's decision changes. If it changes, that counts as evidence of bias. But a neighbourhood indicator correlates with race and at the same time carries real information about environmental exposure. A rational decider should change its decision when such a value changes. The fact that a decision changed therefore cannot, on its own, separate discrimination from sound inference.

A paper posted to arXiv on August 24 puts a number into that gap. It computes exactly how far a prediction-optimal decision rule holding the same information would lean on the variable, and takes that value as the reference. The authors asked four LLMs to pick which of two patients to see first. In the condition where the proxies carry no information about the outcome, so the reference is exactly zero, all four models reading neutrally named fields still leaned on them. The experiment that follows is the more uncomfortable one. Raising the strength of the evidence across four levels barely moved the models' reliance at all.

On the regulatory side that gap turns into a practical problem. Demands to show that bias has been identified and evaluated are growing, and an audit with no quantitative reference will penalise reliance the evidence supports while taking comfort in suppression that a field name alone holds down. False alarms and misses come out of the same run. Computing that reference, though, requires knowing the outcome rule, and observational data alone cannot identify it without strong assumptions. That is why the study runs on a fully specified synthetic generating process. The authors placed seven real clinical datasets on the same structural axes to show that the manipulated range straddles reality, but that comparison covers covariance structure only.

Key figures

Source: Wu & Xiao, arXiv:2608.22887 (2026-08-24)

+19.7pp

Reliance on a proxy with zero information

Claude Sonnet 4.5 at 12 attributes under neutral labels. The evidence-warranted level in this condition is exactly zero

0.05

Slope at which reliance follows the evidence

A slope of 1 would be calibrated tracking. Pooled over the three identically configured models, with all eight fits between −0.08 and +0.15

+15.7pp

Reliance revived by provably clean examples

Measured with examples orthogonalised so the proxy-outcome correlation is exactly zero. A rational learner given the same examples showed −1.0

24pt

Reliance gap at indistinguishable accuracy

Across conditions and models whose accuracy cannot be told apart statistically, proxy reliance differed by as much as this

1

Audits only ask whether the decision changed

Emergency-department triage is a leading route by which LLMs enter decisions about people. Large-scale evaluations on simulated clinical decision tasks have already been run under physician review. In settings like these, the law draws a line that behavioural evaluation has so far struggled to redraw. Using a variable because it predicts the outcome is ordinarily treated as legitimate inference. Using that same variable because it stands in for a protected attribute such as race or gender is, on the dominant legal account, proxy discrimination.

The difficulty is that a single variable can be both things at once. A neighbourhood indicator correlates with a patient's race and also carries real information about environmental exposure. How much reliance on it is justified is a quantitative question rather than a yes-or-no question, and scrutinising which inputs a system receives cannot settle it.

The audits in use today do not answer that question. The dominant paradigm changes a demographic attribute or its correlates, observes whether the model's decision changes, and treats a changed decision as evidence of bias. Yet a rational decider changes its decision whenever the manipulated attributes carry genuine signal. An unchanged decision, conversely, can mean fairness or ignored evidence. Benchmarks of social bias carry the same limitation in a different form. They are largely built so that group membership should not be used as evidence, which leaves them silent about situations where group-correlated information legitimately should matter.

Others have pushed at this limitation. Difference-aware evaluations ask in which situations group information should legitimately matter, and causal audits have been proposed to separate job-relevant pathways from impermissible ones. This paper takes a slightly different route. A recent line of work held LLM confidence against Bayesian updating as an explicit normative standard, and what turned up there was structured bias rather than noise. The authors carry that move over to discrimination. Where the confidence studies used a posterior probability, this study puts a different value in its place.

One thing was missing. A quantitative reference for how much reliance the available evidence supports. Without one, an audit can report a difference but cannot render a verdict. Zengqing Wu and Chuan Xiao at the University of Osaka set such a reference: the degree of reliance on group-correlated attributes that a prediction-optimal decision rule with the same information would show, which the paper calls the evidence-warranted level.

The reference is statistical and task-conditional. It measures the reliance warranted by the specified outcome rule and ranking loss, and it does not by itself establish that such reliance is legally permissible or morally justified. Computing it requires knowing the true relationship between attributes and outcomes, which observational data alone cannot identify without strong assumptions, so the study runs on a fully specified synthetic generating process. That is a precondition for defining the quantity, not a convenience.

2

Every model leaned on proxies with no information

Every experiment runs on one task. Prompted as a triage specialist, the model sees two patients described by named numeric indicators and picks one to prioritise. The clinical vignette is only the surface, and the statistical structure underneath is fully known because the authors generate the patients themselves. Each patient carries 6, 12 or 18 attributes at a fixed two-to-one ratio of legitimate to proxy attributes. Legitimate attributes determine the true risk and proxy attributes correlate with the protected attribute, which never appears in any prompt. The setup mirrors deployments where the protected attribute is withheld but its correlates are not.

The measurement is a causal intervention. For each pair the authors run one factual arm and two counterfactual arms, setting the target patient's protected attribute to each value and regenerating only that patient's proxy attributes through the structural equations with all exogenous noise held fixed. The comparator patient is never intervened on. The difference between the probabilities that the model picks the target under the two settings is the directed proxy-specific effect. Within each structural condition, 99 fixed evaluation pairs are shared across every rule, label, model and evidence condition, so the comparisons are paired.

Four models were tested. Claude Sonnet 4.5, DeepSeek-V4-Flash-0731 and Qwen3.7-max ran on an identical configuration, at temperature zero with forced tool choice producing single-token answers. GPT-5.6 Terra does not accept a temperature setting and runs with reasoning disabled at the provider default, so it is reported as a separate channel throughout the paper. Pooling it with the other three changes no verdict.

The key experimental lever is the proxies' true predictive value. In the zero-information condition the proxies carry no outcome information beyond what the legitimate attributes already hold. The evidence-warranted level is therefore exactly zero, and any proxy effect measured there is unwarranted by construction. With neutrally named fields, Sonnet showed proxy effects of 11.1 points at 6 attributes, 19.7 at 12 and 14.6 at 18. All 95% confidence intervals excluded zero.

At field level, the two proxy fields sit among the strongest correlates of the model's choices, and the ratio of proxy to legitimate correlation rises from 1.3 to 4.2 as the attributes become more entangled. The models are not merely brushing against the proxies. They are putting weight on fields that carry no diagnostic content. Excess reliance at 12 attributes was near identical across the three identically configured models, at 19.7, 22.2 and 20.2 points. The paper reads that as a shared baseline rather than a quirk of one system.

2.1One signal yields three verdicts

Switch to the condition where the proxies genuinely predict the outcome, and the verdict flips with everything the model sees held fixed. Under neutral labels, reliance was not distinguishable from the reference at 6 and 18 attributes and exceeded it by 7.1 points at 12, with an interval from 1.0 to 13.6. An attribute-substitution audit would report a changed decision in both conditions alike, which means flagging one correctly and the other wrongly. In the same condition DeepSeek-V4-Flash-0731 exceeded the reference by 12.1 points and Qwen3.7-max by 10.1.

The third verdict appears when the field names change. Renaming the proxy fields with neighbourhood and occupation terms instead of biomarker terms pushed Sonnet's reliance below the reference by 9.6 to 12.1 points, with every interval across dimensions excluding zero. The suppression carries no measurable accuracy cost where the design can detect one, so the paper does not call it overcorrection. It does deviate from the warranted level in the opposite direction, and an audit that looks only for excess sensitivity never sees this deviation at all.

One behavioural signal splits into three verdicts Evidence-warranted level (the reference) Over-reliance +19.7pp Zero-information proxies Neutral labels, reference exactly 0 Three identical setups all near 20pp Warranted Not distinguishable Informative proxies 6 and 18 attributes match the reference Only 12 attributes exceed it, by +7.1pp Under-reliance −9.6 to −12.1pp Social labels (Sonnet) Fields renamed to neighbourhood terms An attribute-substitution audit reports the same changed decision in all three cases
▲ The verdict structure of §2.1 and Fig.1 of arXiv:2608.22887, redrawn as a concept diagram | Original Pebblous diagram
3

The evidence moved and reliance stayed put

If the verdict changes with the condition, the next question follows naturally. Does the model's reliance move with the evidence at all? The authors scaled the proxies' share of the outcome signal across four levels while holding total outcome variance, and therefore task difficulty, constant. The reference is a Bayesian regression fitted on exactly the same eighty examples the model saw, which the paper calls the ideal learner. The reason for using a decider with the same amount of information rather than a perfect-knowledge one is plain. Holding a finite-evidence learner to the perfect-knowledge standard would mistake ordinary statistical caution for model failure. This ideal learner's own warranted reliance rises from 0 to 23 points across the four levels.

The evaluation items are identical at every level. Each pair of patients is judged under every evidence strength, so the comparison is paired, and that is what gives the design its power. One regression line is fitted per model. A slope of 1 means calibrated tracking of the evidence and a slope of 0 means no tracking at all.

Across four models and two label conditions, the slopes of all eight fits fell between −0.08 and +0.15. Every interval excluded calibrated tracking and none excluded zero. Because the upper bounds still allow weak tracking, up to 0.49, the paper does not write that insensitivity has been proven. The defensible description, it says, is severe undertracking over the tested range. Averaging the six slopes from the three identically configured models gives 0.05.

What each model does instead is operate at a large positive level that the evidence does little to move. The intercepts are scattered between 18.3 and 30.4 points. Across all eight fits, differences between models and between label conditions appear almost entirely in the level rather than in the slope.

The reference climbs while the model stands still 0 10 20 30 Model proxy effect (pp) 0 8 15 23 Ideal learner's evidence-warranted level (pp) Calibrated tracking (slope 1) Observed models (slope 0.05) Intercept +18.3 to +30.4pp All eight fits span −0.08 to +0.15, each excluding a slope of 1 and none excluding 0
▲ The dose-response result of §2.2 and Fig.2 of arXiv:2608.22887, redrawn as a schematic concept diagram | Original Pebblous diagram

3.1The rise in accuracy is close to an illusion

Raw accuracy does improve as the evidence grows stronger. The trouble is that raising the proxies' predictive value also moves the true ranking, so a model that never changed a single decision would score better too. The authors ran a frozen-policy control, taking each model's choices at the weakest evidence level and scoring them against the strongest level's truth. That accounted for 12.1 of Qwen3.7-max's 15.2 points of accuracy improvement, on the neutral arm.

The residual gains that came from actually responding to evidence strength are small and signed in both directions. Only one of the eight fits excludes zero, Qwen3.7-max under social labels at 4.0 points with an interval from 1.0 to 7.6. Two are negative with intervals touching zero at the boundary, Sonnet under neutral labels at −6.6 and GPT-5.6 Terra under neutral labels at −5.1. A decision policy does move under stronger evidence. There is no guarantee the movement is beneficial.

Which verdict an audit returns is not determined by what the model does differently as the evidence changes. It is determined almost entirely by where a deployment's evidence structure happens to sit relative to the fixed level of reliance the model already carries with it. When the reference sits below that level the verdict is over-reliance, and when it sits above, the same model is judged to be under-relying.

4

A few examples undo what the field names held down

Reliance falling under social labels looks like a safety property. The paper reads it as a shortcut the field names switch on rather than a stable commitment. Three strands of evidence back that.

First, the suppression is conditional. It appears only when the proxy channel is open and the task is ambiguous. Across four dependence-structure conditions run without examples, the social-versus-neutral contrast excluded zero exactly where neutral-label reliance was itself nonzero, at −8.25 and −6.7 points. Where the channel was closed or an explicit scoring rule was supplied, the contrast vanished. A model that is told what to compute does not consult the field names.

Second, the suppression is uniform across decision difficulty. Splitting the evaluation pairs by the true risk gap, it shows up in every stratum, with contrasts of −17.5, −25.0 and −22.5 points and all intervals excluding zero. That rules out the reading that social names merely push the ambiguous cases toward caution.

Third, the strength of the suppression is a gradient across providers rather than a property of the model class. It runs to −20.7 points for Sonnet, −15.2 for GPT-5.6 Terra, −9.1 for Qwen3.7-max and −6.6 for DeepSeek-V4-Flash-0731, and the example-induced increase follows the same rank order. Unwarranted reliance under zero-information proxies is nearly identical across providers while the suppression varies threefold, which fits a shared baseline bias with a protection added separately by each provider. A four-point ordering is a pattern rather than a statistical test, but its direction matches independent evidence about differing alignment pipelines.

Suppression strength varies threefold across providers 0 (neutral-label baseline) Sonnet −20.7pt GPT-5.6 Terra −15.2pt Qwen3.7-max −9.1pt DeepSeek-V4-Flash −6.6pt
▲ Recreated from arXiv:2608.22887 §2.3 and Fig. 3(c)'s provider suppression gradient | Original Pebblous diagram

4.1Examples washed clean of correlation still revive reliance

The most decisive result comes from in-context examples. Adding examples to the prompt raises reliance clearly above zero in every model, even under social labels. An innocent explanation is available here, namely that the examples taught the model that the proxies predict the outcome. To close it off, the authors used calibration sets whose proxy-outcome correlation is exactly zero by explicit orthogonalisation.

Even given these provably clean examples, social-label reliance jumped to 15.7 points, with an interval from 9.1 to 22.2. A rational learner given the same examples showed −1.0. Deliberately inducing a correlation adds a further 12.6 points on top, so misleading evidence does compound the effect, but clean examples alone already lift reliance far above zero. All six example-supplied social cells sat clearly above zero, and the social-versus-neutral contrast persisted in only three of the six.

Cleaning few-shot examples of spurious correlations is not a sufficient mitigation. The moment a prompt gains examples, the apparent zero-shot protection is forfeited. An uncomfortable practical conclusion follows. The regime in which models look safest, zero-shot prompts with socially explicit field names, is precisely the regime least like a deployed system, since real deployments typically supply in-context examples and rename or abstract their features.

5

Reliance splits by 24 points where accuracy does not

For an accuracy-only evaluation to catch proxy reliance, accuracy and reliance would have to be pulled by the same underlying factor. Across this study's manipulations they did not move together. Holding dimension fixed and manipulating only the structure moves proxy reliance by as much as 24 points while accuracy stays within estimation noise. Partialling dimension out of the correlation between the error gap and proxy reliance leaves an association indistinguishable from zero, at 0.26 with an interval from −0.12 to +0.56.

The most telling result came from cutting the number of examples. Going from eighty to twelve left accuracy statistically unchanged while lowering proxy reliance by 14.6 points. That means the unwarranted reliance was induced from the examples rather than being a by-product of degraded capability, and it also means improving accuracy cannot be assumed to reduce it. The same dissociation appears across models. On the cell where the four models diverge most, proxy reliance spans 16.7 points while accuracy spans about 2. Conditions and models with indistinguishable accuracy differ by up to 24 points in proxy reliance, which is the quantitative content of the claim that accuracy-only evaluation is blind to the proxy channel.

One observation runs the other way. Within cells, incorrectly answered pairs tended to be the most protected-attribute-sensitive, with a mean correlation of 0.22. This pair-level association cannot identify a causal direction, and it does not alter the condition-level picture that audit design depends on.

5.1Adding attributes can raise reliance or lower it

Attribute count and dependence structure act on different channels. Dimension harms structure discovery, not rule execution. When the model has to infer which fields matter from eighty examples, its regret, the gap to the ideal learner's error on the same examples, rose monotonically from 19.7 to 27.8 to 36.9 points at 6, 12 and 18 attributes, with all twelve intervals excluding zero. Without examples, error barely moves, so the task never became unlearnable. And when attributes are more entangled, the gap in example-based learning shrinks instead, by as much as 13 points.

The suspicion that changing the attribute count simply makes the task harder was blocked at the design stage. True risk is normalised to unit variance at every attribute count so that ranking difficulty is matched across dimensions, and the proxy contribution is normalised as well so the signal does not grow mechanically with the number of proxies. Manipulation checks hold the structure statistic and the proxy-channel strength constant by re-solving the latent alignments per dimension. The 17.2-point rise from six attributes to eighteen also exceeds that contrast's minimum detectable effect of 14.7 points.

The most counter-intuitive result in the paper appears here. The effect of dimension on proxy reliance reverses sign with the dependence structure. Under low dependence, adding attributes raises proxy reliance, moving from 15.7 to 23.2 points under neutral labels. Under high dependence the same manipulation lowers it, from 12.1 to −1.5 points. Under social labels the high-dependence trajectory falls further and crosses zero, from 13.1 points at 6 attributes to −8.6 at 18, with the latter interval excluding zero. The interaction replicates across the three identically configured models with a pooled magnitude of −15.1 points, interval −22.4 to −7.9. The authors state explicitly that both directional reversals are exploratory and post-hoc.

The same manipulation moves opposite ways by structure 0 10 20 Proxy reliance, neutral labels (pp) 6 12 18 Attribute count (k) Low dependence: +15.7 → +23.2pp High dependence: +12.1 → −1.5pp crosses zero
▲ Recreated from arXiv:2608.22887 §2.4 and Fig. 4's dimension × dependence-structure interaction (neutral labels, schematic layout) | Original Pebblous diagram

No single main effect of dimension on proxy reliance can be stated. What growing the attribute space does depends on how entangled the attributes are. Both structural regimes exist in real data. Computing the same structure statistic on seven public clinical datasets places six of the seven inside the manipulated band, three near each level. Deployments built on administrative utilisation counts and deployments built on enzyme panels genuinely fall on opposite sides of this sign reversal. The anchoring compares covariance structure only, and involves no model calls.

6

An audit run unlike deployment measures something else

Audits are rarely run under the conditions of deployment, and two kinds of error follow from the mismatch. When the audit population differs from the deployment population but the model sees the same fields, the error is a distribution shift. Importance weighting, which reweights audit cases to match the deployment distribution, can repair it in principle, provided the audit cases cover the deployment distribution's support. The case where the audit shows different fields, or the same fields under different names, is not like that. The measured quantity itself changes, so no reweighting can repair it.

Measured side by side, the correctable error turns out to be the smallest. Below are the four terms as measured on the neutral arm.

Audit-deployment mismatch Measured bias Does reweighting repair it?
Full proxy omission 13.9pt Residual still excludes zero after reweighting
Mismatched semantic labels 8.25pt Not possible in principle
Partial omission (one of two retained) 4.6pt Residual not distinguishable from zero after reweighting
Distribution shift 2.6pt Unresolved at this sample size

▲ The error decomposition in §2.6 of the paper. Every value is reported on the neutral arm, and the interval on distribution shift runs from 0.4 to 5.1

Label mismatch cannot be reweighted away by construction. Reweighting does not turn it into the same quantity measured under a different distribution, because a different quantity is being measured in the first place. Full omission survives importance weighting with a residual that still excludes zero. Whether the distribution-shift term is correctable remains unresolved at this sample size, since its residual never separated from zero and is therefore not estimated precisely enough to distinguish successful correction from remaining bias.

The ordering of the magnitudes is the substantive point. The two terms reweighting cannot touch, full omission and label mismatch, are 5.3 and 3.1 times the size of the distribution shift, the only term the standard correction targets. The same structure explains why the label-mismatch row alone carries no interval. That value is the difference between the neutral and social full-panel reference cells, so no resampling interval attaches to it.

Mismatch arises in another way as well. Two design families the authors identified in piloting drive the proxy effect to exactly zero. When the prompt states the scoring rule together with explicit field weights, the model executes the rule, error rates approach zero and no semantic manipulation has any effect. When evaluation pairs are drawn without stratifying on the true risk gap, most comparisons are wide enough that the ranking is obvious from the legitimate fields alone and the proxies never enter the decision. The proxy channel opens only when no scoring rule is given and all fields carry homogeneous technical names, so the model must judge relevance itself, and when the comparisons include near ties. A zero measured under an explicit-rule or wide-margin design does not mean the model is unbiased. It means that design cannot detect the effect at all.

The practical rule for auditors reduces to one line. An audit that shows the model fewer fields than deployment does, or the same fields under different names, is not an approximation of the deployment measurement but a different measurement. And that difference was the largest error this study observed.

6.1Where these results do not reach

The paper draws three boundaries around itself. The generating process is synthetic because the evidence-warranted level requires a known outcome rule, an abstraction with documented risks. The real-data anchoring covers structural axes rather than outcome realism, with semi-synthetic designs the natural next step. The evidence sweep covers proxies contributing up to one third of the outcome variance, so the regime where proxies dominate is not covered. The correctability of the distribution-shift term also remains open, as noted above.

Two further things are worth holding in mind when reading the numbers. Cross-dimension contrasts cannot share evaluation items and are therefore between-sample, with minimum detectable effects ranging from 8 to 16 points depending on the cell. Differences across dimensions smaller than that are reported by the paper as directional rather than confirmatory. And because displayed values are standardised within sample, regenerating the target's proxies in a counterfactual arm can move a rendered comparator value by one unit in the second decimal, which affected five of the 99 pairs at 12 attributes. Excluding those pairs moves no headline proxy effect by more than 1.5 points.

The replication side is solid. The core cells were rerun for two models under two entirely fresh seeds, and unwarranted reliance replicated in all 16 zero-information cells, ranging from 7.1 to 26.8 points. The dose-response verdict replicated in all 8 curves. Raw records, code, seeds and an offline design verifier are published on GitHub and Zenodo, so every reported quantity can be recomputed.

The paper also notes the regulatory context. Regulation increasingly demands demonstrations that bias has been identified and evaluated, and is moving toward permitting access to protected attributes for exactly that purpose. The authors point to this trend by citing legal scholarship on Article 10(5) of the EU AI Act. They conclude that such demonstrations need a quantitative evidence-warranted reference to be meaningful. Audits without one produce both false alarms, penalising reliance the evidence supports, and misses, crediting fragile surface suppression.

Editor's Note: There is an illusion Pebblous runs into often in data quality work. The evaluation dataset gets treated as the ruler that measures the model, and what that ruler was built to be able to measure goes largely unexamined. What this paper shows is that the markings on the ruler produce the verdict. How many fields you include, what you name them, and whether you attach examples yield over-reliance, warranted reliance and under-reliance as three different conclusions about the same model. The AI-Ready Data proposition that data curation and labelling design determine the validity of model behaviour evaluation takes a concrete form here, as fairness auditing. It is also why a record of how an audit dataset was designed becomes an output as important as the audit result itself.

The paper is available at arXiv:2608.22887. Raw records and replication code are published in the GitHub repository and the Zenodo archive.

R

References

Primary Source

Related Work