Executive Summary
This article looks at a case where the report card for a brain MRI anomaly detection AI moved because of the scoring procedure rather than the model. The source is a paper posted to arXiv on October 1, and its authors named the evaluation protocol they propose MIRTO. Its subjects are four models that learn from healthy brain scans alone and then mark, on their own, the places that look wrong in a scan that holds a lesion. Models like these are usually ranked by a single score, and behind that score stands a line of choices that almost never get reported.
The two numbers that stand out are 0.873 and 0.583. One diffusion model's anomaly map was read in the wrong axis order, the cropped field of view was never restored, and the voxel-level score fell from 0.873 to 0.583. The model, the weights and the test data were all untouched. In that same accident the slice-level score moved by only 0.112, so anyone watching the summary number saw nothing go wrong.
Sections 1 through 4 follow what the paper measured and the figures its authors recorded. Section 5 carries the finding over to training-data quality, and that move is this article's reading rather than anything the paper claims.
Key figures
All four cards below come from the same paper. The first two set a single accident beside the two scores that watched it, and they disagree. The third is a signal that surfaces only once the overall average is broken apart. In the fourth, changing the metric changed what the score is really tracking.
Source: Kafee Hernashki and Chatterjee, MIRTO: a registration-gated, multiverse-tested evaluation protocol for unsupervised anomaly segmentation in brain MRI, arXiv:2610.02136 (2026-10-01), abstract and text
0.873→0.583
Voxel-level score under a mismatched axis order
The model, the weights and the test data never changed
0.112
How far the slice-level score moved in that same accident
The voxel-level score moved 0.290. By median the slice figure is 0.083
24%
Share of subjects scored backwards
Their per-subject score sits below 0.5, which a coin flip beats
0.95 vs 0.14
How much "which model" explains the score
0.95 for the voxel-level score, 0.14 for lesion sensitivity
What Dropped the Score Was Not the Model
Brain MRI anomaly detection never learns what a lesion looks like. A model is shown nothing but scans of healthy brains until it has absorbed what normal looks like, and then it is handed a new scan and asked to mark the places that depart from that normal. What comes out is a single map the same size as the scan. Every voxel in it carries a score saying how unusual that spot is, and the map is called an anomaly map.
Scoring means holding that map up against the answer key. The answer key is a label a person painted over the tumour region. The test data in this paper is 312 subjects from BraTS 2020, and all four models were trained on the same healthy scans and tested on the same subjects. The four compared are REFLECT, cDDPM, UCCD and AnomalyDINO. One restores a normal scan in a single pass, two are conditioned diffusion models, and one does no training at all and simply remembers the features of an off-the-shelf foundation model.
The usual score is the voxel-level AUROC. Pick one lesion voxel and one healthy voxel at random, and it is the probability that the model gave the lesion voxel the higher score. Close to 1 means the model marked the right places, 0.5 is a coin flip, and anything below 0.5 means the model is pointing the wrong way.
The accident happened where the map gets read. Preprocessing differs from model to model, so each one stores its map in its own way. Which axis gets written first, how far the field of view was cropped, how much the resolution was reduced, all of it varies. Comparing against the answer key means undoing those choices one by one, and the authors wrote that this undoing is easy to get wrong and hard to see. In practice a mapping was used that read cDDPM's map in the wrong axis order and ignored the cropped field of view, and the voxel-level score under it was 0.583. Undoing the same map correctly gave 0.873. The difference is 0.290, with an interval from 0.277 to 0.302.
A mismatched axis order means the whole map is flipped or twisted. A three-dimensional scan is stored as a block of numbers, and which direction was written as the first axis travels only as a convention outside the file. Break that convention and a signal from the front left lands at the upper right. The model marked the tumour correctly, and the scoresheet is looking somewhere else.
Why the Summary Score Stayed Quiet
The coldest part of this paper is not the number that fell. It is the number that did not. In the same accident the slice-level AUROC moved by 0.112, and laid out per subject the middle value is 0.083. All of that while the voxel-level score moved 0.290.
The two scores ask different questions. The voxel-level score asks of every voxel whether it is lesion. The slice-level score takes the MRI one cross-section at a time and asks only whether that slice contains a lesion, with the single highest anomaly score inside the slice standing in for the whole slice. Twist the map and the coordinates of every individual voxel go wrong, but much of which slice looks suspicious survives. So the fine-grained score collapses while the summary score looks fine.
So what caught the accident? A diagnostic that never touches the answer key. Instead of the overall average, each subject is pulled out and scored alone, and the count is of how many land below 0.5. The authors call that share the anti-location rate. Under the misread map, 24% of cDDPM's test subjects fell into it. Packed into a single average it stays invisible, and spread back out per subject it is plain that one in four is inverted.
That signal does not stand alone. The paper sets out five quantities that can be measured without labels. Besides the anti-location rate there is rank stability, which checks whether per-subject scores keep the same order across two conditions that should be pointing at the same tumour, and shared conspicuity, which rests on the fact that how visible a tumour is belongs to the subject, so per-subject scores ought to move together even when the method changes. Neither one looks at the answer key. That means the question of whether the scoresheet is in the right place can be asked apart from how well the model did.
What this measures is not the anomaly detection model's ability. It is whether the scoresheet puts the model's output in the right place. Break that one thing and the score comes down near a coin flip no matter how precisely the model picked the lesion out. And that collapse barely registers in the summary score most papers report.
The Same Accident Happened Twice
Once could be folded away as somebody's slip. The authors recorded that they met the same kind of accident a second time while assembling this benchmark. That time it came while exporting REFLECT's maps. Of the test subjects, 26% were scored backwards, and the voxel-level score came down from 0.936 to 0.616. Through all of it the slice-level score stayed put at 0.915.
The second accident matters not because its numbers are larger. It matters because it happened again while the same people were building a benchmark with the same care. As long as every method stores its maps with its own axis order, its own cropping and its own resolution, every comparison creates a transform that has to be undone, and that transform fails quietly.
That is why MIRTO puts a registration gate in front of every comparison. The geometry gets checked first, and only the pairs that pass go on to be scored. The gate's own power was validated in advance on synthetic volumes whose answers are known. It passes the identity, a field of view with 40% of a sphere removed, a one-voxel erosion and a 20% dilation, and it fails an in-plane flip, a transpose and a two-slice offset. Rank stability and the anti-location rate, two label-free diagnostics by themselves, caught every flip and transpose at a false-alarm rate of 0.05 with as few as thirty subjects. On the real data, 1,342 of 1,344 method–subject pairs passed the gate. That count comes from pairing each of the four methods with 312 test subjects and 24 validation subjects.
Running All 15,552 Scoring Paths
Axis order is only one of the choices hidden inside a scoring procedure. Which scan serves as the reference, whether to look at the whole brain or only where the fields of view overlap, which data the threshold is set on, how many cubic centimetres of false positives to allow, whether small blobs count as lesions, how much overlap counts as a hit, whether per-subject scores are pooled by mean or by median: every one of those is a choice. The paper laid out ten such axes and multiplied the options on each of them together. That gives 155,520 combinations per method.
From there only the combinations anyone could defend were kept. The reference scan is fixed to the canonical one, the voxel volume is the true one, and only three threshold protocols are admitted: validation, conformal and test-matched. Trimmed that way, 2,592 combinations remain per metric, and 15,552 across the six metrics. Every comparison is then repeated over those 15,552 paths, with paired subject bootstrap intervals and a correction for multiple comparisons attached. It amounts to carrying multiverse analysis, which began in psychology, over to medical image scoring.
The idea did not originate with this paper. Multiverse analysis and specification curves, which check a single conclusion across every defensible analysis path, came first, and in the many-analysts design, which hands the same data to several research teams at once, even reasonable paths landed on different conclusions. Medical imaging had its own warnings on the pile. Challenge rankings were reported to shift depending on which metric and which aggregation were used, and choices made after looking at the test data were reported to inflate performance estimates. What the paper adds is the one thing that always comes attached to carving out anomalous regions with a threshold: writing down, rather than hiding, the point at which something gets called an anomaly.
Run all of them, and what holds the score differs from metric to metric. For the voxel-level AUROC and the voxel-level area under the precision-recall curve, "which model was used" explains 0.95 or more of the variation in results. For Dice it is 0.77. For lesion sensitivity it drops to 0.14, and the result is driven instead by how a lesion is defined and how much overlap counts as a hit. The same set of models can hand the win to a different method depending on which metric is used to declare it.
The same approach turned up one more thing. A Dice advantage that was significant when measured at a threshold set on validation data disappeared once the false-positive volumes actually realised at test time were matched against each other. The authors used an exact identity to pin that difference on carrying a threshold across. There is a case running the other way too. Changing how REFLECT aggregates its latent space, with no training involved, raised Dice by 0.052 at the same false-positive burden. It was not a better model that lifted the score but a better scoring convention.
The authors drew a line around their own results as well. They tested nine hypotheses against criteria fixed in advance, and they wrote that because the cohort used to refine the protocol was the same one used in the tests, every inference is exploratory.
Whose Number Is the Performance Number?
When model performance gets reported to us, what we actually receive is one number. Which coordinate frame the predictions were moved into on the way to that number, which data the threshold was picked from, how many separate cases were folded into a single line: almost none of it makes the report. And when an unrecorded choice turns out to be wrong, the number still arrives looking exactly like a number. The axis order accident that dragged 0.873 down to 0.583 was exactly that. An evaluation procedure is not a neutral yardstick held up to a model. It is itself a process that handles data, and processes develop quality defects.
Where the defect hid deserves a closer look. What covered the accident was not a lie but an aggregation. The moment a slice was reduced to one representative value, the fact that its coordinates were wrong got erased along with everything else. Means and maxima are tools for tidying things up and tools for burying defects, both at once. The same thing happens wherever data quality is watched only through overall metrics. If the missing-value rate and the duplicate rate both stay inside their thresholds and the model still fails in one particular range, the packed-together numbers have to be spread back out and counted unit by unit.
The remedy this paper offers is not a smarter metric either. Check the geometry before any comparison begins, count with label-free diagnostics what percentage of cases came out inverted, and change the scoring choices one at a time to see whether the conclusion holds. None of the three touches the model. The same questions can be put to anyone handling training data. What was used to confirm that the labels and the inputs sit in the same coordinate frame? Is the metric on the screen right now a packed one or a spread-out one? Change one preprocessing convention, and does the conclusion flip?
Editor's Note
There is a question Pebblous keeps asking about data quality, and this paper asks it too. What is the number just reported actually measuring? The model's ability, or the conventions laid down in order to measure that ability? When the second one is wrong, the number raises no alarm. It simply comes out a little lower. The axis order accident barely registered in the summary score. That is the sharpest example of the quiet. The overlap, though, is something this article drew out, not something the paper says.
Thanks for reading this far. The full paper is at arXiv:2610.02136, and every figure quoted here was checked against the abstract and the text. If you have ever found a mismatch in an evaluation procedure after the fact, we would be glad to hear what gave it away.
Pebblous Data Communication Team
October 5, 2026
References
- 1.Kafee Hernashki, N. & Chatterjee, S. (2026). "MIRTO: a registration-gated, multiverse-tested evaluation protocol for unsupervised anomaly segmentation in brain MRI." arXiv:2610.02136.
- 2.Steegen, S., Tuerlinckx, F., Gelman, A., & Vanpaemel, W. (2016). "Increasing Transparency Through a Multiverse Analysis." Perspectives on Psychological Science, 11(5), 702–712.