Executive Summary
When a table arrives and a model is graded on it, we usually count one row as one observation. But if rows with identical values appear several times over, the first thing to settle is what that repetition is. It might be a record of the same account transacting normally for three days straight. It might be the mark left by one record copied three times while two tables were joined. A paper posted to arXiv on 27 August put that question to every public dataset in the anomaly detection field. This article looks at what the audit turned up, and at how much the scoreboard moves once the counting unit changes.
An exact-row audit of the 690 datasets in the anomaly detection benchmark collection OddBench found that in 355 of them, more than half, the training set and the test set both contained the very same row. And in 147, two rows that did not differ in a single cell carried the labels normal and anomalous respectively. The author stops short of calling this an error. There are legitimate reasons for a row to repeat, and the released table alone cannot tell you which reason applies — that is the paper's actual claim.
Sections 1 through 4 report what the paper says. Section 5 is this article's own reading of it.
Key Figures
Source: Deng, arXiv:2609.29580, abstract, Table 1, and experimental sections.
355 / 690
Datasets where training and test data overlap
Both sides hold a row whose values match exactly, which is 51.4% of the benchmark
137
Datasets that grade a normal as an anomaly
A row marked anomalous in the test set is identical to a row taught as normal
50–61
Datasets where AUROC moved by more than 0.05
With the model untouched and only the counting unit switched from rows to values, for each of four detectors
30
Datasets where the top detector changed
The same four detectors ran on the same data, and the winner depended on the counting unit
690 Datasets, Compared One Row at a Time
Anomaly detection is the work of picking out what stands apart from the normal crowd. Card fraud, early signs of equipment failure, attempts to break into a server. Researchers in this field score every new method against a collection of public datasets. OddBench is one such collection, holding 690 tabular datasets already divided into training and test portions.
What the author did is simple. He opened all 690 and counted where, and how often, a row appeared whose values did not differ in a single cell. Not approximate matches, not similarity, only exact equality. The result is Table 1 of the paper.
This sits on a different layer from the checks the field has run so far. ADBench, from 2022, compared 30 methods across 57 datasets, and MacrOData, out this year, pushed that scale to 2,446. There has been work cataloguing the traps that splitting procedures set in tabular deep learning, and there are tools that learn a representation and rank the samples most likely to be duplicates or label errors. What grew in all of those was the number of methods compared and the number of datasets. What this paper counts is neither methods nor datasets but the rows inside a dataset. Which is why growing the model side does not reach it, as the author notes up front in the introduction. A detector that swaps only its encoder or prediction head cannot recover missing provenance, entity numbers, time, or exposure, and bringing in a foundation model trained on tables wholesale still maps identical rows to identical scores.
Each line means something slightly different. The 355 are cases where a row from the training set also sits in the test set unchanged. The model meets a question it has already seen. The 147 are cases where rows with completely identical values were split between normal and anomalous, counted across training and test together. Of those, 140 show the same conflict even when you look only inside the test set. And in 137, a row the test grades as anomalous does not differ by a single character from a row the training taught as normal.
That last line is the awkward one because the grading does not hold together. If the model answers normal for that row, it answered as it was taught and gets marked wrong. If it answers anomalous, it gets marked right but answered against what it learned. Neither direction measures ability.
Where rows carry the same values but split labels, no model can tell the two apart. The input is the same, so the output has to be the same, and one side is guaranteed wrong. From here the author computed, for each dataset, the ceiling on the score any model could reach. Twenty-nine datasets cannot touch an AUROC of 0.99 no matter how good the model, 17 are capped at 0.95, and 7 at 0.90. The remaining distance is not a gap in skill but a portion the data deducted in advance.
With Isolation Forest running, the more often this conflict appeared in a dataset, the lower its published AUROC. The correlation is −0.221, not large, yet hard to write off as chance. Overlap showed up on the side that shaves scores rather than inflates them, pointing the same way as the ceiling just described. The author is firm, though, that this is an association and not a causal estimate.
What It Means When the Same Row Appears 500 Times
Read only this far and it sounds like a paper faulting whoever built the benchmark. The author does not go that way. The example the introduction opens with shows what kind of paper this is. Suppose one transaction record sits in the table 500 times. A join may have gone wrong and copied one record 500 times. It may be the trace of an attack that sent the same request over and over. Entity numbers and timestamps may have been stripped, leaving 500 genuinely different records looking identical. Or it may simply be a normal transaction that happened 500 times.
The four demand opposite judgments. If it is a replay attack, the repetition itself is the strongest anomaly signal there is. If it is a join accident, the repetition is noise to be erased. If identifiers were stripped, the 500 really are 500, and collapsing them to one destroys information. Yet all the table records is the pattern of values and a count of the rows carrying it.
| Why the same row appeared 500 times | Then the repetition is | The right handling |
|---|---|---|
| A normal transaction that really happened 500 times | Business frequency | Count it as 500 as it stands |
| An attack sending the same request over and over | The anomaly signal itself | Feed the count into detection |
| A join gone wrong, copying one record | Noise from the data pipeline | Collapse it to one |
| Entity numbers and timestamps stripped away | 500 genuinely different records | Collapsing loses information |
The two right-hand columns differ on all four lines. And nothing inside the table tells you which line the left-hand column is on.
The paper pins this difficulty down as a proposition. How many times a given value is observed on average is written as the product of three things. How long you watched, which is exposure. How often the thing genuinely happens in the business, which is intensity. And by what factor the process that built the data inflated it, which is replication. The trouble is the second and the third. Multiply intensity by any number and divide replication by the same number, and the product is unchanged. Observe as many values as you like and the two shares still cannot be separated, which the author calls the non-identifiability of frequency causes.
Assigning a cause requires information from outside the table. The paper names four. Records of the same object observed repeatedly at several points in time, a provenance history of where and how each row entered, entity identifiers, and interventions that deliberately alter values. Miss any one of them and the meaning of a duplicate stays undetermined. A static table gives no answer.
So the paper's conclusion is not an instruction on how to handle duplicates. It marks out three branches, that multiplicity may be signal, may be nuisance, or may be uninterpretable without further information, and it stops at spelling out the conditions under which each branch holds. This is also why it declines to recommend global deduplication. Wipe duplicates in bulk and the genuine frequency information goes out with them. The author likewise rejects the opposite prescription, writing the duplicate count into a feature column, because the geometry the detector sees gets distorted by exactly that much.
The Same Four Detectors, a Different Winner
A model gets graded whether or not the meaning of a duplicate has been settled. How much that report card wobbles is the paper's second experiment. The AUROC in ordinary use counts one row as one vote. If rows with identical values appear 500 times, that one value casts 500 votes. The alternative the author set against it gives one vote per distinct value. However many times it turned up, the same value is one vote. The paper calls the first row weighting and the second support weighting, and it names the metric computed the second way replication-invariant AUROC, a metric proven to hold its value under any amount of row duplication.
Four classical detectors ran across the 690 datasets with both metrics measured side by side. Isolation Forest, ECOD, HBOS, and COPOD, the ones anomaly detection papers always bring along as baselines. That produced 2,760 detector-dataset pairs. Among them, the cases where nothing but the counting unit changed and AUROC still moved by more than 0.05 numbered between 50 and 61 per detector.
A 0.05 swing is not small. Not one model and not one dataset changed, only what counts as a single vote at grading time, and it moved that far. Look at the averages and nothing appears to have happened. Across the four detectors the two metrics differ on average by 0.0054 to 0.0112. The wobble just described was sitting underneath that average. A signed average does not describe the instability of individual datasets, the paper writes. Nothing was overturned wholesale, of course. The rank correlation between the two metrics is 0.929, so they point broadly the same way. On 64 datasets, though, the ordering of the detectors changed, and on 30 of those the top spot changed hands.
The author takes no side here. If each row really is the intended observational unit, row weighting is answering the right question; if multiplicity crept in while the data was being made, the replication-invariant metric gives the right answer. Which of the two applies is, as section 2 showed, not settled by the table alone. So the paper says to use this metric as a sensitivity check rather than a replacement for row weighting. Where a dataset opens a wide gap between the two values, that is a signal to defer interpreting its score until the counting unit has been declared.
A Detector That Separates Shape From Frequency
The report card is not the only thing that wobbles. A detector trained row by row is skewed in what it learns. An anomaly detection model learns the shape in which normal data is scattered, and if one value sits there 500 times, it takes a distribution thickened 500-fold around that value as the picture of normal. The more often a value was replicated, the more normal it looks, and a value that appeared once gets pushed to the margins. The paper describes this as row-level training learning a law biased in proportion to multiplicity. It is the third consequence, following the evaluation ceiling and the metric sensitivity.
So the back half of the paper turns to the question of how to train instead. SCOUT, the author's proposal, splits one model into two channels. One looks only at the shape of the values. It fits the model once per distinct value, so however many times a row is replicated, this channel's score does not change. The other looks only at how often something appeared. This one, though, switches on only when information such as the observation window or the exposure is supplied alongside. Keeping shape and frequency out of the same space is the point of the design.
In a controlled simulation each channel caught only what it was assigned. On anomalies that stand out by shape the shape channel scored 0.9854, while on anomalies that stand out by frequency alone it scored 0.5012, no better than guessing; the frequency channel did the reverse, at 0.9946 and 0.5351. Isolation Forest, measured alongside, came in at 0.5086 on frequency anomalies and 0.9859 on shape anomalies. The kind of anomaly that arrives as the same request repeating is, in other words, effectively invisible to an ordinary detector.
The condition for switching the frequency channel on is where the design's honesty lies. Exposure means how much opportunity there was for the value to appear. The months an account was watched, the hours a machine ran. Ten occurrences out of a single day and ten out of a full year are different stories even at the same count. Without that information there is no basis for interpreting a count, so the channel simply stays off. It is the non-identifiability of section 2 transposed into model architecture.
Both channels pass through a calibration step that guarantees a bound on the false alarm rate. It works by holding out normal data unused in training, looking at how the scores spread there, and setting the alarm threshold from that, which holds without assuming an infinite supply of samples. Running it on 100 datasets across 5 seeds for 691,385 scored tests, the actual rate tracked almost exactly the false alarm rate set at 1%, 5%, and 10%. This is the first number anyone asks about when putting an anomaly detection tool into the field.
The question that matters is whether counting this way costs performance. On 686 datasets across 5 seeds, with four dropped for holding fewer than 40 distinct normal values, SCOUT using only the shape channel went up against a row-trained Isolation Forest: ordinary AUROC came out 0.0026 lower for SCOUT. The confidence interval straddles zero by enough to be judged non-inferior, and measured with the replication-invariant metric SCOUT was ahead instead, by 0.0049. Time per fit is much the same, 0.387 seconds against 0.333.
The difference surfaces when replication is deliberately applied. Copying some normal rows artificially several times over and measuring how far the score ordering shifts, the row-trained Isolation Forest had a rank correlation with a median of 0.9825 and a worst case of 0.8883. Six datasets fell below 0.95. Under the same intervention the ordering from SCOUT's shape channel was exactly as before. Not approximately alike but identical to the limits of numerical precision.
On real data which side comes out ahead varied by table. In the 172 datasets in the top quartile by duplication rate in the training data, SCOUT's ordinary AUROC was 0.0104 lower and its replication-invariant metric 0.0198 higher, opening a clear gap between the two approaches. In the remaining three quarters that gap stayed within 0.0006. Where there are few duplicates, it hardly matters what you count as one, and this discussion turns urgent in tables where identical values have piled up thick.
To the question of whether it would not be better to keep both channels on at all times, the paper answers no with its own experiment. On anomalies that stand out by shape alone, the combined score came in 0.0146 below the shape channel used by itself. Splitting the test in two costs that much, so the default was left at the single channel. What the frequency channel adds is also conditional. Where the rate of occurrence varies sharply with the features, it climbs to between 0.057 and 0.083, but where that dependence is weak it comes to between 0.003 and 0.006. All six comparisons were statistically significant; what settled their usefulness was the size of the gap, not the p-value.
The paper puts three limits on itself. Both the audit and the detector handle only cases of exact equality. A one-character typo, a rounding position, the same entity differing in a few cells — none of these are caught this way. The guarantee the calibration provides rests on the assumption that normal data behaves the same whatever order it arrives in, so it breaks if the behaviour drifts over time or if anomalies contaminate the calibration data. The frequency channel's usefulness was confirmed only in semi-synthetic experiments, because OddBench carries no exposure information at all.
An extension swapping the backbone for a Deep Isolation Forest lowered AUROC by 0.0036 across 100 datasets, and the confidence interval failed to clear the non-inferiority bar, so the result was left on the record rather than presented as a successful extension. Another detector that looked promising in simulation reached 0.679 on 20 real datasets, short of the 0.705 from an Isolation Forest built the same way.
Why Pebblous Is Watching This Paper
From here on this article rereads the paper through the lens of data quality.
When we say we are inspecting data quality, what we mostly look at is values. Are there blanks, do the digits line up, is anything outside its range. What this paper points at is the layer below that. The values are all perfectly fine, and nowhere is it written whether that one row stands for a single event or for many. With a gap on this layer, every value check passes. And with it passed, the model gets trained and the score gets assigned.
Put the non-identifiability proposition into working language and it goes like this. Once data has been made, no amount of staring at it will produce the answer. The answer was back where it was being made. Which join was applied, which column was dropped for privacy reasons, how the aggregation window was drawn — if that is kept on the record, the meaning of a duplicate is settled, and if it is not, it can never be recovered. Keeping a provenance history and entity identifiers turns out to be not paperwork but the only key that can be used later.
And this paper does not recommend the deduplicate button. That is the first prescription anyone working on data reaches for, and the moment identical rows are erased the genuine frequency information goes with them. What has to be settled first is what counts as one, and only after that declaration does keeping or deleting follow. If deduplication sits in your toolkit as basic data hygiene, one field needs to go in front of it, recording the observational unit.
In the experiments here not a single model changed, and the top spot on 30 datasets changed hands. If two candidates on an internal model comparison sheet sit within 0.05 of each other, there is nothing on hand to tell whether the difference came from the model or from repetition inside the table. Scoring the same data twice, once by row and once by value, and seeing how far the two values diverge is a day's work. Which tables make that check worth running was already shown in section 4. Where duplicates are few, the result came out nearly the same whichever unit did the counting. So what to measure first is not the score but the duplication rate in your own table.
The observational unit lives in the head of whoever built the data, and it disappears when that person moves teams. What single thing does one row of this table stand for, if the same value appears in several rows how many events is that, which tables were joined to build this one. Three sentences are enough, and they need to sit somewhere that travels with the data, whether beside the schema or in a data description document. Whether the person scoring a model on this table later can interpret their own number depends on that.
The integrity of evaluation sets is ground Pebblous has covered before. We took up benchmark contamination in large language models, where the exam questions were already inside the training data, and we looked at a preterm birth prediction benchmark whose performance had been inflated because records from one mother were split across training and test. This story sits a layer below both. Nothing leaked out of the exam, and no patient was split across the divide. It sits somewhere that fixing the splitting procedure does not reach. Nobody wrote down what unit a single row counts.
So there is one question that opens your own table. How many pairs of rows with completely identical values sit across our training data and our evaluation data? If nobody has counted, that is worth counting, and if the count comes back positive, the next question is whether that repetition is a number of events that really occurred on the books or a mark left while the data was being moved. Without records that can answer the second question, a score assigned on that data is still short of interpretation.
Thank you for reading this far. The paper this article follows is available in full and free on arXiv. We would be glad if you counted how many pairs of exactly identical rows sit in the evaluation data your organization uses, and told us what you found.
References
Primary Source
- 1.Deng, J. (2026). "When Identical Rows Disagree: From Benchmark Identifiability to Replication-Robust Anomaly Detection." arXiv preprint.
Compared Benchmarks
- 2.Han, S., Hu, X., Huang, H., Jiang, M. & Zhao, Y. (2022). "ADBench: Anomaly Detection Benchmark." Advances in Neural Information Processing Systems, Vol. 35.
- 3.Ding, X., Klüttermann, S., Wen, H., Chen, Y. & Akoglu, L. (2026). "MacrOData: New Benchmarks of Thousands of Datasets for Tabular Outlier Detection." Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining.