Executive Summary
When you go looking for a bump in data without knowing in advance where it sits, you end up choosing its position from the data as well. Two things happen the moment you choose. If a genuine signal is there, its strength is measured larger than it really is. If no signal is there, one background fluctuation starts to look like a signal. A paper Tommaso Dorigo of Italy's National Institute for Nuclear Physics (INFN) posted to arXiv on September 3 places the two side by side in a Gaussian matched-filter model and shows that the same local geometry produces both.
The bookkeeping turns on a single number, the dimension of the freedom you took. When D coordinates such as position or width are chosen from the data, the fitted signal strength is pushed upward by D/(2Q) on average. Q is the signal-to-noise ratio you would have with the position held fixed, so the weaker the signal, the bigger the bias. The same D turns up when no signal is present at all. The trials factor you pay when converting a local significance into a global one grows as the D-th power of the threshold u. The paper's result is that the two quantities are joined exactly by a single logarithmic derivative, and the author himself writes in his conclusions that the two problems are not identical phenomena but complementary limits of the same profiling and random-field geometry.
Sections 1 through 3 follow what the paper actually measured and what it held back. Section 4, which carries the result over to dashboards and threshold tuning, is this article's reading rather than the paper's.
Key Numbers
Sources: Dorigo, The Greedy Bump Bias: Local Profiling Geometry and the Look-Elsewhere Effect, arXiv:2609.03581 (2026), Secs. 3, 8 and 9; ATLAS (2016) abstract
D/(2Q)
How much a genuine signal inflates
1/(2Q) for each coordinate chosen from the data, and larger the weaker the signal
3.9σ → 2.1σ
ATLAS's 750 GeV excess of 2015
Counting the chance of seeing it at other masses too pulled the significance down this far
0.53 → 1.05
From local bias to global bias
At a weak signal of Q = 1, once the search half-range is widened to 14σ
Q ≈ 5 to 6
Where the search range stops mattering
Above this the bias is the same no matter how wide a range you scanned
Choose where to fit and the bump rises
Finding a new particle means finding a bump somewhere in a mass distribution. The trouble is that nobody knows which mass it sits at. So the standard move is to let the signal amplitude and the position float together and take whichever combination maximizes the likelihood. Shape parameters such as the width usually go into the fit as well. This operation is what the paper calls profiling.
The intuition is simple. Even when the signal model is correct, its position and its amplitude are estimated from the same noisy data. Letting the position float lets the fitted template adapt to that day's noise realization. It may move toward a positive fluctuation or away from a negative one. The displacement itself has no preferred direction, but which way it goes is settled by whatever improves the fit. So the change in the optimized peak height is positive on average.
The paper splits the fitted amplitude into two pieces. One is the ordinary fluctuation you would get with the template pinned to the true position, and that piece has mean zero. The other is the gain obtained by moving the template across the manifold. The entire bias comes from the second piece. Measure with the position held fixed and the bias never arises at all.
This story starts twenty years ago. In an unpublished internal CDF study carried out in 2003, the author observed that the fitted yield of a Gaussian signal was biased upward when its position was allowed to float, and that the size of the effect fell as the signal-to-noise ratio rose. A 2009 blog post gave the effect its name, the greedy bump bias. That original study established the effect numerically, the paper notes, but gave no analytic account of the mechanism behind it.
1.1What the simulations confirmed
With D coordinates free to be chosen, the likelihood gain from profiling converges to a chi-square distribution with D degrees of freedom, whose mean is D. The bias in the fitted amplitude is D/(2Q). The author checked both predictions with ROOT pseudoexperiments. The one-dimensional model, in which only the position floats, was run 200,000 times at each signal strength; the two-dimensional model, which floats position and width together, was run a million times. Multiplied by Q, the bias converges to 1/2 at D = 1 and to 1 at D = 2.
Even the rate of convergence lines up. Fitting the leading correction that survives at finite Q in the two-dimensional model gives 2.263 ± 0.159 against an analytic value of 9/4, that is 2.25. The bias in the width coordinate and the correction to the likelihood gain matched their predicted values of 1 and 4.5 as well. The author attaches one caveat, though. The quantities that measure how sharply the position itself is pinned down converge more slowly, and over the accessible range of Q the higher-order terms in those remain heavy.
The same dimension sets the odds of a bump that isn't there
Now move to the case where there is no signal at all. Background alone still has a highest fluctuation somewhere, and measuring at that spot produces a rather convincing bump. Particle physics calls this the look-elsewhere effect. The name says that when you assess an excess seen at one mass, you have to count the chance that a similar fluctuation could have shown up at some other mass. This is where the correction that discounts a local significance into a global one comes from.
Just how large that correction can be became widely known at the end of 2015. ATLAS and CMS both reported an excess near 750 GeV in the invariant mass distribution of photon pairs. ATLAS quoted a local significance of 3.9σ in a search assuming spin 0, and CMS quoted 3.4σ. The global significances the two experiments wrote into the same papers were 2.1σ and 1.6σ. Theory papers followed in a flood, and then the bump vanished in the 2016 data. This episode does not appear in Dorigo's paper; we bring it in to give a sense of how big the correction gets.
Under the practical approximation Gross and Vitells published in 2010, the trials factor in a one-dimensional search grows roughly in proportion to the local significance. In D dimensions the same structure becomes the D-th power of the threshold u. The constant out front absorbs the size and geometry of the search region. So the trials factor splits into two shares. The constant effectively carries how many independent regions were swept, while u to the D carries what was gained by continuously optimizing position and shape inside each of those regions. This split between global multiplicity and local optimization, the paper notes, is already present in the Gross and Vitells analysis.
That is why the same integer D shows up in three different places.
| Where it appears | Form | What it means |
|---|---|---|
| Trials factor under background only | TF0(u) ∝ uD | The higher the threshold, the more the multiplier on your local significance grows, as the D-th power |
| Profile likelihood gain | E[Δq] → D | The mean gain obtained purely by choosing position and shape |
| Bias in the fitted signal strength | 2Q · E[Q̂ − Q] → D | How much a genuine signal is pushed upward |
Q is the signal-to-noise ratio with the position held fixed, Q̂ is the value measured with the position chosen too, and Δq is the likelihood-ratio gain from profiling. Source: arXiv:2609.03581, Eq. (9).
The paper's first contribution comes out of this. Take the logarithm of the trials factor and differentiate with respect to the threshold, and the constant out front disappears. How wide the search region was gets erased, and only the local optimization share remains. Evaluate that quantity at u = Q and you get exactly the bias in the fitted amplitude. Half the relative rate at which the trials factor grows with threshold is how much a genuine signal inflates.
EQ[Q̂ − Q] = ½ · d(log TF0(u))/du | u = Q
The result the paper puts in a box (Eq. 158). The right side is the rate at which the trials factor grows with threshold, the left side is how much a genuine signal inflates.
You could ask whether the matching dimension is a coincidence. The paper's answer is that both quantities come out of the same matrix. When a genuine signal is present, the matched-filter peak narrows in the tangent directions in proportion to Q. Condition a background-only field on a rare peak of height u and the same directions return the same curvature. The difference is that the signal-side calculation takes the inverse of that curvature and reads off its trace, while the background-side calculation takes its determinant. Only the use differs; the metric they are built from is a single one.
The phrase "same metric" can be turned into a picture. The stronger the signal, the narrower the patch of positions where the template could plausibly sit. The covariance of the fitted coordinates goes as the inverse square of Q, so the volume of that patch shrinks as the inverse D-th power of Q. Take the logarithm of that volume, differentiate with respect to Q, halve it and flip the sign, and out comes exactly D/(2Q). And the cell that narrows this way is precisely the cell whose shrinking multiplies the number of distinguishable search positions under background alone by Q to the D. One geometry, local resolution, seen from the signal side is an upward push and seen from the background side is a growing count of places to look.
There is a condition attached to the two sides matching this neatly. In a general non-centered Gaussian field, the author notes separately, the two agree only when the rate at which the signal curvature changes with amplitude equals the covariance of the noise gradient. In the normalized matched-filter problem the curvature is QG and the covariance is G, so the condition holds automatically. The identity is a property of matched-likelihood geometry, not something that attaches to every bump-hunting procedure.
The author himself drew the line around what is new. Positive optimization bias, the chi-square likelihood gain with D degrees of freedom, the coefficient D/(2Q), and the u-to-the-D form of the trials factor are each, taken in isolation, not claimed as novel. Siegmund and Worsley had both calculations inside one 1995 paper, and in astronomy the same bias has been treated under the name optimization bias. The South Pole Telescope's galaxy cluster analysis optimized two position coordinates and one angular scale, then applied the corresponding 3/(2Q) correction empirically.
The closest prior work by the author's own reckoning is a December 2025 paper by Alden Green and Jonathan Taylor. Their calculation for inferring the location and height of a selected peak already contains the signal-side peak-height correction in the form of half the trace of the inverse curvature times the noise-gradient covariance. It handles the background side too. Conditioning on a peak of height v makes the curvature grow in proportion to v, so the determinant that enters the density of local maxima produces v to the D. The two peak geometries were, in effect, already lying side by side. What is not there, the author writes, is the step that connects that relation to the global trials factor in the Gross and Vitells sense.
So there are two new things. One is the identity obtained by differentiating the global trials factor. The other is the treatment beyond the Poisson limit, which the next section takes up.
When the signal weakens, the search range steps in
The arithmetic of the previous section holds when the signal is strong enough. You cannot recover the background-only case from it by sending Q to zero, because once the signal disappears there is no true position left to expand around. The paper calls that spot the apex of the cone. It reaches instead for a different variable, the competition between peaks.
In the one-dimensional model, climbing uphill from the point where the signal was injected lands on the maximum associated with the signal. The highest of the remaining maxima inside the search interval is counted separately. The global maximum is exactly the larger of those two. The global bias then splits exactly into the local profiling contribution attached to the signal and the contribution from a remote peak winning. This decomposition assumes neither independence of the maxima nor a Poisson process.
Once split, it becomes visible where the search range does its work. At a weak signal of Q = 1, the local bias attached to the signal is 0.527. Widen the search half-range to 6σ, 10σ and 14σ and the global bias climbs to 0.777, 0.939 and 1.053. The local bias has not moved; only the share from remote peaks winning has grown. The probability that a remote peak beats the signal runs 0.354, 0.498 and 0.582 in the same order.
The decomposition itself is exact and assumption-free, but an approximation does enter where the remote peak is swapped for a background unconnected to the signal. The author measured that approximation three ways and compared them: keeping the event-by-event pairing intact, shuffling only the pairing between the signal peak and the remote peak, and drawing the remote peak from a separate background-only simulation. At a half-range of 10σ with Q = 1, the three come out at 0.9392, 0.9472 and 0.9278. The value with the pairing intact is slightly lower, so event-by-event dependence pulls the global bias down a little, and remote peaks drawn from signal-containing toys are slightly more competitive than those from background-only toys. The two discrepancies have opposite signs and largely cancel.
As the signal strengthens, this share becomes exponentially rare. By the time Q reaches roughly 5 to 6, the global and local biases are indistinguishable on the scale of the simulation. How wide a range you swept does not affect the bias of a strong signal at leading order, and conversely, for a weak signal the search volume enters the value directly. This is why the paper keeps three layers apart. The dimension D fixes the universal local effect, the curvature of the manifold fixes the model-dependent correction at the next order, and the search volume, the boundaries and the correlations among maxima enter once the competition begins.
In the range where the threshold is not arbitrarily high, the approximation that treats background maxima as a Poisson process also wobbles. The paper organizes that departure with factorial cumulants. At a half-range of 10σ and a threshold of 1.0, the probability of finding no competing peak measured over 500,000 pseudoexperiments is 0.25336, where the Poisson approximation gives 0.26827. Including the second cumulant gives 0.25223 and including the third gives 0.25349, which effectively coincides with the simulated value. The sign of the second cumulant flips with threshold, though, so the author writes that calling this correlation repulsion is less accurate than calling it a non-Poisson correlation.
Once the boundaries and the labelling conventions are stripped away, what remains is the clean interior-maxima process, and there the story sharpens. A two-point Kac-Rice calculation yields the result that the density of two maxima appearing very close together is suppressed as the fourth power of their separation. The reason is that in a smooth field two nearby peaks must be nearly equal in height, must both have vanishing gradients and must turn direction in between, which cuts the permitted configurations sharply. The second factorial cumulant from this calculation was checked against 500,000 pseudoexperiments counting interior maxima only, and it matched at the percent level or better at all four thresholds, with every difference falling inside twice the simulation standard error.
Dashboards pay the same price
That is where the paper ends. The author was explicit about how far the result reaches. The calculations were all done on a Gaussian noise model, and that was a choice made to expose the geometry cleanly. Real resonance searches are more often built on Poisson or extended likelihoods. The conclusions call it an important next step to establish which of the tangent-space results, the local bias and the finite-threshold corrections survive under those conditions.
The author flagged one more limit, on the residual dependence seen in Section 3. That dependence is small in the examples studied, but in the rare reversals that happen at high signal strength there remains a tendency for the winning competitor to cluster near the signal. That clustering shows up more the shorter the search interval, and the author names it as the reason an independent background model does not fully account for remote-win rates at high Q. Quantifying that share properly requires a two-point calculation for a non-centered field. The third factorial cumulant is still an empirical number pulled from simulation, so getting its analytic counterpart means going to a three-point calculation. The acknowledgments state that OpenAI's ChatGPT (GPT-5.6 Sol) was used in developing and checking the analytic derivations, the numerical studies and code, the literature analysis and the preparation of the manuscript, and that the scientific direction, the interpretation and the final responsibility remain with the author.
Step outside the paper and this structure is not a physics matter. Any procedure that finds something in data without having fixed the position or the shape beforehand pays the same price. Scrolling a dashboard and clipping out the most striking interval to report is one. The start and the end of that interval were set by looking at the data, so the dimension is at least two. Sliding a threshold until the anomalies come out cleanest is another. Running an A/B test and picking, after the fact, whichever of several metrics came out best is a third. All three choose where the signal sits by consulting the data, and a finding chosen that way is bigger than one that was not.
This correspondence is not in the paper and is our reading. Even so, three questions are worth working through in the same order.
- How many things in this finding did we set by looking at the data? Count the interval boundaries, the threshold, the metric, the segment and the time window, and that count is the dimension.
- Did the report say what that count was? In the paper's arithmetic the constant out front drops out only under differentiation, and the reader still needs to know how wide a range was swept.
- Is the signal on the weak side or the strong side? The weaker it is, the larger the bias and the more the search range enters the value. The more marginal an anomaly looks, the more it deserves suspicion.
As to whether that count can go straight into D, the paper attaches a caveat in advance. A quantity that cannot be negative, such as a width or an amplitude, truncates its tangent direction, and two variables that do nearly the same job cut the number of effectively independent directions by that much. So the nominal D may have to be replaced by an effective dimension that reflects boundaries and degeneracy, the author writes among the open directions. The count you take is closer to an upper bound than an answer.
Fail to count the degrees of freedom in your search and the finding will always be a little too big. How much too big is countable, and this paper wrote down in a single line where that number comes from.
Editor's Note
There is a scene Pebblous runs into often while diagnosing data quality. Someone finds the stretch where one metric spikes unusually and makes the next decision on the strength of it. If the record does not say who picked that stretch and on what basis, reopening the same data gives you no way to work out how big the finding really was. Keeping a record of how something was found, as much as what was found, is in the end what decides what the data is worth.
Thanks for reading this far. The paper is at arXiv:2609.03581, the numbers in this article were checked against Secs. 3, 6, 8, 9, 10 and 11, and the 750 GeV significances were taken directly from the abstracts of the 2016 ATLAS and CMS papers. If your team already has a way of recording the degrees of freedom in a search, we would like to hear which fields you keep.
References
Primary Source
- 1.Dorigo, T. (2026). "The Greedy Bump Bias: Local Profiling Geometry and the Look-Elsewhere Effect." arXiv:2609.03581.
Prior Work
- 2.Siegmund, D. O., Worsley, K. J. (1995). "Testing for a Signal with Unknown Location and Scale in a Stationary Gaussian Random Field." The Annals of Statistics 23(2), 608–639. doi.org/10.1214/aos/1176324539 — Sec. 2.2, the earlier case that already held both calculations in one paper.
- 3.Gross, E., Vitells, O. (2010). "Trial Factors for the Look Elsewhere Effect in High Energy Physics." The European Physical Journal C 70, 525–530. doi.org/10.1140/epjc/s10052-010-1470-8 — Sec. 2, the high-threshold approximation to the trials factor and the split between global multiplicity and local optimization.
- 4.Green, A., Taylor, J. (2025). "Inference for Location and Height of Peaks of a Standardized Field after Selection." arXiv:2512.04059 — Sec. 2.4, the Kac-Rice calculation the author names as the closest prior work.
- 5.Vitells, O., Gross, E. (2011). "Estimating the Significance of a Signal in a Multi-Dimensional Search." Astroparticle Physics 35, 230–234. doi.org/10.1016/j.astropartphys.2011.08.005 — Sec. 2, the extension to multidimensional searches.
- 6.Zubeldia, Í., Rotti, A., Chluba, J., Battye, R. (2021). "Understanding Matched Filters for Precision Cosmology." MNRAS 507(4), 4852–4863. doi.org/10.1093/mnras/stab2461 — Sec. 2, the astronomy work that systematized the same bias under the name optimization bias.
- 7.Davies, R. B. (1987). "Hypothesis Testing When a Nuisance Parameter Is Present Only Under the Alternative." Biometrika 74(1), 33–43. doi.org/10.1093/biomet/74.1.33 — Sec. 2, the statistics literature that first formulated the problem.
- 8.Dorigo, T. (2009). "Bump Hunting II: The Greedy Bump Bias." Science 2.0 blog, April 25, 2009 — Sec. 1, where the name was first attached.
Experimental Reports
- 9.ATLAS Collaboration (2016). "Search for Resonances in Diphoton Events at √s = 13 TeV with the ATLAS Detector." JHEP 2016(9), 001. arXiv:1606.03833 — Sec. 2, the source of the local 3.9σ and global 2.1σ.
- 10.CMS Collaboration (2016). "Search for Resonant Production of High-Mass Photon Pairs in Proton-Proton Collisions at √s = 8 and 13 TeV." arXiv:1606.04093 — Sec. 2, the source of the local 3.4σ and global 1.6σ.