Executive Summary
Observation alone does not decide what a telescope discovers. The weakest signal you agree to accept as a candidate decides it too. PLATO, the European Space Agency's next transit survey, has not launched yet, and a paper posted to arXiv on September 4 has already priced what moving that line costs. Its author is a single independent researcher working outside the PLATO consortium.
Moving the line from 6 to 8 takes the share of planet-free light curves that still return a candidate from 79% down to 0.6%. Over the same range, the share of injected planets the pipeline hands back falls from 61% to 51%. Fewer phantoms means more missed planets, and this paper's contribution is the exchange rate for that trade, written out in numbers while there is still time to choose.
Sections 1 through 4 follow what the paper measured and the reservations its author states himself. Section 5, which carries the result over to data pipelines, is this article's interpretation.
Key Numbers
Source: Yohann Tschudi, Habitable-zone Earths at the detection frontier, arXiv:2609.04887 (2026), Sects. 3 and 4 and Fig. 2
79% → 0.6%
False-alarm rate as the floor moves from 6 to 8
Share of the 465 planet-free light curves that returned a candidate
61% → 51%
Completeness over the same range
The share of injected planets recovered comes down with it
6.0% / 54%
The two values at the adopted floor of 7.5
False-alarm rate and completeness at the point chosen for discovery
4.0
Habitable-zone Earths expected as candidates
Out of about 19 statistically present inside the footprint (+4.2/−1.8)
What Happened When the Floor Went Up by Two
A transit search looks through a record of a star's brightness for the moments when it dims very slightly, over and over on a fixed period. A planet crossing in front of its star blocks a little light and leaves a shallow notch in the light curve. The trouble is that starspots, stellar oscillations and leftover instrument error carve notches of their own. So search software scores every signal it finds and promotes it to the candidate list only when that score clears a line set in advance. The paper calls that line the acceptance floor, and the statistic it applies to accounts not only for white noise but for the red noise that stays correlated across time.
With the floor at 6, 79±2% of the 465 light curves that had no planet injected return at least one candidate. That 79% is not the share of the candidate list that is fake. It is the share of stars with nothing there at all that produced a phantom anyway. At a floor of 8 the count drops to 3 of 465, or 0.6%. Noise-made signals are largely confined below a score of about 7.5 to 8, while real planets extend above it. Over the same range, though, completeness, the share of injected planets the pipeline actually recovers, falls from 61±2% to 51±2%.
The operating floor the author adopts for discovery is 7.5. At that point the false-alarm rate is 6.0±1.1% and completeness is 54±2%. For work that needs statistical purity he recommends 8 instead. The three operating points sit side by side below.
The author reports the two numbers together for a reason. On correlated noise, part of the completeness that a low floor appears to buy is itself made of false alarms. The 61% at a floor of 6 therefore contains a portion where the pipeline counted a noise signal as a real recovery. A pipeline that advertises only its completeness never says how many phantoms it waved through to get there.
A second test behind the floor brings the numbers down further. This chain checks, event by event, whether a transit really does come back on a steady clock, and counting only the candidates that earn a coherent-clock verdict leaves a false-alarm rate of 9 of 465, or 1.9%. With the setting that lets the test abstain rather than rule, it is 7 of 465, or 1.5%.
He also checked separately that this abstention setting does not move his results. It changed 29 of 435 verdicts, all of them at periods of 10.6 days or less, and none at all in the habitable-zone band. Recovery in that band is 37 of 75 whether the setting is on or off. The habitable-zone forecast further down therefore does not depend on how it is configured.
The 465 Twins With No Planet in Them
PLATO is the exoplanet survey satellite the European Space Agency is preparing now. It will hold its first long-pointing field, LOPS2, in view for at least two years, which brings the habitable zones of its M-dwarf sample within reach of a transit search for the first time. Around these dim stars, the orbits where water stays liquid have periods between 19 and 112 days. Planets injected into that band collected between 7 and 37 transits over the 717-day baseline, with a median of 18.
Forecasts of how many planets PLATO will find already exist along three lines. Heller et al. injected Earth-like planets around Sun-like stars and recovered them in 2022; Matuszewski et al. computed mission-wide statistical yields for the FGK samples in 2023; Schlecker et al. ran demographic survey simulations reaching the M dwarfs in 2024. Instead of adding a fourth forecast, this paper produces the raw material for one: the two quantities that only injection and recovery through a real search-and-screening pipeline can supply, which are the fraction of injected planets the pipeline returns as candidates and the rate at which the same pipeline returns a candidate from a light curve with no planet in it.
Rather than observe these stars, Tschudi generated them. He produced 717 days of photometry at a 25-second cadence with the PLATO Solar-like Light-curve Simulator, then binned it to 300 seconds. Binning shrinks a light curve from 2.5 million points to 200,000 and divides the search cost by an order of magnitude, at a cost in transit signal-to-noise below 1%. Three kinds of star were chosen to represent the faint half of the sample.
| Cell | Host and cameras | Noise (per 300 s bin) | What this cell measures |
|---|---|---|---|
| M4 bright | M4, V=14.5, 12 cameras | 1,308 ppm | The completeness transition at the low-noise anchor; HZ band 19–52 d |
| M2 | M2, V=15.0, 12 cameras | 2,076 ppm | HZ band 43–112 d, the long-period extreme |
| M4 faint | M4, V=16.0, 6 cameras | 6,200 ppm | The hard end, where only six cameras cover the star |
The three stellar cells of the campaign, taken from Table 1 of the paper, single-planet cells only. Three further cells imitate multiplanet systems. The median magnitude of the field M dwarfs is V=15.5, so these cells sit in the faint half of the sample. Source: arXiv:2609.04887.
The design of the campaign starts by building twins. Each of the 465 light curves carrying an injected planet has a partner built from the same star and the same noise class with no planet at all. The injections number 525 planets in total, and the 465 twins carry the entire false-alarm measurement. Because that data is known to contain no answer, anything the pipeline finds there is by definition a phantom.
The campaign runs in two arms. One assumes the instrument has been corrected perfectly and injects only stellar and photon noise; the other adds a pre-launch model of the residuals left after correction. The two arms share every input except those residuals, so the difference between them is the price of the instrument. The first arm ran 930 light curves and the second 790, for 1,720 in all. With the residuals in place, the point where completeness reaches half moves from a signal-to-noise of 7.72 to 8.26. The loss sits on the transition around a signal-to-noise of 10, peaking at about 14 percentage points of recovery, and disappears where detection is easy.
The penalty the residuals impose differs in size from star to star: +1.5 in signal-to-noise on the quietest cell, +0.6 on the intermediate one, and indistinguishable from zero on the noisiest. The contrast between quietest and noisiest is itself significant at 3.0σ. The author is explicit that this gradient should not be read as a prediction of the per-star penalty. The residual template was transplanted at fixed amplitude from the single configuration for which official tables exist, and those tables stop at magnitude 13, short of the campaign's cells at V=14.5 to 16. On faint stars, he adds, the real residuals worsen with charge-transfer inefficiency and would likely reverse the near-zero point.
The floor itself survives the residuals intact. At the adopted 7.5, the systematics arm actually shows a lower false-alarm rate of 3.5±0.9% at 46±2% completeness, because the residuals push noise candidates below the line. The habitable-zone yield does not change either, since the long-period injections that feed it lie in the low-penalty regime. In the other direction, the +0.54 shift should not be read as an upper bound on what the instrument will cost. This pipeline has no dedicated systematics-correction stage at all, which overstates the cost, while treating the residuals as independent between cameras understates it.
This pipeline has its roots in TESS. Tschudi carried over intact the chain he had developed and validated on M dwarfs there, and the scale of that validation is in the paper: applied uniformly to the 461 M-dwarf hosts of TESS objects of interest listed by ExoFOP, it re-detected 165 of the 193 confirmed planets inside its period range, misclassifying none of them as a false positive.
The chain removes stellar variability with Gaussian-process detrending, pulls out signals with an iterative transit-least-squares search, resolves harmonic aliases, and finally tests whether the event times are coherent. One of the settings changed for PLATO is a veto that rejects candidates within 0.5% of 90.004 days and its low harmonics. The satellite rotates its attitude each quarter and leaves regular structure at that period, an instrumental clock rather than a transit, and the veto carries a price of its own. It caught three injected planets, one of them at a signal-to-noise of 11.8 and a period of 29.90 days that the chain would otherwise have detected. The comb crosses the habitable-zone band of the M2 cell, leaving three narrow blind periods.
Kepler Measured Completeness and Reliability After the Mission Ended
Fixing a threshold before launch is nothing new. Kepler did exactly that. Its mission requirement was at least three valid transit events with a multiple event statistic of 7.1σ or above before a signal could be promoted. The difficulty was that the line was frozen before anyone had seen the real data, and the pre-launch noise model behind it turned out to be wrong.
As the DR25 candidate catalogue paper puts it, the typical noise level for 12th-magnitude solar-type stars is closer to 30 ppm than the 20 ppm expected before launch. That gap forced Kepler to need a longer baseline to find a meaningful number of Earth-like planets, and because different stars carry different noise, the depth to which the search was sensitive varied from star to star.
The consequence shows up in the final search. The paper reporting that run, which covered all 17 quarters, took in 198,709 stellar targets and reported signals on 17,230 of them, and re-searching the same stars added enough to reach 34,032 detections in total. Against a set of 3,402 well-established Kepler objects of interest, the recovery rate was 99.8%. Its abstract attaches one sentence to that achievement: "The high recovery rate must be weighed against a large number of false-alarm detections." The choice behind it is stated in the same place, which records "a strategic emphasis on completeness over reliability for the final Kepler transit search."
Completeness and reliability were actually measured only after the mission was over. The DR25 candidate catalogue injected artificial transits into light curves to measure completeness, and simulated false alarms by inverting and scrambling light curves to measure reliability, the share of candidates that are not the work of noise. Below 100 days, completeness came out above 85% and reliability above 98%. In the low signal-to-noise corner between 200 and 500 days around FGK dwarfs, however, completeness was 76.7% and reliability 50.5%. One candidate in two there was noise rather than a planet. The same catalogue holds 4,034 planet candidates, of which the authors counted "ten high-reliability, terrestrial-size, habitable zone candidates." That is what remains after four years of photometry have been searched and the reliability of that corner has been measured.
How far a reliability assumption travels is on record in earlier work the same catalogue paper cites. Summarising Burke et al. 2015, it reports that the occurrence of small, Earth-like-period planets around G dwarf stars changed by a factor of about 10 depending on the reliability of a few planet candidates. A threshold does not merely set the length of a list. It changes the universe you count with that list.
None of this says Kepler picked the wrong line. Kepler had to start before it knew what its own data would look like, and the long tail of correlated noise only revealed itself once the data had accumulated. The new paper differs in one respect, which is that it moves the measurement of that pair ahead of launch. It is a simulated measurement, of course, so whatever the simulator fails to capture is missing from the numbers as well.
So How Many Will Show Up as Candidates?
The completeness curve splits depending on which planets you measure it over. Measured across all injected planets, recovery reaches half at an expected signal-to-noise of 7.72 (+0.15/−0.14), and above 9 the pipeline catches 165 of 176, above 10 it catches 124 of 129. Restrict the measurement to planets inside the conservative habitable zone of their host cell, though, and the half-way point moves out to 8.61 (+0.33/−0.28). That is 0.9 harder in signal-to-noise, and a permutation test that reassigns the habitable-zone labels at random and refits the shift returns p=0.012.
The reason is not that these planets are in the habitable zone. Longer periods give longer transits, and long notches are eroded more heavily by Gaussian-process detrending and correlated noise. The evidence is that planets outside the habitable zone with periods of 19 days or more shift by the same amount. The forecast below therefore uses this long-period curve rather than the overall completeness curve.
Laying the measured curve over the real stars of the LOPS2 field produces the forecast. Start with what is already known. Inside the field there are 36 M-dwarf hosts carrying 50 transiting TESS objects of interest, 19 of them confirmed planets and 31 still candidates in the ExoFOP snapshot of 29 June 2026. Their recovery probabilities are all effectively unity; a median predicted signal-to-noise of 96 puts them far above the floor.
For the planets that have yet to be found, the situation is entirely different. The calculation uses the 35 of those 36 hosts that have a catalogued stellar radius. If a planet exactly one Earth radius across were orbiting one of these stars inside the conservative habitable zone with its orbital plane happening to face us, the pipeline would return it as a candidate with a probability of 41±1%. For a super-Earth of 1.5 Earth radii the figure rises to 65±1%. The bright end of the field is recovered at either size, and the faintest hosts stay out of reach wherever the floor is set.
Across the whole footprint, the same calculation becomes a yield forecast. The PLATO input catalogue holds 16,929 M dwarfs inside the field, and applying known occurrence rates puts about 19 transiting habitable-zone Earths among them, or between 10 and 38 across the occurrence-rate interval. The number this pipeline would be expected to pull out as candidates is 4.0, with a 68% interval of +4.2/−1.8, plus 6.6 habitable-zone super-Earths. For comparison, the PLATO consortium's own performance assessment quotes ten planets below two Earth radii in the habitable zone at an intermediate occurrence rate, so two calculations resting on different stellar and occurrence assumptions land in a similar place. One caveat belongs beside that comparison. Those ten come from sensitivities stated for three other samples rather than for the M dwarfs treated here, so this is not the same target measured twice.
The denominator behind that 4.0 invites a misreading. The 16,929 is a population count inside the footprint, "not the list of stars P4 will observe," and the yield scales linearly with the fraction actually observed. Most of that reservoir sits at an expected signal-to-noise of 2 to 6, around stars too faint for any method at all.
4.1What the Author Marks Down Himself
The forecast arrives with a long list of caveats, all of them the author's own. The heaviest one is flares. The simulator does not produce them, and for an M-dwarf sample that is not an omission to wave past. He writes that it leaves both completeness and false-alarm rate optimistic by an amount no control arm in this study can bound. Yaptangco et al. applied the same family of chain to real TESS photometry of M dwarfs in 2025 and found the 50% threshold moving from a signal-to-noise of 7±1 on inactive hosts to 14±1.5 on active ones. The numbers in this paper describe the quiet end of the population.
That said, not everything missing from the simulator should be counted as an omission. Granulation and the stochastic activity term were switched off because they are calibrated on solar-like and evolved stars, and solar-like oscillations were switched off because the modes of main-sequence M dwarfs lie far above the frequency limit that 300-second bins can see. Those were disabled with a reason attached. Flares were not, which is why the paper singles them out as the consequential one.
The remaining caveats mostly point the same way. Completeness and yield are scored at a floor of 7, so at the adopted 7.5 they are a few points lower. The completeness curve was measured only out to periods of 100 days, whereas about a third of the real habitable-zone hosts sit between 100 and 160 days, where the baseline delivers only 4 to 6 transits. That is the genuine few-transit regime the campaign never sampled, and the long-duration argument given above no longer holds there.
The 41% reflects the statistical width of the completeness curve and nothing else, and the author lists three effects that push it towards an over-estimate: the camera-count scaling, the extrapolation in period, and the extrapolation of the noise model towards brighter stars. The last of the three carries the most weight. The noise interpolation is anchored at P-band magnitudes of 13.5 to 15.0, so 15 of the 35 habitable-zone hosts are handled by extrapolation, and those 15 carry roughly 80% of the integrated probability. The impact parameters of the injections point the same way, sampled only from 0 to 0.6 while the forecast draws them from 0 to 0.9, which leaves about a third of the forecast population outside the range where completeness was actually measured, in the unfavourable direction. Even so, he notes that the limits point in both directions and that their relative weights are not established in this work.
One more thing. Every yield here is a candidate count. Pixel-level effects and blends are out of scope, so the class of false positive where a background eclipsing binary mimics a planet is never probed at all. Going from candidate to planet still takes separate observations.
Does Our Pipeline Have That Curve?
That is where the paper ends. From a data practitioner's seat, the trade is familiar. Turn up the sensitivity of an anomaly rule and alerts multiply while missed events fall. Tighten the rejection criterion in label review and what passes gets cleaner while sound records come back with the bad ones. Missing-value checks, duplicate detection, personal-data scanning, document filtering: wherever a pass criterion is set, the same exchange rate is running.
The difference is that most lines are drawn without knowing the rate. A threshold gets chosen by convention, or because the alert volume looked manageable on the first day it ran. Lines set that way are not wrong, but when someone proposes moving one there is nothing to argue from. This paper leaves behind a curve rather than a conclusion. Wherever you put the floor, you can read the price of that position as a number.
The study also shows what drawing such a curve requires. The four questions below are not in the paper, and they are worth working through in this order when we check the thresholds in our own pipelines.
- Do you hold data that is certain to contain no answer? This study deliberately built 465 twins with no planet injected. A false-alarm rate can only be measured on data guaranteed to have nothing wrong with it.
- How varied is the data whose answer you do know? The 525 injected planets were placed with expected signal-to-noise between 3.5 and 13, clustered around the threshold. Plant only easy cases and completeness will always look high.
- Do you report the two numbers as a pair? A dashboard that shows only a pass rate leaves you no way to tell how many phantoms came through with it.
- Have you written down how the axis of the curve is defined? The signal-to-noise axis in this paper uses the limb-darkened window depth, measured at 1.20 times the geometric depth, and "using the geometric depth would displace the axis by 20%." A threshold copied across without its definition becomes a different number in the next team's hands.
The reproducibility terms come with the paper. The author published the scored tables from both arms, the generator scripts and random seeds that regenerate all 1,720 light curves, the scripts that draw the figures, and the detection chain frozen at the version used for the campaign. His acknowledgements state that the pipeline and analysis were developed by the author with assistance from Claude (Anthropic), and that all code, results and text were reviewed and validated by the author. That combination is part of how one person produced work at a precision comparable to the performance assessment of a large space mission.
Editor's Note
A question Pebblous hears often while diagnosing data quality is what number to set the criterion at. The answer usually depends on the data, so instead of an answer we suggest building the curve. The cost of a miss and the cost of a phantom differ from team to team, and it is better to choose once you hold the exchange rate between them.
Thank you for reading this far. The full paper is available at arXiv:2609.04887, and the completeness and false-alarm figures in this article were checked against Sect. 3 and Fig. 2, the yield forecast against Sect. 4, and the reservations against Sect. 5. If your team has set its pass criteria from a measured curve, we would be glad to hear what control data you built to get there.
References
Primary source
- 1.Tschudi, Y. (2026). "Habitable-zone Earths at the detection frontier: Measured completeness and false-alarm rate of a transit pipeline for the PLATO M-dwarf sample." arXiv:2609.04887.
PLATO yield-forecast literature
- 2.Heller, R., Harre, J.-V., & Samadi, R. (2022). "Transit least-squares survey. IV. Earth-like transiting planets expected from the PLATO mission." Astronomy & Astrophysics, 665, A11.
- 3.Matuszewski, F., Nettelmann, N., Cabrera, J., Börner, A., & Rauer, H. (2023). "Estimating the number of planets that PLATO can detect." Astronomy & Astrophysics, 677, A133.
- 4.Schlecker, M., Apai, D., Lichtenberg, T., et al. (2024). "Bioverse: The Habitable Zone Inner Edge Discontinuity as an Imprint of Runaway Greenhouse Climates on Exoplanet Demographics." The Planetary Science Journal, 5, 3.
- 5.Cabrera, J., Csizmadia, S., Montalto, M., et al. (2026). "Assessment of PLATO Science Performance." Experimental Astronomy.
Kepler post-hoc diagnostics & related work
- 6.Thompson, S. E., et al. (2018). "Planetary Candidates Observed by Kepler. VIII. A Fully Automated Catalog With Measured Completeness and Reliability Based on Data Release 25." The Astrophysical Journal Supplement Series, 235, 38.
- 7.Burke, C. J., et al. (2015). "Terrestrial Planet Occurrence Rates for the Kepler GK Dwarf Sample." The Astrophysical Journal, 809, 8.
Methodology & caveat support
- 8.Yaptangco, D. C., Ballard, S., & Dittmann, J. (2025). "Quantifying the Effect of Short-timescale Stellar Activity Upon Transit Detection in M Dwarfs." The Astronomical Journal, 169, 153.