Executive Summary
US regulation lets a water utility settle a pipe's material with a statistical or predictive model rather than digging it up. New York is one of the few states that also publishes, address by address, which method produced each call — which puts those model outputs next to what the same utilities' crews actually found in the same neighborhoods. An independent researcher read the statewide file that way. Of every method that made a call, only the predictive model recorded neither lead nor unknown, on all 43,000-odd addresses. What is missing is not the pipe. It is the value written in the ledger.
The strongest evidence is not a comparison with other cities but one held inside a single city. New York City writes lead hundreds of thousands of times and writes unknown hundreds of thousands of times, and never once in the stretch the model handled. Two further facts need no estimator at all. State guidance permits that value only where written records show construction after June 1986, yet thousands of these addresses sit under buildings that predate 1940 and the installation-date column is empty on every one of them. And in the 2025 snapshot the public-side determinations did not exist. The values sitting there now came across from the customer side — a different pipe with a different owner. A village of a few thousand people, 550 km from New York City and served by an unrelated utility, reproduces the same pattern on its own.
None of this says the model was wrong. The paper reports no accuracy figure and names no vendor. A confidently mistaken model, a sensible workflow that sends the ambiguous cases to a crew and records only the cleared ones as model output, a pipeline that discards uncertainty before submission — all three fit this data, and a reader of the ledger cannot tell them apart. When how a value was produced and how uncertain it was are not written down beside it, the only way left to recover that later is to pull a year-old snapshot out of the Internet Archive and diff it.
43,215
NYC addresses classified by model
One material value on every one of them; zero lead, zero unknown
0.0085%
95% upper bound on that bucket's lead rate
Records-based calls in the same city find lead on 19.80%
7,782
of them under pre-1940 buildings
Guidance permits the value only with post-June-1986 records; the date column is 100% empty
1,150–1,450
Expected lead public-side lines, six estimators
Not a confidence interval — the spread of disagreement between specifications
The rule allows models, with conditions
The US Lead and Copper Rule Revisions (LCRR) required every community water system to publish a service line inventory by October 2024. Small amounts of lead matter here. In October 2021 the Centers for Disease Control and Prevention lowered its blood lead reference value for children to 3.5 µg/dL while restating that no safe blood lead level has been identified. Where a system does not know a line's material, the rule leaves open a path other than the shovel: a statistical method or a predictive model. A commercial market grew on top of that permission. Utilities buy models for a plain reason — digging up a single line costs an order of magnitude more.
The size of that gap appears in material the US Environmental Protection Agency circulates. A per-line cost table by identification method puts excavation at $1,120, sequential water sampling at $715, and visual field inspection at $29. Predictive modeling carries no per-line figure in that table, because no primary source publishes one. Replacement, which follows the call, costs far more: EPA works from about $4,700 per line (range $1,200–$12,300), while the American Water Works Association puts it between $8,247 and upward of $12,000. The distance between those two numbers is itself a long-running argument in the sector. Through the infrastructure law the federal government set aside $15 billion for lead service line replacement across fiscal years 2022–2026, of which New York State has received $369 million through the third allotment.
A model saves money not because the call itself is cheap but because it shrinks the number of lines anyone digs up. In cases its own marketing describes, one leading vendor estimates that the Detroit Water and Sewerage Department avoided $165 million by adopting predictive modeling, and reports that South Bend physically verified 125 of roughly 50,000 lines (0.25%) and handled the rest with the model. Both figures are vendor-published and should not be taken at face value, but they show exactly where the savings come from. The intended use of a model is to shrink the population that gets verified. It is not to remove verification.
How regulation governs a call with that much money attached is the baseline for everything below. Version 3 of the New York State Department of Health's service line inventory guidance lists seven ways to identify material: utility or public records, field inspection by system staff or a professional plumber, excavation, sampling, statistical analysis or a predictive model, customer self-identification confirmed by staff, and other methods acceptable to EPA or NYSDOH. Only excavation and field inspection look at the pipe. The rest are desk methods, and the guidance attaches different conditions to each.
1.1A model output is not presumptively acceptable
The item covering predictive models asks the question bluntly: "Is a predictive (probability) model or statistical analysis acceptable to become a known service line without physical verification?" The answer is conditional.
"A model's output typically needs physical verification due to an inherent inaccuracy of any model or statistical analysis. However, on a case-by-case basis, some of the model and statistical analysis results will be accepted without physical verification. You must provide sufficient information to the State to evaluate how much physical verification is adequate." Two of the examples the guidance lists are the heart of this audit: "random physical verification process such as the proposed number of SLs that will be physically verified" and "confidence interval for the model." The item then closes: "Note that a State's initial determination for a required physical verification rate can be revised based on the accuracy of physical confirmation results."
That closing sentence writes a feedback loop into the rule. The verification rate is meant to respond to what the shovels find. Whether it does is a question we can put back to the file in Section 2.
1.2Eight values separate "not lead" from "might be lead"
A system picks one of eight values in the material column. In the template's own order: Lead including lead-lined galvanized, Copper, Galvanized, Plastic, Known Other, Unknown but could be lead, Unknown but unlikely lead, and Unknown. Three of the eight are explicit unknowns, and one of those three flags that the line could be lead.
The value that recurs most in this report is the fifth one, Known Other, and the word "other" is misleading if you read it casually. In the state template it does not sit with the three unknowns; it sits alongside named materials like copper and plastic, and it asserts a line whose lead status is known to be negative even though its material is not named. Aggregate the inventory and it maps to non-lead. The hedge Unknown but could be lead is the value at the other end: an unknown that is explicitly flagged as possibly lead.
The guidance also pins down when Known Other may be used. A system needs written records showing the entire distribution system was constructed after June 1986 or after that municipality's own lead ban, and the entire length of the customer-owned line must postdate the same date. A fallback is attached in the same item: "If you do not have such records, you need to verify service line material with one or more methods included in Item 14." Absent the records, in other words, go back to a method that looks at the pipe.
1.3The same guidance gives sampling a number and the model a procedure
Read the guidance straight through and something appears on the page just before the predictive-model item. Water sampling — the other indirect method, sitting right beside the model — carries a numeric verification standard. "For systems that do not add a corrosion inhibitor, sequential sampling for SL material identification is acceptable only when it is part of a study approved by the NYSDOH. Up to 20 percent physical verification of SL materials tentatively identified with the sampling will be required. If the accuracy of the physical verification result is less than 90%, the sampling should not be used without physical confirmation."
Put the two items in one table and the asymmetry shows.
| Requirement | Water sampling (Item 16) | Predictive model (Item 17) |
|---|---|---|
| Prior approval | Only as part of a study approved by NYSDOH | Case-by-case |
| Amount of physical verification | Up to 20%, stated as a number | "Provide sufficient information to evaluate how much is adequate" |
| Failure threshold | Below 90% verification accuracy, may not be used | No numeric threshold |
| What must be submitted | An approved study plan | A random verification process + a confidence interval |
Source: New York State Department of Health, Service Line Inventory Guidance (LCRR) v3, August 2025, Items 16 and 17. Quotations checked against the source PDF for this report
Reading this as "the rule went easy on models" gets it wrong. The guidance does make demands of the model side, and they are not light ones: a random verification process and a confidence interval. What is missing is any way to check from inside the inventory whether those demands were met. Sampling gets a sentence you can adjudicate — below 90% accuracy, you may not use it. The model gets a procedural instruction to submit sufficient information. That difference is the backdrop to everything observed below.
One aggregation swept 153 localities
The audit starts from a single test. For each locality, how many distinct material values do its model-based classifications take? And what did that same locality's own physical verification find? The first question is a one-line aggregation; the answer to the second is already sitting in the same file. The paper ran both stages across all of New York in seconds.
The scope is the 153 localities that classified at least 100 addresses with a model. Locality is the finest reporting unit the state file publishes, so that is the unit the screen uses. Of those, 75 (49%) record exactly one distinct material value, and those 75 cover 125,990 addresses — 57.0% of all model-classified lines screened.
This is where the paper's restraint begins. A single value is not by itself misconduct. For a system serving housing built after 1986, one value is what you would expect. That is why the screen has two stages: zero variance flags, and only contradiction confirms. Of the 75, 68 are consistent with their own physical verification or have too little of it to check at all, and the paper defends those 68 explicitly rather than leaving them under suspicion.
The second-stage threshold was fixed in advance: localities with at least 200 physically verified lines of which at least 1% are lead. Seven localities clear that bar and still show a single value across the model bucket. Five are boroughs of New York City, one is East Rochester 550 km away, and the last is Troy.
| Locality | Model-classified | The one value recorded | Physically verified | Verified lead rate | Expected lead | P(observe zero) |
|---|---|---|---|---|---|---|
| Staten Island | 16,434 | Known Other | 20,500 | 2.31% | 380 | 10-167 |
| Brooklyn | 10,308 | Known Other | 52,024 | 7.54% | 777 | 10-351 |
| Queens | 8,513 | Known Other | 66,460 | 14.19% | 1,208 | 10-566 |
| Bronx | 4,867 | Known Other | 17,769 | 5.46% | 266 | 10-119 |
| Manhattan | 3,315 | Known Other | 11,751 | 3.40% | 113 | 10-50 |
| East Rochester | 472 | Known Other | 557 | 9.69% | 46 | 10-21 |
| Troy | 187 | Copper | 325 | 1.54% | 3 | 0.055 |
Localities whose model output takes a single value and whose own crews contradict it. The last two columns are "how many lead lines the model bucket should have produced if that locality's own verified lead rate applied to it" and "the probability of nonetheless observing zero." Source: arXiv:2608.19922, Table 1
The last row is one the paper itself tells you to set aside. Troy's expected count is three and the probability of observing zero is 0.055, which is unremarkable. The paper explains why it stayed in the table anyway: the screen's threshold was fixed in advance, and dropping a case after looking at its p-value is precisely the practice this paper objects to. For the other six, the probability of observing zero runs from 10-21 to 10-566, so no argument here leans on Troy.
Five of the seven rows are New York City boroughs, which invites an obvious objection: one city counted five times. The unit the paper uses is the locality, the finest geography the file publishes, not the regulated entity. To the Department of Health the five boroughs are one water system (NY7003493). Collapse them and the count becomes 149 units, of which 71 (48%) take a single value and three rather than seven are contradicted. The 125,990 addresses covered by the single-value bucket do not change. The paper reports both figures and notes that nothing downstream turns on which one you use.
2.1East Rochester takes the obvious objection off the table
If only the five boroughs had been flagged, the finding could be read as a big-city peculiarity or one utility's exception. East Rochester closes that door. It is a village of a few thousand households, 550 km from New York City, served by an unrelated utility, with nothing in the data indicating a shared vendor relationship. Its model classified 472 addresses and recorded Known Other on every one of them. Crews in the same village opened 557 lines and found lead on 9.69% of them.
One more detail matters. Those 557 were all excavations — not field inspections, which see only the accessible portion of a line, but holes in the ground with the pipe in view. The contradiction does not rest on the weaker of the two physical methods.
2.2More than 40,000 addresses cannot be checked at all
To reach stage two, a locality needs physical verification of its own. Most have very little. The median locality performs 0.14 physical verifications per model classification. Twenty-three of the 153 have performed none at all, and those 23 cover 40,468 model-classified addresses. No screen can check that stretch.
The clearest example is the locality of Greece. It classified 16,135 addresses by model, all with one value, against 74 verified lines available for comparison. So it never reaches the contradicted list. Not contradicted is not the same as confirmed.
The natural follow-up — do systems that lean harder on models verify less? — is one the paper tests. Expressed as shares of each system's total line count, the two are unrelated (Spearman −0.076, p=0.35). Expressed as model share against verifications per model classification, a relationship appears (−0.433, p=2×10-8). Both forms share a term between numerator and denominator and so carry a ratio artifact, and the paper draws a conclusion from neither. It says only that verification levels are low across the board.
2.3The feedback loop is not closing
Section 1 ended on the guidance sentence that makes the verification rate responsive to what the shovels find. Two snapshots let us look for that adjustment. Between June 2025 and August 2026, model-classified addresses statewide moved from 298,641 to 298,704, and the count of those recorded as lead or as the hedge went from 28,971 to 28,974. Over fourteen months: 63 addresses and three classifications. Over the same period, 664,259 physical verifications were performed across the state.
Crews opened the ground 664,000 times in fourteen months, and the record for 300,000 model-classified lines is effectively frozen. The learning circuit the rule assumes is not turning inside this file.
One ledger, one city, two distributions
New York State is estimated to hold about 494,000 lead service lines, sixth in the country by EPA's drinking water infrastructure survey. The case this audit examines in most detail is a single city, and narrowing to New York City addresses that carry usable coordinates leaves 813,911 of them. Split those addresses by the basis on which each was classified and you get the table below.
| Basis of classification | Addresses | Share | Lead | Unknown | Median year built |
|---|---|---|---|---|---|
| Records review | 532,906 | 65.5% | 19.80% | 9.26% | 1931 |
| Field inspection | 95,087 | 11.7% | 8.45% | 0.01% | 1931 |
| Excavation | 72,174 | 8.9% | 9.86% | 2.60% | 1930 |
| Not verified | 70,529 | 8.7% | 0.00% | 100.00% | 1930 |
| Predictive model | 43,215 | 5.3% | 0.00% | 0.00% | 1984 |
New York City public-side classifications, 813,911 addresses. "Unknown" is the sum of the three unknown values. Median year built comes from a MapPLUTO tax-lot join. Source: arXiv:2608.19922, Table 2
Two rows show a lead rate of zero, and running them together would collapse the whole point, so they have to be separated first. The Not verified row is zero because nothing was classified: all 70,529 of those addresses are recorded as unknown. Where no call was made, the absence of lead is trivial. The Predictive model row is a different object. It carries a call on all 43,215 addresses and still records neither lead nor unknown. Stated precisely: of the methods that made a call, only the predictive model never once used either value.
The physical side carries a caveat of its own. The table separates field inspection from excavation, and although this report pools them as "physical verification," they are not one instrument. Restricted to clean-label addresses, excavation finds lead on 10.13% and field inspection on 8.48%, a difference that is statistically sharp (χ², p=2.8×10-30). The odder direction is in the unknowns. Excavation records an unknown material on 2.60% of the addresses it touches; field inspection on 0.01%. A method that sees less of the line returning far fewer unknowns is the wrong way round. The guidance lists the two separately and field inspection may observe only the accessible portion, so the pooled label is not ground truth in the laboratory sense — it is closer to the weaker of the two.
3.1Same filer, same template, same reporting period
The hardest form of this finding needs no second utility. New York City has values available for lines whose lead status is unsettled, and it uses them heavily. Unknown appears on 121,779 addresses, of which 49,365 were classified from records and 1,880 had already been excavated. Lead appears on 120,692. In the model bucket both counts are zero — zero across all 43,215.
Same filer, same template, same reporting period. Every other basis in the table produces both values. From zero events in 43,215, the 95% upper bound on the lead rate is 0.0085%. Note that this is a bound, not a confidence interval. It does not say the true rate is 0.0085%; it says that given zero observations, the true rate is unlikely to be higher than that.
One more feature of the city's file is worth flagging. Across 817,375 addresses and all five methods, the public-side material vocabulary runs to four values: Known Other, lead, Unknown and Galvanized. The other four of the eight never appear under any method. That fact is both the raw material for the strongest counter-explanation and the reason it falls short.
3.2The rest-of-state comparison comes with two weaknesses attached
Outside New York City the same method uses a wider vocabulary. Across 176,888 addresses it records outright lead on 3.57% and the hedge Unknown but could be lead on 8.64%, 12.21% combined, and copper, plastic, galvanized and "unknown but unlikely lead" all appear as well. On the numbers alone the contrast with the city is dramatic. The paper attaches two reasons not to lean on it immediately after quoting it.
First, 93% of the outright lead is one city, Poughkeepsie — 5,885 of the 6,320 rest-of-state model lead findings — and Poughkeepsie has physically verified nothing at all. The paper writes plainly that "Poughkeepsie may be over-calling by as much as we argue New York City is under-calling." Exclude it and the outright rate falls to 0.26% and the combined rate to 9.26%. Second, the hedge that supplies 8.64 of those 12.21 points is a value New York City has never used under any method. You cannot build on the absence of a value the filer does not use.
So the 12.21% serves only as supporting evidence that the template's full range is in live use. What the argument rests on is the within-city comparison in the previous subsection, and that comparison depends on neither Poughkeepsie nor the hedge.
3.3The "the pipeline is binary" explanation
The paper builds the strongest defense the filer might offer, on the filer's behalf. If New York City's pipeline is effectively binary, then every call that is neither lead nor unknown collapses into Known Other. The four-value vocabulary is exactly what that would look like, and the model may well have produced a distribution that the pipeline then flattened. The paper accepts that far, and says so explicitly: it makes no claim about what happens inside the model.
What the explanation does not cover is what remains. Even in a binary pipeline, Unknown is still an available value, and the same city used it 121,779 times — never once in the model bucket. And records review, the other desk method with no site visit, works within the same four values and finds lead on 19.80%. A narrow vocabulary explains why the model bucket never says "copper." It does not explain why that bucket alone never says "unknown" or "lead."
Construction era explains a third
The paper concedes up front that the model was not assigned to addresses at random. The model-cleared population has a median year built of 1984 against 1930 for the physically verified one, and the share built after 1960 is 77% against 24%. New York City banned lead pipe in 1961; the federal ban came in 1986. A genuinely lower true lead rate in the model bucket is therefore expected, and any comparison that ignores the difference overstates the finding.
The question this section asks is what survives once construction era is held fixed. Joining every address to New York City's MapPLUTO tax-lot data and splitting by era gives this. Records review — a desk method with no site visit, and the city's largest bucket at 532,906 addresses — finds lead in every era, at rates from 4.32% to 31.85%. Physical verification finds lead in every era too, from 1.45% to 14.45%. The predictive model finds it in none of them. There are 4,681 model-classified addresses in buildings raised between 1920 and 1940 alone, an era in which every other method in the same city turns up lead on 14% to 32% of what it touches.
This is where "a desk method is naturally conservative" stops working. In the same city, the other desk method is the one that finds the most lead. Records review does not look at the pipe either.
4.1The 7,782 addresses that need no estimator
Of the 43,215 addresses the model cleared, 7,782 (18.0%) sit under buildings raised before 1940: Brooklyn 2,679, Manhattan 1,991, Queens 1,572, the Bronx 861, Staten Island 679. New York State's own replacement funding program uses the count of pre-1939 housing as a proxy for lead risk. And the physically verified lead rate for pre-1940 New York City buildings is 12.93% (95% CI 12.72–13.14%, n=98,749). Apply that rate to those 7,782 alone and roughly 1,000 lead lines are currently recorded as Known Other.
What makes this subgroup strong is not the arithmetic. It is the guidance condition from Section 1, which bites here directly. The value is permitted where there are written records of construction after June 1986. These buildings predate that by half a century, and the installation and replacement date fields are 100% empty. The kind of record the guidance contemplates does not exist inside the published inventory. And the same item already told us what to do in that case: verify with a method that looks at the pipe. The fallback is triggered by the guidance's own condition.
In fairness, the paper's caveat travels with the finding. A missing date in a submitted inventory is not proof that the utility holds no such record. It is proof that the published inventory does not show one.
A caveat cutting the other way is recorded too. MapPLUTO's year-built values heap on round numbers: 44.9% of joined addresses fall on a multiple of ten and 71.0% on a multiple of five. The era boundary sits on top of one of those heaps, and because the bins are left-closed, a lot recorded as 1940 is counted post-1940. The 7,782 figure is more likely an undercount than an overcount.
4.2Sizing the residual
Five estimators that do not control for era put the expected number of lead lines in the model bucket at 1,640–2,775. Six that do control for it give 1,152–1,461. Construction era, in other words, accounts for about a third of the gap and not the rest. The figure the paper reports is the era-controlled one, 1,150–1,450 lines, or 2.7–3.4% of the model-cleared population — five boroughs, public side only.
The nature of that band is easy to misread, so the paper gives it a subsection. It is not a confidence interval; it is the spread of disagreement between specifications. Bootstrapping the physically verified sample 400 times puts the sampling error of any single estimator at around ±3%. The borough × era estimator gives 1,461 with a 95% interval of 1,419–1,503; borough × building type × era gives 1,379 (1,331–1,421); ZIP × era gives 1,275 (1,227–1,330). Almost the entire width of the band is methods disagreeing with one another.
And a larger uncertainty sits outside the statistics. The whole band rests on conditional exchangeability: the assumption that within a stratum, which addresses got dug is no longer related to true material. That is an assumption, not a randomization, and the paper names the direction in which it fails. A utility that trusts a model enough to skip the shovel plausibly leaves the easier-looking cases to the model and sends crews where something already looks off, which biases the verified sample's lead rate upward. "There are 1,450 lead lines hiding in there" is not a claim this paper makes.
The paper also re-runs everything with excavation as the only physical method. The screen is unchanged at 75 single-value localities and the same seven contradicted, and East Rochester's 557 lines were all excavations to begin with. The pre-1940 lead rate rises from 12.93% to 13.35%, lifting that subgroup's implied count from about 1,000 to 1,040, and the estimator band widens downward to 850–1,325 — a minimum of 20 observations per stratum leaves fewer strata supported when only excavations count.
There is also a crack in the era proxy itself. Half the model bucket sits in post-1980 buildings, and physical verification finds lead on 2.98% of post-2000 buildings (n=13,683) — higher than the 1.45% for 1980–2000. Old pipe left in place under new construction shows up in exactly this shape. Treat the reversal as spurious and force the post-2000 rate down to the preceding band, and the era-controlled estimate drops about 11%, to 1,140–1,300. That is how much the answer moves on whether building age is a usable proxy in the newest era.
4.3A floor built from public data alone
The paper closes by building a baseline model of its own. It trains on the 165,063 physically verified New York City addresses with clean material labels (15,154 of them lead, 9.18%), using only coordinates and building type. Customer-side fields are excluded because some of them are themselves model-derived, and using them would leak vendor model output into a supposedly independent baseline.
The result is a range rather than a point. Depending on how far apart training and test data are held, AUC moves between 0.64 and 0.74. The paper treats that diagnostic as essential to any spatial prediction claim and notes it could not find it in the existing literature. It also leaves implementers a warning: include ZIP as a feature while also splitting folds by ZIP and measured AUC drops by 0.08, because a feature that is dead at test time eats the splits. A pipeline that does this, then retrains on all the data and publishes a headline number, is reporting an accuracy that does not describe the model it deployed.
Two diagnostics matter more than the AUC. Predicted probabilities summed over observed counts give 0.875 overall and 0.866 in the region where the model bucket sits — the estimator undercounts by about 13%, which makes the reported band conservative. And a classifier trained to tell the model-cleared population apart from the verified one reaches AUC 0.753. The two groups separate cleanly on covariates, which is precisely why the era join was necessary.
How to read 0.64–0.74 is spelled out as well. It is a floor on the difficulty of the task, not a ceiling on what a vendor model with work orders, permit histories and tax records can do. Adding MapPLUTO's year-built and built-form columns — one free municipal dataset — lifts AUC in the same design from 0.720 to 0.755. And that cuts back toward the outputs under audit: if coordinates and building type alone separate lead from non-lead this well, an output with no variation at all is harder to explain, not easier.
Setting 0.64–0.74 beside the numbers the industry publishes is not straightforward. As the paper summarizes the field, published performance figures run from roughly 73% to 99.9% with no shared benchmark — a spread that is itself evidence these numbers are not comparable to one another.
Independent validation against excavation exists exactly once in the whole literature. A 2017 geospatial model in Flint was checked against 26,750 excavations and returned AUC 0.9 for copper and galvanized material but 0.6 for lead, barely better than chance, improving to 0.8 only after the kriging method was replaced with a compositional variant. The paper does not treat that one study as closing the question. Flint is one city and the most intensively studied water system in the world since 2016; the model validated there predates the LCRR by seven years and is not one of the commercial products now sold into the compliance market; and it was checked by the author of the model being checked.
More than that, the failure mode is different in kind. In Flint the model attempted "lead" sometimes and aimed poorly. Here, across half the New York water systems that use one, the recorded output does not vary at all.
The copy trace left in the file
Everything so far leans to some degree on comparison populations and estimators. The two facts in this section do not. They are marks left in the file itself, and they carry dates. Of the two findings the paper says it would defend first, one is the pre-1940 subgroup above; the other is here.
Reading the file at all requires clearing some traps first. The paper documents five in the source data, one of which inflates New York City's apparent size twofold.
- Rows are not lines. The August 2026 snapshot's 4,618,115 rows resolve to 3,744,223 distinct locality–street–ZIP keys, and the excess is concentrated in one place. Every New York City borough has a rows-per-address ratio of 2.001, while no other locality above 5,000 rows exceeds 1.14. The city files two rows per address: one carrying the public-side columns, one carrying only the customer-side columns.
- Read the material, not the category. The address-level category field collapses to four classes and, outside New York City, reports the worse of the two sides. On 9,625 addresses the public-side material is definitely not lead while the category says lead, driven by the customer side.
- The method and material fields are free text. Locality strings are case-split, and the city's boroughs appear as the codes QN, BK, SI, BX and MN — a literal filter on "Queens" returns 11 rows statewide. The model method appears in its canonical form, Statistical Analysis/Predictive Model, 219,899 times, and under four further spellings a further 4,603 times.
- The published column descriptions are off by one. In the dataset's own metadata, the description attached to the public-side material column actually describes the next column, and the error propagates down the list. Authority over what a field means sits with the state guidance document, not the portal metadata.
The free-text trap forces a choice on the auditor. Matching the spelling variants broadly moves the rest-of-state comparison from 12.21% of 176,888 addresses to 11.93% of 181,065. The paper matches broadly in the screen and uses the canonical spelling alone in the rest-of-state comparison, and publishes what that choice costs. No New York City figure depends on it: none of the variant spellings occurs in the city at all.
5.1The public-side call was copied from the customer side
The second trap leads to a fact as important as the headline. On New York City's public-side rows, the public verification method equals the customer method on 100.0% of rows — a perfect diagonal across all five methods — and the public material equals the customer material on 100.0% as well.
Agreement alone does not say which side is the original. What fixes the direction is the June 22, 2025 snapshot retrieved from the Internet Archive. At that point New York City's 817,375 addresses appeared once each, and the public-side columns were blank on every one of them. All 43,440 model classifications existed only on the customer side. Fourteen months later the same addresses carry the identical value on the public side.
The public-side determination was not made independently. It was created from a model output produced for the customer's pipe — which under the rule is a different pipe with a different owner. That one value came across from a prediction about another pipe is invisible to anyone reading the ledger, and recoverable only by pulling a year-old archived capture and diffing it.
This is the reconstruction cost of not recording provenance. In this case the cost happened to be payable: the source was a public Socrata dataset, the Internet Archive held a capture from a year earlier, and the paper pinned both snapshots by SHA-256 so the comparison reproduces. Remove any one of the three and the direction of the copy stays undetermined.
5.2Two audits, one file, errors pointing opposite ways
This inventory has already been through one official audit: report 2024-S-9, issued by the Office of the New York State Comptroller on January 5, 2026, covering March 2018 through June 2025. Its stated objectives include whether the inventory is accurate and was completed on time. The paper notes that the report nonetheless did not examine the predictive-model path.
What it examined instead sharpens the contrast. The Comptroller found that the Department of Health's controls and guidance were inadequate; that a substantial share of state funds went unspent or was used for administrative costs rather than replacement; and that some areas with elevated childhood blood lead levels were left out of prioritization. On the inventory itself, it found that about a third of covered water systems missed the October 2024 deadline, that 140 remained non-compliant as of August 2025, and that the data was inaccurate or incomplete. Its example of a data error: service lines already replaced that were still listed as lead or unknown. The Department largely disputed the findings and replied that compliance reached 90% within months of the deadline.
Two audits read the same file and found errors pointing in opposite directions. The Comptroller found overcounting — lines recorded as lead that are not. This paper finds undercounting — lines recorded as not-lead that might be. Neither is wrong. Neither was looking along the other's axis.
Search a ledger for errors in a direction you have already chosen and the other direction stays invisible. What makes both directions visible at once is the basis-of-classification column. Without it you cannot ask where a value came from, and if you cannot ask, you check in one direction only.
Why Pebblous is watching this
Nothing this audit asks for is a new practice. Statistics Canada's quality guidelines on imputation put it this way: "Good imputation processes are automated, objective and reproducible, make an efficient use of the available auxiliary information, have an audit trail for evaluation purposes and ensure that imputed records are internally consistent." The guidelines are specific about what has to survive: "Information on the imputation process should be retained on the post-imputation files… Such information includes variables indicating which values were imputed and by which method, variables used to indicate which donors were used to impute recipients and so on."
Recording how each value was produced is an official quality requirement of a national statistical agency, documented since at least 2009. New York's template making the basis-of-classification column mandatory is not exceptional foresight; it is baseline statistical practice. It is also the only reason this audit could happen.
1This is not an argument against putting predictions in a registry
Without a counter-case, this piece reads as "keep model output out of official records." That is not the claim. In medical administrative data the opposite result is on record. Ross and colleagues' 2024 cystectomy model study reports that determining case status by probabilistic imputation reduced misclassification bias relative to using procedure codes, and recommends doing so. In the authors' words: "Accuracy of administrative database research can be increased by using probabilistic imputation to determine case status instead of individual codes."
A model's prediction can be more accurate than the code on file. The problem is not that a predicted value exists; it is that the value is not marked as predicted. Ross and colleagues' result holds only because it is known which values came from the model. What disappeared from New York City's inventory is not accuracy. It is that mark.
2One decision moves the national estimate by more than half
To see the size of the problem, step outside one city. EPA long estimated the national count of lead service lines at about 9.2 million, and that figure was the basis for allocating federal money. In the 2025 update it became about 4 million: 3 million confirmed lead plus 1 million predicted to be lead from the unknown pool. The revision reflects the first measured data to arrive after LCRR made inventory submission mandatory.
For the 24 million-plus lines still recorded as unknown, EPA assumed fewer than one in twelve would be lead when producing the final estimate. The Natural Resources Defense Council objected publicly to the revision: systems that submitted no data at all were treated as entirely non-lead, and where earlier estimates assumed a far larger share of unknowns were lead, this one used one in twelve. In their reading, that is a change of assumption rather than better data. Which side is right is not this report's call. What is clear is that a single decision — what percentage of unknowns counts as lead — moves the national estimate by more than half, and where that assumption is written down is not something you can learn by reading the inventory.
The auditability gap shows here too. At least 34 states accept predictive modeling. Far fewer publish, as an inventory field, how each address arrived at its classification. The federal standard requires the material classification and not the basis for it. Two states are confirmed to publish the basis — New York and Washington — and with no exhaustive survey available, the exact count cannot be stated. The paper's third conclusion follows from that: this one column is what makes the file checkable, and requiring it nationally would cost nothing.
3Three things you can put into a ledger today
Carried over to organizational data, the case reduces to three devices. All three cost almost nothing, and all three are things a data quality practice can do before any regulation asks.
| Device | What happens without it |
|---|---|
| Record how each label was produced | If you cannot tell which method produced a value, no audit is possible at all. Without that column in New York's file, there is no paper |
| Keep a random physical verification sample | New York already required this. The problem was never the absence of the requirement but the absence of any way to confirm compliance from inside the ledger |
| Do not discard the uncertainty values | Having a hedge value in the schema and having the pipeline actually populate it are two different things. The second needs its own test |
New York's Item 17 requirements and the paper's three conclusions in §9, restated for organizational data
The third is the cheapest of the three. New York's template already contains the hedge Unknown but could be lead, used 15,285 times elsewhere in the state, and a deployment covering 43,000 addresses never used it once. In the paper's words: "An inventory that provides for uncertainty and receives none from a 43,215-address deployment is recording less than it was designed to." We see the same shape in labeling pipelines constantly. A serving path that keeps the argmax and throws the distribution away exists in every domain.
The zero-variance screen itself transfers directly as a test. "Do the labels from one source take a single value?" is one aggregation; "does that bucket contradict what another method found?" is the second stage. The paper ran both across an entire state in seconds, and correctly declined to flag 68 systems. Cheap tests are usually cheap because a basis-of-classification column already exists. The same structure — one line of metadata creating auditability — showed up in our report on the metadata line that turns a research agent into a fraud vector, and again in the proof that was right with no record of how it got there. Distributions splitting slice by slice inside one file is what we covered in the safety dataset language-slice audit.
4How far to trust this paper
A few things belong on the record for balance. This is a single-author preprint that has not been peer reviewed, posted on August 20, 2026. As of this writing no press coverage or commentary on it has appeared. At the same time, every input is public, both snapshots are pinned by SHA-256, and the code and reproduction instructions are in a public repository. Readers can open the file and check the judgment themselves.
The regulatory picture is worth stating too. The successor rule, the Lead and Copper Rule Improvements (LCRI), was published in the Federal Register on October 30, 2024 and took effect that December 30. The rule text requires systems to comply with these provisions "no later than November 1, 2027," by which point a baseline inventory must be published and a replacement plan in place. A trade-association challenge has oral argument scheduled for fall 2026, so the final outcome is open, but the rule is currently in force and not judicially stayed. The values being written into these ledgers now become that baseline.
Finally, the paper states plainly what would falsify it: a per-address record linking a model prediction to a subsequent excavation; or documentation that the predictive-model label is applied only after confirmation by another means; or the confidence interval and verification plan the guidance asks systems to supply. Any one of the three would do. And it closes like this.
"Any of these could be produced by a utility or vendor in an afternoon."
That the material needed to refute the audit is half a day's work also means that, had it been in the ledger from the start, there would have been nothing to audit.
Editor's Note. Pebblous builds data quality diagnostics and provenance design, so we have an interest in this subject. This report is not written to recommend a product. It is written to put on record a measured instance, outside our own domain, of what happens when an unverified prediction is absorbed into an official record as a settled value. The demand that every value carry a record of how it was produced is not our invention; it is a requirement national statistical agencies have had in writing since 2009, and the fact that one column honoring it made this audit possible is the whole of the case.
References
Primary source
- 1.Sohail, M. S. (2026). Auditing Recorded Predictive Lead Service-Line Classifications Against Physical Verification: A Statewide Study of New York. arXiv:2608.19922v1 [cs.LG], 2026-08-20, CC BY 4.0. (Single-author preprint by an independent researcher, not peer reviewed. Most figures in this report come from its §2–§10, Tables 1–3 and Figure 2)
- 2.Sohail, M. S. (2026). predictive-lead-service-line-audit (reproduction code and checksums). Both snapshots are pinned by SHA-256
Regulation and policy
- 3.New York State Department of Health (2025). Service Line Inventory Guidance, Lead and Copper Rule Revisions (LCRR), Version 3. August 2025, 15 pp. (Item 14's seven identification methods, Item 16's 20% / 90% sampling standard, Item 17's model conditions, Item 18's conditions for Known Other. Quotations checked against the source PDF for this report)
- 4.New York State Department of Health (2026). New York State Lead Service Line Inventory. Socrata dataset j63k-4n92. (The audited file itself; snapshots of 2026-08-11 and 2025-06-22)
- 5.US Environmental Protection Agency (2021). National Primary Drinking Water Regulations: Lead and Copper Rule Revisions (LCRR). 40 CFR Parts 141 and 142.
- 6.US Environmental Protection Agency (2024). National Primary Drinking Water Regulations for Lead and Copper: Improvements (LCRI). Federal Register, 2024-10-30. (Effective 2024-12-30. §141.80(a)(3) requires compliance "no later than November 1, 2027." The Federal Register full text was read directly for this report)
- 7.Office of the New York State Comptroller (2026). Lead Service Line Replacement Program and Lead Service Line Inventory. Report 2024-S-9, 2026-01-05, audit period 2018-03 to 2025-06. (The report PDF itself was inaccessible; findings were taken from the press release and secondary coverage, and this report does not quote it verbatim)
- 8.Statistics Canada. Quality Guidelines — Imputation. 12-539-X. (The audit trail requirement and "variables indicating which values were imputed and by which method." Source text read directly for this report)
Academic literature
- 9.Goovaerts, P. (2023). Geospatial model of composition of water service lines in Flint, Michigan. AWWA Water Science 5(2). doi:10.1002/aws2.1331. (The only independent validation against excavation, covering 26,750 digs: AUC 0.9 for copper and galvanized, 0.6 for lead, 0.8 after the method change. This report did not open the original and attributes the account to §1 of the audit paper)
- 10.Hensley, K., Bosscher, V., Triantafyllidou, S., Lytle, D. A. (2021). Lead service line identification: A review of strategies and approaches. AWWA Water Science 3(3). (A qualitative review confirming that no independent quantitative benchmark exists)
- 11.Nigra, A. E. et al. (2023). Geospatial assessment of racial/ethnic composition, social vulnerability, and lead water service lines in New York City. Environmental Health Perspectives 131(8):087015. (Prior work that independently established the same MapPLUTO join; its data carries no basis-of-classification field)
- 12.Abernethy, J. et al. (2018). ActiveRemediation: The search for lead pipes in Flint, Michigan. arXiv:1806.10692. (Active learning to choose where to dig — a different purpose from replacing the dig)
- 13.Ross, J., Lavallee, L. T., Hickling, D., van Walraven, C. (2024). Development of the multivariate administrative data cystectomy model and its impact on misclassification bias. BMC Medical Research Methodology. (The counter-case in which model imputation reduced misclassification bias relative to procedure codes. Source text read directly for this report)
- 14.Kutzing, S. et al. (2023). Journal AWWA 115:22–33. (Per-line cost by identification method: excavation $1,120 · sequential sampling $715 · visual inspection $29. Cited via EPA CESER FAQ material)
Statistics and industry sources
- 15.US Environmental Protection Agency (2025). 7th Drinking Water Infrastructure Needs Survey and Assessment: 2025 Update. (National lead service line estimate revised from 9.2M to 4M — 3M confirmed plus 1M predicted from unknowns. New York's roughly 494,000 lines, sixth nationally, is the same survey's April 2023 tally)
- 16.NRDC. The EPA Now Says There Aren't 9 Million Lead Pipes—There Are 4 Million. Be Skeptical. (Objects to the revision as a change of assumption: the one-in-twelve assumption over 24 million unknowns, and non-reporting systems treated as entirely non-lead)
- 17.CDM Smith. What states accept predictive modeling for inventory development? (34 states accepting predictive modeling, 13 without guidance, 3 awaiting federal guidance; 2026-08 snapshot)
- 18.Washington State Department of Health. Service Line Inventory Guidance. (The second confirmed state requiring the basis of classification to be recorded and the inventory published)
- 19.Centers for Disease Control and Prevention. Updates to the Blood Lead Reference Value. (Updated 2021-10-28 to 3.5 µg/dL, the 97.5th percentile of the NHANES 2015–18 distribution for children aged 1–5; CDC's position remains that no safe blood lead level has been identified)
- 20.BlueConduit. Lead Service Line Predictions · No Lead Validation. (Detroit's estimated $165M avoided cost; South Bend's 125 physical verifications out of roughly 50,000 lines. Both are vendor-published figures, cited here only as an illustration of where the savings come from)
Earlier Pebblous reports
- 21.Pebblous (2026). The One Line of Metadata That Turns an AI Research Agent Into a Science-Fraud Vector.
- 22.Pebblous (2026). The Proof Was Right. How It Got There Wasn't Recorded.
- 23.Pebblous (2026). The AI Safety Dataset Missing Every Hausa Self-Harm Entry. 2026-08-18.