Executive Summary
This report traces, back to the primary sources, the two slides from State of Markets II that went round the internet in the week Andreessen Horowitz (a16z) published it on 30 September 2026. The deck runs to 90 pages of charts. One of those two says 69% of the S&P 500 point to a live AI deployment while only 2% disclose a metric they track over time; the other says 98% of US households are not yet paying for AI. Neither number is wrong. In both charts, though, the group named on the axis is not the group that was counted.
a16z did not produce the 69% and the 2%. They come from an outside index by way of Apollo, and that index draws its sample from the S&P 500 and the FTSE 100 combined, minus the chipmakers and AI software vendors — 532 companies. The disclosure ladder it uses has five rungs; the rung that travelled is the fourth, and the top one has been empty two quarters running. The consumer side has the same shape. Where the chart says 'US households', the source says what PNC wrote in the body of its report: PNC households. And that name had already widened in a news story five and a half months before the a16z deck appeared.
Everything above can be checked by anyone who puts the public documents side by side. What follows is this report's reading of them. Calling a bank's customer data 'US households' is not a change of wording. One bank institute that set out to do exactly that built an income estimation model, reweighted its results to the national distribution, and then printed a mean absolute error of 41% in the technical note; the US Bureau of Labor Statistics ruled in January this year that data of the same kind cannot substitute for a survey. What a16z skipped is not a word but that procedure. And PNC never claimed to be speaking about the country.
0 firms
Report AI value as a line of its own, every quarter
Of the index sample of 532 · Q1 and Q2 2026 alike. This is the fifth rung of the ladder
12 firms
Disclose the same measure again the next quarter
2% of the same 532. This division is the only published basis for the 2% that travelled
2.2%
The source figure where a16z wrote 'US households'
PNC customer households · May 2026. PNC's own text calls it the share of PNC households
41%
Income estimation error left after taking transaction data national
The mean absolute error another bank institute printed in the technical note for the same job
Only two of the 90 pages made it out of the deck
State of Markets II is a 90-page slide deck published on 30 September 2026. Its contents run to eight parts, sweeping through a year of markets in charts: macro indicators, semiconductor capital expenditure, the IPO market, consumer AI. Follow-on coverage introduced it as "over 100 charts", a figure a16z never states anywhere. Count the pages yourself and you get 90. One disclosure before anything else: a16z is a venture capital firm with heavy exposure to AI companies, and this deck is market material it published under its name.
Two of the 90 pages broke off and circulated on their own that week: the right-hand chart on page 27 and the top-left chart on page 38. Here is what those two charts actually have printed on them. Everything that follows stands on those words.
69% of the S&P 500 Point To A Live AI deployment. Only 2% Disclose A Metric They Track Over Time. Source: a16z, State of Markets II, p. 27, chart title. Subtitle "Share of SP500 companies" · y-axis "% of S&P 500 Companies" · source footnote "Apollo Daily Spark (9/11/26)"
More and More Households are Paying for AI Source: a16z, State of Markets II, p. 38, top-left chart title. Subtitle "Share of US Households with paid AI subscriptions" · final bar 2.2% · source footnote "PNC Research, Internal Data (July 13, 2026)"
The route the two charts travelled also lies outside the deck. a16z's own summary post restated both figures in prose with no source attribution, and the official social card posted on the evening of 1 October consists of one line — "98% of US households aren't paying for AI yet" — and a link. That is where a positive statement flips into a negative one about the complement. Everyone whose payment has simply not been observed is reclassified as someone who does not pay. Two days later the tech press picked up the sentence, and on 5 October a Korean-language article rendered it as "only 2.2% of US households subscribe to paid AI".
The next question is whether these two pages were the exception. Rendered as ten low-resolution contact sheets and sorted by the source footnote on each chart, the largest group is commercial data providers. Next come panel and internal data, roughly twenty pages, and they cluster in two places. One is the front block that handles AI metrics; the other is the last nine pages of the deck. Pages 82 through 90 are, without exception, first-party data from a16z portfolio companies, and nearly every one carries a footnote pointing to the firm's list of investments. Since the counting was done by scanning contact sheets, "roughly twenty pages" is rounded, but the range of the final nine is fixed page by page. We have read this firm's material once before, in a16z's hardware-only fund.
a16z did not do the counting
Follow the footnote on the page 27 chart up one level and you reach Apollo. The chief economist at the asset manager Apollo published a short note called the Daily Spark on 11 September 2026, and its chart title matches the sentence on a16z's slide almost word for word. So the label "S&P 500" was attached by Apollo, not by a16z; a16z inherited the sentence intact. Go up one more level and you reach the outside index Apollo cites: an AI value realisation index that runs under the name The AI Value Gap.
The index is new and close to a one-person operation. Its author is named, so is his methodology adviser, it states that its evidence comes from earnings calls, 8-K filings, major press and securities filings, and it opens a verdict page for each company by ticker. There is no basis for discounting a source that publishes its methodology and shows its individual judgements simply because the name is unfamiliar. One thing should be plain, though: this is not official S&P 500 statistics. How the index operator assembled his sample is a question for the next section.
Lay out the dates and it becomes clear that the figure had already done a lap of the press before it reached the deck. Forbes ran a story on the same numbers the day Apollo's note went up, 19 days ahead of a16z. After the deck appeared, the direction reverses. The number stays the same from hand to hand while the sourcing comes off one layer at a time.
| Date | Who | What is written there |
|---|---|---|
| 2026-09-11 | Apollo Daily Spark | "69% of the S&P 500 point to a live AI deployment." No sample size, no mention of the FTSE 100. Source given as The AI Value Gap |
| 2026-09-11 | Forbes | "Nearly 70% Of S&P 500 Companies Deploy AI—But Few Track Metrics Over Time". 19 days ahead of the deck |
| 2026-09-28 | Index operator, Q2 update | "532 companies across the S&P 500 and FTSE 100" · "twelve, still just 2% of the sample" |
| 2026-09-30 | a16z deck, p. 27 | Axis "% of S&P 500 Companies", footnote "Apollo Daily Spark (9/11/26)" — the source is named |
| 2026-09-30 | a16z summary post | "nearly 30% of SP500 companies report some 'quantifiable impact' of AI, only ~2% are reporting any tracked metric" — no source given |
| 2026-10-01 | a16z social card | "98% of US households aren't paying for AI yet" — no source given |
Each row reproduces the sentence actually printed in that place. The 28 September update puts figures 2–3 percentage points higher on every rung for the same quarter (see section 4).
The conclusion of this section is therefore not that a16z hid anything. It is closer to the opposite. The deck names its sources on both page 27 and page 38, and those footnotes are what allowed us to climb back to the primary material. The label grows widest and the source disappears not in the deck but in the formats that break off from it. The summary post and the social card drop the qualifier and the footnote at the same time. When the word count shrinks, the first thing cut is precisely the condition you need in order to read the number.
The top rung has been empty two quarters running
The headline that went round quotes two numbers, but the chart it came from has five pairs of bars. The index places companies on five rungs, and each rung carries a one-sentence definition. The table below gives those definitions first and then attaches the Q1 and Q2 2026 figures. Definitions come first because what this ladder measures is not AI adoption but what a company wrote in its earnings calls and filings.
| Rung | The index's own definition | Q1 | Q2 |
|---|---|---|---|
| 1. Intent | An ambition, target or planned investment, with no result yet | 68% | 74% |
| 2. Live | A live deployment with usage, adoption or spend to show for it, but no result | 64% | 69% |
| 3. Quantified once | The first real, quantified result | 26% | 29% |
| 4. Tracked | A result tracked over time — a defined metric, such as cost or margin, that can be followed | 1% | 2% |
| 5. Broken out | AI value reported as its own metric or P&L line, in the same place every quarter, so an outsider can track the same figure over time | 0% | 0% |
Definitions are verbatim from the index; the figures were re-read off the a16z page 27 chart at 600dpi, down to the footnote text. The index caps pilots at rung 2 and states separately that it does not count them as proof.
Apollo puts the top rung in one sentence — "none break AI value out as its own KPI or P&L line". Not one company has pulled AI value out as a separate metric or income statement line.
So "only 2% measure the effect of AI" is not an accurate summary. The share that put a number on a result is 29%. What separates 2% from 29% is not whether they measured but whether they used the same measure again the next quarter. And above that sits one more rung nobody has reached. What made the headline was the fourth rung; the one worth being surprised by was above it. The point that deciding what to count comes before the number is one we took up in the country that counts tokens like GDP.
The same source yields one more ratio, this time with the unit of counting switched from the company to the claim. Take every AI claim pulled out of the filings and ask how many come with evidence attached: 8% as of Q1 2026. There is a per-company version too. Block substantiates 12 of 219 claims, or 5% — the company rated the best at proving things stands behind one-twentieth of what it says. The index then states separately that this figure does not enter its score: it is a companion measure, not part of the index, and not a multiplier. A ratio that counts companies and a ratio that counts claims run side by side, and only the first becomes the score. Both are right; the unit of counting differs. For completeness, the index's own headline reading is 15.7 out of 100 in Q2, up from 13.5 in Q1, and neither of the two slides that broke off carries that number.
Apollo also recorded which side the evidence piles up on. Of the disclosed evidence of effect, 70% is on the cost side and 22% on the revenue side. What the remaining 8% consists of appears nowhere in the source, so this report does not guess. The cost skew by itself does not mean AI fails to lift revenue. Apollo notes the possibility of a lag in the same paragraph, and cost savings are the easier thing to measure in the first place.
"Easier to measure" sounds loose, but the index has measured that too. It tags each disclosed deployment as either machine taking over work a person used to do or machine assisting a person, and across 1,410 cases over two years, the assisting side is 55%. Narrow it to the cases that reached rung 3 and the assisting side drops to 52%; among those that did not get that far it is 61%. Nine percentage points open up. The line the index attaches to that gap is the one that matters most here: proven value rewards what can be measured. Work that assists a person is harder to turn into a number, the index writes, so this classification may not reflect the real mix. The sector cut repeats it. Information services produced a quantified result in ten of eleven companies, while utilities, staples and materials do not clear one in five, and the reason the index gives is not a difference in AI skill by sector but that data products and transactions make it easy to carve out the AI share. Climbing this ladder therefore means not only that a company used AI well but that the work happened somewhere a number could fall out of it.
3.1The leaderboard does not measure the fourth rung
Open the index's leaderboard to see who those twelve companies are and there is something to know first. That table is not scored on the fourth rung. The index publishes the formula on its site. It adds up the highest rung reached in each of three areas — cost, revenue and customer — scales the total to 100, and counts only rung 3 and above. The worked example on the site says so directly.
Block discloses a quantified result in all three areas, each at Level 3… 3 + 3 + 3 = 9, out of a maximum of 15. Scaled to 100, that is 9 ÷ 15 × 100 = 60. Block scores 60, the highest today. Source: The AI Value Gap, "How the index is built", worked example. Same page: "Only Level 3 and above counts."
The seven companies sitting on 60 points are therefore not "companies on the fourth rung" but companies that filled rung 3 in all three areas. The index says as much on its company pages. S&P Global is described as broad evidence across three areas that nonetheless rests on one-off figures rather than tracked ones; The Trade Desk, all three one-off quantified results; Iron Mountain likewise one-off results. On Block the Q2 update is blunter still: it remains at Level 3 and has not yet produced an AI-linked economic metric over time.
Open those pages and each company's supporting figure is printed alongside. For Iron Mountain it is $540M of FY2025 digital revenue — an annual figure, not a quarterly one. Set that down separately. Read the leaderboard as a list of fourth-rung companies, go digging through that company's 2026 quarterly releases, and you will find nothing; not because nothing is there, but because the figure was never on that cycle. Read one table against the wrong standard and you manufacture a false absence on the spot. The thing this report keeps describing happens on the checking side too.
So where is the list of the twelve on the fourth rung? No public document gathers them in one place. The Q2 update does name three in passing: RELX, Auto Trader and FactSet track results with recurring KPIs. The first two are FTSE 100 companies. Section 4 comes back to that.
The same update illustrates, case by case, how disclosure quality divides. This is where it is clearest that what separates rung 3 from rung 4 has nothing to do with how much data a company has.
- Booking Holdings — the index's model case. It puts out a recurring cost metric, customer service cost per booking, and reports that the cost fell by a double-digit percentage while bookings rose about 10% and satisfaction held. At the same time it marks the places where it cannot yet speak. On five AI features in testing, management flagged "a very small sample"; and when it raised the company-wide transformation savings target to $650M a year, it did not break out the AI contribution. For good measure, the same call carries a figure for LLM-referred traffic: still under 1% of room nights.
- Block — the counter-example, and one that mirrors this report's subject. Management attributed to AI a projection that gross profit per employee would roughly double this year, to about $2M from $1M, over a period in which headcount fell by roughly 4,000 people, about 40%. Shrink the denominator and that ratio rises mechanically. The projection was also made in a magazine interview rather than on an earnings call, and the three preceding calls explain gross profit growth by payment volume and software adoption. The index operator called the mismatch a red flag.
- S&P Global and eBay — between the two. S&P Global said in Q2 2026 that its enterprise data organisation had achieved about 60% of a targeted $100M annualised cost saving, but that is one unit's figure rather than the whole company, and on the same call it declined to give an annualised number when asked for the AI contribution in another division. eBay recorded a 50% rise in listings per seller in Q1 2026 and a figure in the same family the quarter before. Of the five company pages we opened, eBay is the only one the index describes in its evidence line as tracked across quarters.
None of this is grounds for marking the index down. It states for itself that it counts disclosure language, and it publishes both the formula and the worked example. There is one sentence to take from it. There are two yardsticks inside the same index, and the 2% that travelled and the published leaderboard are measured with different ones. How far an indicator built only on what documents say can be pushed is something we examined in Anthropic's robot automation forecast.
The only published basis for that 2% has British companies in it
Scraping and classifying listed-company AI disclosure end to end is not new ground for us. Two days ago we read a study that classified 9,821 UK annual reports. That one measures how specific the language of risk is; this section asks which population an index that counted the language of results drew its number from. Different thing counted, different country.
The index is explicit about its sample. It combines the S&P 500 and the FTSE 100 and then removes the "AI value chain" — chipmakers and companies selling AI models, platforms and software. That leaves 532 companies.
That number will not sit still either. The index homepage says 536, the Q1 report says 536, an analysis piece in between says 535, and the Q2 update of 28 September says 532. Four documents, three values. The site does explain why: the S&P 500 is refreshed quarterly and the FTSE 100 half-yearly, each with a different reporting period. Two membership lists pulled on different cycles and then combined give a total that shifts with the refresh date. So none of the figures is a typo; what it shows is that the denominator moves. This report uses the 532 from the 28 September release, because that piece is the most concrete — it writes denominators into its sentences, as in "361 of the 532 (68%)".
▲ Pebblous original diagram — the same index's own published sample count across four documents. The denominator itself is a moving number.
Here is where the chart and the number come apart. The y-axis on a16z page 27 reads "% of S&P 500 Companies" and the subtitle reads "Share of SP500 companies". But the index published exactly one basis for that 2%, and that one basis uses a denominator of 532 that includes British companies.
Three more companies reached L4… That brings the total to twelve, still just 2% of the sample. Source: The AI Value Gap, Q2 index update, 2026-09-28. The same piece describes the sample as "532 companies across the S&P 500 and FTSE 100"
What makes it odder is that the index splits the two countries on other rungs. The same piece reports the share that quantified a result separately: 43% for the FTSE, 30% for the S&P. The party with data it can split chose, for the fourth rung alone, to publish on the combined denominator. Apollo pushes the other way, describing its material as "Data from the S&P 500 second quarter earnings season". One document reads as US-only; the other counts Britain in.
So it is worth separating what this section can settle from what it cannot. This much is settled: the name on the chart and the number on the chart do not come from the same public document. One of the two has to give way, and nothing in the public record says which. What is not settled comes next.
- We cannot say the 2% is wrong. There is no basis for calling either figure wrong.
- Nor can we write that "all twelve are US-listed, so the real figure is 2.7%". That division would be our inference rather than a published value — and this time there is more than inference against it. The same Q2 update names RELX and Auto Trader as companies that track results with recurring KPIs, and both are FTSE 100. The fourth rung is not a US-only rung. It is true that all seven companies on 60 points are US-listed, but as section 3.1 showed, that table does not measure the fourth rung.
- Nor is it that the index is shoddy. It publishes its methodology and opens a verdict page per company. The usable fact stops at this: it is not official S&P 500 statistics.
One more version exists. Apollo's 11 September snapshot reads 74, 69, 29, 2, 0; the index operator's update for the same quarter, posted on 28 September, reads 76, 71, 32 — two to three points higher on every rung. On 30 September a16z printed the 11 September version. Whether more earnings calls landed in those three weeks, whether the sample was rebalanced, or whether the two pieces use different denominators to begin with cannot be settled from the public material. A different value already existed at the moment of printing. These are moving snapshots, not settled quarterly figures.
The values are not the only thing that moved. The index swapped out the ladder itself once. It used to grade by the type of metric disclosed; now it blends in whether a reader outside the company can follow the same figure quarter to quarter, and it filters out targets, pilots and one-off boasts. And it writes that change into its posts and keeps a separate changelog on the site. As section 7 will show, that is exactly what the US Securities and Exchange Commission asks of a company that discloses a metric: if you have changed the yardstick, say what changed and how. An indicator that scores companies on whether they kept the same yardstick changed its own and said so. That is not a flaw; it is the condition that let this report use the number at all.
What the source says where the chart says 'US households'
The consumer chart has a footnote too: "PNC Research, Internal Data". PNC is a large US regional bank, and its economics team pools card and account transactions from its own customers into a short monthly report called the Consumer Health Check. Open that report and the 2.2% on the a16z chart is right there. What differs is the name in front of the number.
The share of PNC households paying for a Gen AI subscription reached 2.2% in May. Source: PNC Economics Research, Consumer Health Check, 2026-06-15, p. 4. The body of that report writes "PNC households" twice
And yet the qualifier already drops off inside that same report. The figure that plots the number is titled "Share of households paying for Gen AI subscriptions". No "PNC". Nor is it only that one figure. Of the five figures in the June edition dealing with generative AI, four count households — share of subscribing households, year-on-year growth, by generation, by income tercile — and all four say only "households" in the title, with no "PNC" attached. The body attaches it twice; the figure titles drop it four times. The April edition titled the same figure the same way, so this is not a one-off in June. The first place the label widens is neither a16z nor Apollo. It is the gap between the original report's body text and that report's own figure titles.
PNC had also already written down, inside that report, the standard this piece would otherwise be applying from outside. These are the first two sentences of the methodology box on page 7.
The data is based on aggregated and anonymized selections of PNC data. The data may have a degree of selection bias due to selected populations and data availability. Source: same report, p. 7, Methodology. The items that follow note that card spending is based on a fixed cohort of retail customers and that categorisation follows the merchant category codes (MCC) common to the financial industry
You could read that as a one-line boilerplate disclaimer. In this context it is hard to. The comparisons that come later took that same sentence as a starting point and went on to build models and reweighting procedures. PNC wrote the sentence and stopped, and stopping is not in itself a fault. What PNC meant to describe was its own customers.
5.1The places where that name widened
The next step is the press, and that step comes five and a half months before the a16z deck. A CBS story on 17 April 2026 rendered the same figure as "only about 2% of all U.S. households". The deck appeared on 30 September. The table below sets out, in order, which name the one number wore in which place.
| Where | The name written there |
|---|---|
| PNC report, body text | PNC customer households ("the share of PNC households…") — twice in the body |
| PNC's own figure titles | Households ("Share of households…") — the qualifier drops in all four generative AI figures, and has since the April edition |
| CBS, 2026-04-17 | All US households ("only about 2% of all U.S. households") — 5.5 months before the a16z deck |
| a16z deck p. 38, 09-30 | US households ("Share of US Households with paid AI subscriptions") — PNC named in the footnote |
| a16z social card, 10-01 | "98% of US households aren't paying for AI yet" — no source. A positive statement flips into a negative one about the complement |
| Korean-language article, 10-05 | "Only 2.2% of US households subscribe to paid AI" — PNC is credited as the source, but the qualifier "PNC customers" is gone |
Each row was checked against the original document. The number holds; the label in front of it does not.
The sharpest sentence in this chain turns out to be inside the CBS story. To show how low the AI subscription rate still is, the story compares it with streaming — and the denominator switches mid-sentence.
The 2% of households that subscribe to generative AI services is far lower than the roughly 25% of U.S. consumers who pay for monthly streaming subscriptions. Source: CBS News, 2026-04-17. The first figure counts households, the second counts consumers
The 2% in front is a share of households and the 25% behind it is a share of people. Two numbers in different units, joined by "far lower". And the person making the comparison is the chief economist at the bank that owns the panel. This is not something to prosecute. It is what people do whenever they talk about numbers, which is why it is the subject of this report. Even the person who knows the data best switches denominators inside a single sentence.
5.2What the 98% erases
Something else drops off besides the qualifier. The flat single sentence of 98% erases the gradient inside it. Page 5 of the PNC report splits the same data by income and by generation. Upper-income households pay at over 4%, middle income at about 2%, lower income at under 1%. By generation, Gen Z, millennials and Gen X sit at 3 to 3.5% while baby boomers are around 1%. What subscribing households spend per month has also risen from about $22 two years ago to about $31. Who is paying is already stratified, and "more than nine in ten don't pay" flattens that distribution into one layer.
One more thing gets erased. PNC's two figures put May 2025 and May 2026 side by side, and every bucket roughly doubled over the year. By generation: 1.5 to 2.9, 1.4 to 3.2, 1.4 to 3.4, 0.4 to 1.0. By income: 0.4 to 0.8, 0.9 to 1.7, 2.0 to 4.1. In ratio terms the doubling is even; in gap terms it is not. The upper-income bucket gained 2.1 points, the lower-income bucket 0.4. What this data supports is therefore closer to "growing fast, with the growth concentrated at one end" than to "not used yet". The single figure of 98% erases both sentences. In the same place, PNC even plotted the year-on-year growth rate separately, and the vertical axis on that figure runs to 200%.
The other charts on the same slide share the property. Page 38 holds, besides the two PNC transaction panels, charts drawn from e-receipts and desktop traffic records, and the y-axis on one of them reads "Total Observable Subs" — the number of subscriptions that can be observed. The data side wrote honestly on the axis what it is looking at. What widened was the title, not the axis. And not one chart on that slide measures anything by survey.
5.3The prose misread the tick marks on its own chart
There is one more discrepancy, of a different kind, so it goes separately. a16z's summary post writes "As of April" where the PNC report says "in May". Render both charts at high resolution and line the bars up pixel by pixel and the explanation appears. In both charts the x-axis ticks are spaced by quarter while the bars are monthly, so the highlighted final bar stands unlabelled in the slot just past the last tick, "Apr '26". In other words, both charts plot the May value, exactly matching PNC's body text. What went wrong is not the chart but how the prose describing it read the tick marks. The preceding three issues are about who was counted; this one is about how someone read their own picture, and lumping them together would blur both.
Step back from all of this and the two charts turn out to have the same shape. The diagram below puts the two chains side by side. The top row is the corporate side, the bottom row the consumer side, and under each box is only the wording that changed at that step.
Every box in both chains is a public document. At each step the number survives; the name of the group it points to does not. At the last box the source attribution goes with it.
The four answers to one question count different things
Ask the one-sentence question of how many people in the US pay for AI and four substantial sources published in 2026 give four different answers: 2.2%, about 3%, 55% and 41%. Before setting them side by side, here is each one's denominator. Denominators diverging under a single metric name is something we took up in recounting AI usage statistics against an independent corpus.
| Source | What it counted | Denominator | Value | As of |
|---|---|---|---|---|
| PNC Economics Research | AI subscription payments observable in card and account transactions | PNC customer households | 2.2% | 2026-05 |
| Bank of America Institute | AI service payments observable in payments data | Active BofA customer households | about 3% | through 2026-02 |
| Menlo Ventures survey | Self-reported "I pay for at least one AI product" | US AI users | 55% | fielded 2026-07 |
| Federal Reserve survey | Self-reported "I use generative AI for work" — not payment | Individuals aged 18–64 | 41% | fielded 2025-11 |
The units are mixed. The top two count households, the bottom two count people. And the bottom figure asks about use, not payment.
The good news first. The two bank panels agree with each other. The 2.2% and the roughly 3% come independently from different customers at different banks and land in the same place, which leaves little reason to doubt the panel figures themselves. One caveat attaches to the convergence, though. A Federal Reserve Board study of supervisory card data notes that single-bank spending series can be volatile even when the aggregate tracks well. One PNC and one BofA are exactly that single bank. Two figures agreeing is not permission to call either of them "US households".
The size of the denominator is written down on only one side. Bank of America does not publish its panel household count, only its inclusion rules: at least five transactions a month, an account held since January 2019, business cards excluded, a fixed cohort. On the PNC side a number does surface once. The same June report, explaining households receiving unemployment benefits, refers to "the 4 million household cohort PNC tracks". Whether that 4 million is the cohort the generative AI share came from, the report does not say. The methodology box lists card spending and balances as separate fixed cohorts; it does not pair each figure with its cohort. Either way, the rules themselves are the bank writing down that the denominator is not all US households.
▲ Pebblous original diagram — the same four figures regrouped by denominator (households vs. people). The dashed line is why this report does not compute a multiple between them.
6.1A multiple between the four figures mixes units
Look at the table and you want to ask how many times 2.2% goes into 55%. We do not do that division. The denominator of 2.2% is households and the denominator of 55% is people, so dividing one by the other yields a number in mixed units. The direction tangles too. A household counts if any one person in it pays, so household shares usually come out higher than individual shares; here it is the other way round. The multiple is therefore not the size of a gap but evidence that the two numbers do not sit on the same axis. A piece arguing for denominators ends the moment it mixes denominators in its own body.
The direction of the gap still matters. The usual finding in the consumption measurement literature is that surveys capture less than transaction records: people remember and report less of what they spent. Here the survey side is overwhelmingly higher. That literature measured spending amounts while what diverges here is whether someone pays at all, so the two do not overlay directly — but ordinary under-reporting does not explain the direction of this gap. The distance between surveys, averages and individuals is something we examined in validating models asked to answer surveys in people's place.
A good part of the gap is explained by the survey itself. Among those in the Menlo tally who said they pay, only 48% pay the full amount themselves; 34% have family or friends paying for them, and about 20% have an employer or school paying (multiple responses allowed). What a bank panel sees is money leaving that household's card for an AI company, not whether someone in that household pays for AI. Those are different questions. The payment channel adds to it. App store purchases and bundled subscriptions show up on a card statement as 'Apple' or 'Google' rather than the AI company's name. That is not a guess; it follows directly from a rule PNC wrote into its methodology box, which states that spending is categorised using the merchant category codes common to the financial industry. Group by merchant and what you bought through Apple becomes Apple. The US Bureau of Labor Statistics raised the same point as a limit of transaction data. How large that share is, in numbers, we did not find in a source during this research. We note the absence and move on.
6.2Menlo's two figures were never the same kind of thing
Menlo Ventures' 2026 report records this year's 55% and recalls last year's figure like this: "Last year, AI looked no different: Just 3% of AI users paid." Open the 2025 edition, though, and that 3% is not a survey measurement.
1.8 billion users at an average monthly subscription cost of $20 per month equals $432 billion a year; today's $12 billion market indicates that only about 3% pay for premium services. Source: Menlo Ventures, 2025: The State of Consumer AI
That is revenue divided by a hypothetical ceiling, and the denominator is users worldwide. The 55% in 2026 is a survey value from 5,067 US adults, with US AI users as the denominator. So the account of going from 3% to 55% is not the same thing measured twice. It is also a different kind of mismatch from the ones in the preceding sections. Those were cases where who was counted changed; this is a case where what was counted changed. Filed under one name, both get blurred. And the two cannot be divided into each other anyway, one being global spending and the other a US survey. The word to avoid closing on here is contradiction. What belongs here is that the two do not divide.
One more interest to declare. Menlo Ventures is a venture capital firm too, and its portfolio companies appear in the report. Noting a16z's position and omitting this one would be unfair.
6.3The price of moving 'our customers' to 'US households'
How big a job it is to take a bank's customer data national is not something anyone needs to guess at. An institution that set out to do it wrote down what it did in a technical note: the JPMorgan Chase Institute.
We aim to publish generalizable insights that are representative of the overall US population. To do this, we require a method to reweight research based on key characteristics, with income foremost among them. Source: JPMorgan Chase Institute, Estimating Family Income from Administrative Banking Data, 2018-12
So what did it do? It expanded 400 raw variables into roughly 800 candidates, built a ground truth set of 250,000 people a year stratified by the income quintiles of a national survey, and fitted a machine learning model. Even then, the paper records a mean absolute error of 41% on the income estimate and a 55% rate of placing a household in exactly the right income quintile, rising above 90% if adjacent quintiles count. It also states that the ground truth set itself skewed towards higher incomes, and pins the use down: these estimates are for analytical and research reweighting, not for business decisions.
The Federal Reserve Board goes a step further. Because it works with card data collected from many banks at once for supervisory purposes, it can ask whether single-bank data is representative by calculation rather than by guess. It plotted the distribution of bank-level series over the aggregate series. The answer is mixed. For most of the period the bank-level values cluster tightly around the aggregate, but around 2016–2017 and again around 2022 the 10th-to-90th percentile band visibly widens. In some periods, which bank's data you use changes the movement of consumption meaningfully. The caution the paper then writes down is the same sentence as this report's subject.
Researchers using data from individual banks, fintech platforms, or payment processors should therefore validate their spending measures against external benchmarks and, to the extent possible, assess whether their results are driven by particular customer segments or geographic markets. Source: Federal Reserve Board, arXiv:2607.08759v1, §7.3 "Bank-Level Heterogeneity and Sample Representativeness"
Who that caution is addressed to matters here. Whoever calls their own customers their own customers has no occasion to benchmark against anything external. The need for that arises only where the number acquires the name "US households" — and the place where the name is applied is not the place that holds the data. The party that applied the name has no data to validate against, and the party with the data has no reason to validate.
In January 2026 the US Bureau of Labor Statistics reached a conclusion about transaction data of the same kind. It is not geographically representative, it captures only 10% of all card transactions, it is organised by merchant type rather than product type, and it carries no demographic information. So it closed with this: useful for trends, but not a substitute for the Consumer Expenditure Survey. That is an official statistical agency's ruling in January, and the chart in question appeared in September.
The axis of this comparison is not which institution is better. The JPMorgan Chase Institute also sits under a bank and also works from proprietary data nobody can reproduce. There is one axis: what you have to do before calling something national. One side went through that procedure and published its own error; on the other side the step was skipped entirely while the name widened. And PNC never claimed to be speaking about the country. The part of a panel that demographic reweighting does not reach shows up again in behavioural representativeness in mobility data.
The rules put up a tollbooth rather than a roadblock
There is an explanation running the other way for the zero on the top rung in section 3. It may not be that companies are lazy but that the rules are shaped that way. This section examines that possibility separately. The conclusion up front: half right. The rules do not block it, they do not require it, and the moment you volunteer it they put a price on it.
7.1Nothing blocks it
The 2020 interpretive guidance on Management's Discussion and Analysis from the US Securities and Exchange Commission carries, in a footnote, sixteen examples of the kind of metric companies disclose voluntarily: operating margin; same store sales; sales per square foot; total customers and subscribers; average revenue per user; daily and monthly active users; active customers; net customer additions; total impressions; number of memberships; traffic growth; comparable customer transactions increase; voluntary and involuntary turnover rate; workforce composition percentages; and then total energy consumed and the number of data breaches. Non-financial technical and operating metrics are squarely in scope, and the list itself is left open with "include, but are not limited to". None of the sixteen, though, is an AI item. The road is not blocked; the name simply has not been called.
There is a precedent too. In the first quarter of 2015, Amazon changed its reportable segments to North America, International and AWS — pulling cloud out as a line of its own. And it recast the prior figures.
Beginning in the first quarter of 2015, we changed our reportable segments to North America, International, and AWS. … Certain prior period amounts have been reclassified to conform to the current presentation, including recasting the segment financial information. Source: Amazon.com, Form 10-Q, 2015 Q1, Note 8
The index defines its fifth rung as reporting the figure as its own line, in the same place every quarter, so it can be followed from outside — and what Amazon did in 2015 is precisely that. It broke the line out, recast the history, and reported it in the same place every quarter afterwards. The reason it gave was not a regulation but an internal one: the way the company evaluates and operates its business had changed.
7.2Nothing requires it either
The same guidance pins down the nature of metric disclosure in a single word: voluntary.
Some companies voluntarily disclose specialized, company-specific sales metrics, such as same store sales or revenue per subscriber. Source: SEC Release No. 33-10751, 2020-01-30
Breaking a segment out works much the same way. Segment disclosure under US accounting standards follows the management approach: what becomes reportable is the line the chief operating decision maker already receives on a regular basis. If a company does not run an AI profit and loss internally, the standard is silent. Accounting standards cannot create an AI line. They only follow one into existence.
7.3Writing it down attaches a price
So is volunteering it the end of the matter? No. A list follows the moment you write it down, and that list is the most important thing in this section. The guidance says it would generally expect a disclosed metric to be accompanied by a clear definition of the metric and how it is calculated, a statement of why it provides useful information to investors, and a statement of how management uses it in managing or monitoring the business. It adds that any estimates or assumptions underlying the metric should be disclosed too. And changing the method of calculation gets a passage to itself.
If a company changes the method by which it calculates or presents the metric from one period to another or otherwise, the company should consider the need to disclose, to the extent material: (1) the differences in the way the metric is calculated or presented compared to prior periods, (2) the reasons for such changes, (3) the effects of any such change on the amounts or other information being disclosed and on amounts or other information previously reported, and (4) such other differences in methodology and results that would reasonably be expected to be relevant to an understanding of the company's performance or prospects. Depending on the significance of the change(s) in methodology and results, the company should consider whether it is necessary to recast prior metrics to conform to the current presentation and place the current disclosure in an appropriate context. Source: same guidance, SEC Release No. 33-10751
One more thing attaches. For metrics derived from the company's own information, the guidance says the company should consider whether it has effective controls and procedures over the process that produces the number, to ensure "consistency as well as accuracy". Read the list again and a familiar shape appears. Fix the definition; if you change the yardstick, say what changed and how; recast the history if necessary; put controls on the process that makes the number. That is the same thing the fourth and fifth rungs of the ladder ask for. The distance between putting a number on a result once (29%) and using the same measure again the next quarter (2%) is not a volume of data but the weight of that list.
So this section lands on neither of the two obvious conclusions. Not "the companies are therefore blameless", and not "not disclosing is hiding". The rules did not block the road; they put a tollbooth on it. Twelve companies out of 532 have paid the toll, and none has yet gone the whole way.
There is one further explanation outside the rules, and it is the other side of the toll: is there anything to be had for paying it? The index operator measured that directly. Holding sector effects fixed and controlling for size, growth and margin, he regressed sector-relative forward price-to-earnings on a talk score and on a proof score. All three results point the same way. Companies that can prove AI is producing results do not trade at a premium to their peers; the premium for talking a lot sits inside the noise; and over the past three years, quarters in which a company increased its AI mentions neither beat nor lagged the market. This is one analyst's regression, so it is not to be read as settled. But if the estimate holds, it adds one more explanation for the empty fifth rung. Fixing the definition, recasting the history, putting controls on the process — none of it is currently priced.
Outside listed-company disclosure, examples of the toll being paid are easier to find: an internal scorecard that sets rework rate next to AI spend per employee and measures both on the same basis each quarter, for instance. Rippling putting employee AI spend and performance in one table is close to that shape. Where there is no disclosure obligation, fixing the definition costs far less.
Why This Matters to Pebblous
An interest to declare first. Pebblous does not disclose the effect of its own AI adoption at the fourth-rung standard of this ladder either. This is a piece about reading other people's disclosure, so we start by recording where we stand.
What broke is not a value but the name of a column
What Pebblous does in DataClinic is ask of an incoming dataset what it is a count of. AI-Ready Data puts the same question at the stage before training. What this episode shows is that the question breaks first not inside the dataset but on its label. PNC's customer household data is clean. No missing values, no duplicates, no outliers. The column's name is what broke: it stretches with every hand it passes through.
Survey methodology already has a name for this mismatch between the population you can actually reach and the population you mean to describe. The name matters less than the sentence the literature attaches to it: the mismatch exists before the sample is drawn. Enlarging the sample does not shrink it, and neither does any amount of cleaning. Quality checks look at values, not at names. Which is why a defect of this kind survives intact in a dataset with a perfect quality score.
The ladder's fifth rung asks whether you can measure it again next quarter
Reread the fifth rung's definition: reported consistently each quarter so an outsider can track the same figure over time. The condition written down for a good metric is not accuracy but whether it can be measured again. And the list the guidance in section 7 asks of a voluntarily disclosed metric has exactly the same shape. Write the definition and the calculation; if you change the yardstick, say what changed and how; recast the prior figures if needed; put controls on the process that makes the number. That is what the data quality literature calls reproducibility and lineage, and it touches where the international data quality standards (the ISO 5259 family) deal with measurability. The number of companies meeting that condition is 0 out of 532.
Three things you can do today
First, hold your current AI performance reporting up against the five rungs. Most of it will sit on the second or the third, and what divides the third from the fourth is not the volume of data but whether the definition was fixed and preserved. What to fix is already written out in the list in section 7.
Second, ask for the denominator before citing an outside number. "A percentage of what" comes before "what percentage". The table in section 6 is the exercise. All four figures are correct, and they are not counting the same thing.
Third, audit the labels on the numbers you send out. Are you calling a rate from your own customer sample "the industry"? What maps onto this directly is that in this episode the first place the name slipped was the original report's figure titles. Attaching the qualifier in the body and dropping it in the chart title happens in every organisation.
A number being right and a number belonging to that name are different things
The argument about AI investment has split along whether it works. What this case shows sits one step earlier. Whether the yardstick that measures the effect survives into the next quarter has to be settled before that argument can even be had. At the moment both sides hold the panel that suits them, and neither side is wrong. The distinction is what Pebblous can put into Korean first. A number being right and a number belonging to that name are different things. And the check is not hard. Follow the footnote up one level.
Every verbatim quotation in this report was checked directly against the original a16z deck PDF, Apollo's note, four public posts by the index operator plus the index site's per-company pages, the full PNC report PDF, both Menlo Ventures editions, the full text of the SEC guidance, Amazon's filing, and the JPMorgan Chase Institute, Federal Reserve and Bureau of Labor Statistics documents. Chart figures were re-read off the deck pages at 600dpi, down to the footnote text. Whether the sample is 532 or 536 is explained by the refresh cycles the index site states, and section 4 says so. Several things remain unresolved: whether Apollo's 69% is US-only or includes Britain; the full list of the twelve on the fourth rung; whether PNC's 4 million household cohort is the cohort behind the generative AI share; the share of payments a bank panel cannot see; and what the remaining 8% of the 70% cost figure consists of. We could not open the segment accounting standard itself, so we describe its intent without citing the text, and we did not observe the view count on the social card, so we do not use it. Sections 1 through 7 are what the public documents report, with one exception; the later part of section 6.3 and this section are the part those documents do not cover, so please read them separately. Thank you for reading this far.
References
The subject of this report — primary
- 1.a16z, State of Markets II, 2026-09-30. A 90-page slide deck. The body quotes pages 27 and 38; page count and chart figures were confirmed by rendering the original PDF at 600dpi.
- 2.a16z, State of Markets II summary post — restates both figures in prose with no source attribution.
- 3.a16z post on X, 2026-10-01 18:30 — "98% of US households aren't paying for AI yet". Direct access was blocked as of 2026-10-08, so the view count could not be observed and is not used in the body.
- 4.Torsten Slok, Apollo Daily Spark, 2026-09-11 — "69% of the S&P 500 point to a live AI deployment", cost 70% and revenue 22%. The remaining 8% does not appear in the source.
- 5.The AI Value Gap (AI Value Realisation Index) homepage and per-company verdict pages (
/c/IRM·/c/EBAY·/c/SPGI·/c/XYZ·/c/TTD), accessed 2026-10-08 — headline sample of 536, the five rung definitions verbatim, the Proven Value formula and worked example (3+3+3=9/15→60), "Only Level 3 and above counts", refresh cycles (S&P quarterly, FTSE half-yearly), the statement that Substantiation is not part of the index, the AvAI companion index (assisting side 55%; 52% at rung 3 and above versus 61% below, across 1,410 cases over two years), and Iron Mountain's FY2025 $540M. These are live figures, so the access date applies. - 6.Amin Mrini, No.44 — Q2 Index Update, 2026-09-28 — "532 companies", "twelve, still just 2% of the sample", index at 15.7/100 (QoQ +17%), the 76/71/32 ladder, "361 of the 532 (68%)", the seven companies on 60 points by name, RELX, Auto Trader and FactSet tracking via recurring KPIs, S&P 30% versus FTSE 43% and its cost skew, and the Booking Holdings and Block cases.
- 7.Amin Mrini, Q1 2026 AI Swimsuit Report, Q1 2026 — index at 13.5/100, "across 536 large US and UK companies", the claim-level 8%, Block's 12 of 219, and the sector distribution with its stated reason. · No.35 — AI talk is cheap, AI proof is scarce — how the ladder was revised ("a far harder version of my first attempt in No.33") and the three regression results controlling for sector fixed effects, size, growth and margin.
- 8.PNC Economics Research, Consumer Health Check, 2026-06-15 — "the share of PNC households… reached 2.2% in May". Full text verified from a local copy: "PNC households" twice in the body; Figures 11, 12, 14 and 15 on generative AI all omitting "PNC" from their titles; the generation and income breakdowns against May 2025; "the 4 million household cohort PNC tracks" (p. 3); and, on p. 7, the Methodology line "The data may have a degree of selection bias due to selected populations and data availability" together with the MCC categorisation rule.
- 9.PNC Economics Research, Consumer Health Check, editions of 2026-04-13, 2026-07-13 and 2026-08-10 — confirming that the April edition titles Figure 13 the same way the June edition does, and that the July and August editions carry no generative AI figures. What the "July 13, 2026" in the a16z footnote refers to cannot be settled from the public material, so the body does not take it up.
- 10.Menlo Ventures, 2026: The State of Consumer AI, 2026-09-15 — 55%, 5,067 US adults (Morning Consult, 2026-07), paying in full themselves 48%, family 34%, employer 20% (multiple responses).
- 11.Menlo Ventures, 2025: The State of Consumer AI, 2025-06-26 — the calculation behind the "3%" (1.8 billion users worldwide × $20 a month × 12 months, against a $12 billion market).
- 12.Bank of America Institute, Consumer Checkpoint (methodology appendix), 2026-03 — about 3%, only customers with at least five transactions a month, business cards excluded. The panel household count is not published.
- 13.Deloitte, Digital Media Trends, 20th ed., 2026-03-25 — about 90% self-reporting a paid video subscription (US consumers aged 14+, n=3,575). Being a self-report survey, it is not set directly against the card panel figures.
Methodology and rules — primary
- 14.U.S. Securities and Exchange Commission, Commission Guidance on Management's Discussion and Analysis, Release Nos. 33-10751 / 34-88094 / FR-87, 2020-01-30 — voluntary disclosure, the accompanying disclosure list, the four items and the recasting question when the method changes, internal controls, and the sixteen examples in footnote 11.
- 15.Amazon.com, Inc., Form 10-Q, 2015 Q1, Note 8 — the segment change and the recast of prior figures.
- 16.JPMorgan Chase Institute, Estimating Family Income from Administrative Banking Data: A Machine Learning Approach, 2018-12 — mean absolute error of 41%, income quintile accuracy of 55%, 400 raw variables, a ground truth set of 250,000 people a year.
- 17.Aladangady, A., Duque Gabriel, R., & Wix, C. (Federal Reserve Board), Measuring Consumption with Credit Card Data: Benchmarking and Beyond, arXiv:2607.08759v1, 2026-07 — §7.3 "Bank-Level Heterogeneity and Sample Representativeness": the observation that the 10th-to-90th percentile band of bank-level series widens around 2016–2017 and 2022, and the verbatim caution to users of individual-bank data.
- 18.Erhard, L., & Montag, H., The value of transactional data to a consumer spending survey, BLS Monthly Labor Review, 2026-01-26 — lack of geographic representativeness, 10% of all card transactions captured, organisation by merchant type, and the finding that it cannot substitute for the survey.
- 19.Sen, I., Flöck, F., Weller, K., Weiß, B., & Wagner, C., TED-On: A Total Error Framework for Digital Traces of Human Behavior on Online Platforms, arXiv:1907.08228 — the statement that the mismatch exists before the sample is drawn.
- 20.Groves, R. M. et al., Survey Methodology (2nd ed.), Wiley, 2009 — ⚠️ we did not open the book itself and used only the sentence quoted in reference 18 above.
- 21.FASB, ASU 2023-07, Segment Reporting (Topic 280) — ⚠️ we could not open the text, so we describe the management approach as intent without citing the provision.
The chain of transmission — secondary
- 22.Megan Cerullo, CBS News, 2026-04-17 — "only about 2% of all U.S. households", and the comparison sentence that switches from households to consumers mid-sentence. Quotes PNC's chief economist by name.
- 23.Conor Murray, Forbes, 2026-09-11 — reported the same figures 19 days before the a16z deck.
- 24.techstartups.com, 2026-10-02 — a rare case that notes in the body that the two figures come from different populations.
- 25.The Neuron, 2026-10-02 to 10-04 — combines the two charts in a single story.
- 26.AI Times (Korean), 2026-10-05 — "only 2.2% of US households subscribe to paid AI". Credits PNC as the source but drops the qualifier.
Company disclosure — primary
- 27.S&P Global Q2 2026 earnings call (2026-07-28) · eBay Q1 2026 (2026-04-29) and Q4 2025 (2026-02-18) · Iron Mountain Q1 and Q2 2026 earnings calls — the three cases in section 3.1. Iron Mountain's supporting figure on the index page is FY2025 digital revenue of $540M, an annual figure, which is why the 2026 quarterly calls carry nothing of the kind.
Related Pebblous Blog pieces
- 28.Denominators diverging under one metric name (section 6) · Classifying every UK listed-company annual report (section 4) · Scoring built only on what documents say (section 3.1) · The panel gap demographic reweighting does not close (section 6.3) · Models asked to answer surveys in people's place (section 6.1) · A scorecard that sets performance next to AI spend (section 7.3) · Deciding what to count (section 3) · Earlier material from the same firm (section 1) · Anthropic's forward revenue multiple · The distance between adoption and results