Executive Summary
Song Yao, a professor at Washington University in St. Louis, went through Kaggle's public record and measured how long a medal keeps predicting the performance that follows it, across 444,698 contest entries. The answer was one year. A medal less than a year old predicted the next result clearly, and anything older predicted almost nothing. The pattern was already in place long before ChatGPT arrived.
The part of the story that generative AI seems to own is the upload-format medal. That stock lost 82% of its predictive power in the AI era, but once the author broke the loss apart, somewhere between half and three quarters of it had nothing to do with AI making medals cheap. The contest format itself had already been retired, no new medals were being issued, and the remaining stock simply aged at its usual rate.
No evidence turned up that the people winning medals had changed. What changed sat on the institutional side, in how the signal is issued and displayed. So when you read a credential on a résumé, when it was earned matters more than what it was.
Key Numbers
Source: Song Yao, arXiv:2608.17111 (2026-08-17)
10 to 12 pts
The gap a fresh medal opens
A competitor holding one medal less than a year old finished that many leaderboard percentile points above one holding none
99% · 95%
Explanatory power sitting in the first year
Share of the performance spread that the under-one-year bands capture on their own, before AI and in the AI era
82%
Predictive power lost by upload medals
Over the same span, code-format medals moved the other way, from 0.065 up to 0.087
43 to 87%
Share of that loss explained by aging
43 to 62% under the attribution least favorable to aging, 64 to 87% under the most favorable one
Where the Belief Came From
The claim that generative AI destroyed the value of credentials rests on real evidence. On freelance platforms, proposals started reading better while telling clients less about the person who sent them. In unproctored exams, scores inflated. In both cases the reason was the same: unverified words became cheap. Once a polished cover letter and a clean exam answer are things anyone can produce in seconds, the habit of picking people by reading them starts to wobble.
A Kaggle medal is not that kind of signal. A competitor submits predictions, those predictions are scored against answers nobody can see, and the ranking comes from a direct comparison with everyone else working the same problem. Polished writing does not carry anyone to the top of a leaderboard. Generative AI did make plausible-looking submissions cheap, but it did not make a top finish cheap.
So the question this study asks is a narrower one. Does a verified signal survive the AI era? The author's answer is close to yes, though it is not the reassuring kind of yes. Something else was eating the signal, and verification could not stop it.
Why Kaggle Works as a Laboratory
The decisive condition is that Kaggle ran two scoring regimes side by side for years. In an upload-format competition, a competitor submits predicted values and nobody checks how those predictions were produced. In a code-format competition, the competitor submits code, and the platform executes that code on data it keeps hidden before assigning a score. Scoring that reads only the output and scoring that runs the process coexisted on the same leaderboards, among the same people.
The record is deep enough for that comparison. Kaggle has logged 18.67 million contest submissions since 2010, and it awards medals along with lifetime tiers named Expert, Master and Grandmaster. Plenty of data scientists put those tiers on a résumé or a profile. The panel the author analyzed covers 214 medal competitions that closed between 2018 and 2025, with 197,561 competitors and 444,698 entries. The quarter in which ChatGPT was released, the fourth quarter of 2022, was treated as a boundary and left out of both eras, and the years from 2010 through 2017 entered the study only as the medal history each competitor had accumulated by then, never as performance.
Performance is measured as the final leaderboard percentile, and the predictor is the count of medals a competitor had accumulated before entering that contest. Medals are sorted into bands by how long ago they were won and by which format produced them. The slope of each band is how much information that medal carries about performance.
Some things were filtered out while building the panel. Three formats whose standings depend on data generated after the deadline were dropped: simulation competitions where competitors' programs play against each other, two-stage competitions where the final test data arrives after the close, and forecasting competitions that score predictions about events yet to happen. Seventy-three percent of entries came from solo competitors, and team results were credited equally to every member, though the conclusions held when the panel was restricted to solo entries or when results were credited only to the person who actually submitted.
A Medal Lasts One Year
For medals less than a year old, the slope ran between 0.15 and 0.18 per log medal. In plain terms, a competitor with one such medal finished 10 to 12 leaderboard percentile points above a competitor with none. For medals past their first year, the slope drops to somewhere between 0.01 and 0.03, which is not meaningfully different from having nothing at all.
Measured as information, the concentration is even starker. When medals of every age and both formats are used together to predict performance, the two under-one-year bands alone deliver 99% of what the full set explains. Recomputed on AI-era results, the figure is 95%. Adding every remaining band back in leaves very little on the table.
The decay took the same shape in both formats, and it survived controls for the other medals and the participation history a competitor carried. That is a different mechanism from the standard labor-economics story, in which a degree loses value because the employer watches the worker and learns what they can actually do. There is no watching employer here, and a year was enough for the value to go.
Where Did the 82% Go?
In the AI era the two kinds of medals moved in opposite directions. On pre-ChatGPT results, upload-format and code-format medals had slopes of 0.067 and 0.065, which is effectively the same number. On contests that closed from 2023 onward, the upload slope fell to 0.012 while the code slope climbed to 0.087. Against its own baseline, the upload figure is an 82% loss.
Read only that far, it looks like a story about unverified predictions becoming untrustworthy because of AI. The author reaches first for the evidence pointing the other way. If AI had shaken the whole field, every signal should have moved in the same direction, yet two stocks moved oppositely inside the same competitions. The same fall and rise reappeared across 344,181 entries in community competitions, which award no medals at all: the upload slope dropped by 0.046 with a P value of 0.001, and the code slope rose by 0.074. Whatever this is, the medal label is not producing it.
The decisive fact is in the mix of competitions. In 2016, Kaggle ran 26 upload-format medal competitions and zero code-format ones in a year. By 2024 that had flipped to 3 and 24. The platform moved toward execution-based scoring because of unverifiable submissions and leaderboard gaming, and the move was underway several years before generative AI arrived. Upload medals became a stock that nobody was replenishing, and that stock aged along the decay curve from section 3.
The decomposition runs like this. Of the decline, the share created by the stock growing older accounts for 43 to 62% under the attribution least favorable to aging, and 64 to 87% under the most favorable one. Every bootstrap draw left that component the largest. The explanation that AI repriced predicted values only fills in what remains. The author calls the result a stranded asset. A power plant stops running not because it got worse but because the grid changed, and here it was the institution issuing the medals that exited, not the ability of the people holding them.
None of this means the AI effect was zero. Pulling apart the 0.055 the upload stock lost, the largest piece, 0.028, comes from the stock's age mix shifting older, while the slopes of the bands themselves fell by 0.015, a quarter of the total. Looking only at fresh upload medals won while the format was still issuing them, roughly a quarter of the slope really was shaved off in the AI era: it fell by 0.033 from a base of 0.122, with a P value of 0.002. Fresh code-format medals show no such decline under the same test. Verification did not fully shield the signal from AI, and the size of the effect is far smaller than the number 82% suggests.
There were also attempts to find the change in the competitors themselves. The author registered two predictions before running the analysis: that a working style typical of AI users would produce different results in code contests than in upload contests, and that more recent entrants would fail execution more often. Both missed. The working-style index came in at 0.006 with a P value of 0.28, and the comparison held within the 2,431 people who competed in both formats. Execution failures did look like they were rising in the raw yearly data, but the increase disappeared once contests were held fixed. Failures went up because the contests got harder, not because a different crowd showed up.
There is one more pattern, in which old upload medals appear to have become more valuable in the AI era, and it turns out to be an artifact of how they are read. Counting badges alone makes the value look like it went up, but the effect vanishes once the person's other medals and participation history enter the model. An old medal was standing in for a different signal, namely a long career. Survivorship was checked as well. Among pre-AI competitors, only 7.8% of those without a medal reappeared in the AI era, against 56.5% of those with six or more, and that gap is wide, yet restricting the analysis to people who competed in both eras left the results unchanged.
What the Tier System Throws Away
Kaggle's official tiers are a deterministic function of lifetime medal counts. Expert requires two medals, Master requires one gold plus two more of silver or better, and Grandmaster requires five golds including one won solo. When a medal was earned and which format produced it never enter the calculation, and a running total collapses into four ranks. The author reconstructed the rule closely enough to match 97.7% of the actual tiers, then built an index from the same medal histories that weights recent medals more heavily, and set the two side by side. If the decay in section 3 is real, a total that ignores age has no choice but to discard information.
The index predicted performance 15% better than the tier before AI, and 19% better in the AI era. Turned around, the official tier throws away 13% and 16% of the information already in hand. One objection is that the index simply had more parameters to work with, and the author answers it twice. Entered into the same regression, the explanatory power the index adds on top of the tier was three to four times what the tier adds on top of the index, and the gap held when the sample was limited to users whose medal dates are complete enough for the tier reconstruction to be exact.
The next result is the more striking one. An index whose weights were learned on pre-AI data and then frozen still beat the official tier at predicting AI-era performance. The explanatory power was 0.053 against 0.047, and the bootstrap resampled at the competition level put the 95% interval for the difference between 0.004 and 0.009. The platform did not need any information from after the AI transition in order to display its signal better.
The author closes the paper with four properties of a credential. It is informative, it is perishable, it is institution-bound, and it is interdependent with other signals. The first is the answer this paper gives to the popular belief, and the design lessons come out of the other three.
- • Credentials perish. Because the predictive power sits in the first year, any display that shows a lifetime total structurally overvalues old signals. This is not specific to Kaggle tiers; badges on professional profiles work the same way.
- • Credentialing systems die when issuance stops. When a scoring format exits, the outstanding stock ages at its usual rate even though nobody holding it got any worse. The governing question becomes which formats an institution can keep running.
- • Signals live inside a portfolio. When one of them stops being issued, the statistical weight moves to the signals next to it. An audit that inspects signals one at a time will misread that movement as a change in the signal itself.
The author draws his own lines around the result. Everything measured here is an association observed inside a single scoring system, and the AI-era estimates rest on the minority who kept competing. The decay rate found on Kaggle cannot be copied onto another credentialing system. What travels is not the rate but the question.
Editor's Note: Pebblous meets the same question in data assets. A well-built labeled dataset or evaluation benchmark earns its credibility not from the date it was created but from the date it was last refreshed. Just as a credential quietly ages once issuance stops, data that stops being updated loses value while the people using it notice nothing. Metrics that count the size of the stock cannot show that aging.
The original paper is available at arXiv:2608.17111.