Executive Summary

Seven in ten HR leaders in the United States say that reviewing AI-generated applications has slowed their hiring down. Applying has never been easier, and the work of choosing has only gotten harder. A hiring model posted on September 24 by researchers at Stanford and Cornell explains that mismatch in one line. AI lowers the cost of writing an application and, at the same time, lowers how much information the application carries. This article looks at where the model says the heaviest loss lands.

The model sorts applicants four ways, by whether they have visible experience and by how well they fit the job. As applications blur, a firm leans away from what an application says and toward experience, a coarse but solid marker. The group whose score falls furthest is the one that fits the job and has no experience to show. They lose the individual information, and they have no track record to make up for it. There is also a market where the same thing has already been observed. After Freelancer.com opened a tool that drafts cover letters, the correlation between how well a bid was tailored to the job and the chance of a reply fell 51%.

Sections 1 through 5 follow what three papers and one survey actually built and measured. The passage at the end of section 5 that sets the survey's added interview rounds alongside the model's intermediate evaluation, together with section 6, is this article's own reading.

Key Figures

Sources: Cui, Dias and Ye, Signaling in the Age of AI (arXiv:2509.25054v2) · Robert Half, release of March 10, 2026.

+51%

Callback rate for bids using the tool

A rise of 3.56 points from a base of 7.02%. It shows up only at the 10% significance level, and it fades after about two months

−51%

Correlation of tailoring with callbacks

Same period, same data. The fact that an application was written to fit the posting lost half its power to predict a reply

+5%

Correlation of callbacks with reputation score

Where firms moved instead. The platform builds this score from a worker's past jobs, and it sets the default sort order of the bid list

38%

Employers adding interview rounds

From a survey of more than 2,000 US hiring managers. It counts those adding rounds per candidate; 42% instead spend more time on résumé review

1

Applying Got Easier and Hiring Got Slower

The staffing firm Robert Half surveyed more than 2,000 hiring managers in the United States in November 2025 and released the results in March 2026. In it, 67% of HR leaders said that reviewing AI-generated applications had slowed their hiring process, and 20% of that group put the delay at more than two weeks. Another 65% said the crush of applications made it harder to confirm what a candidate could actually do. The share reporting a heavier workload for their team was 84%.

On the applicant's side of the market everything runs the other way. Tools that turn a posting into a finished application are ordinary now, and work that used to take half a day ends in a single click. A barrier vanished on one side while the labor piled up on the other.

The paper that puts this mismatch into a formal model went up on arXiv on September 24. It is "Can Labor Markets Function in the Age of AI? The Evaluation Bottleneck in Hiring," written by Itai Ashlagi, Ramesh Johari and Anushka Murthy of Stanford together with Jon Kleinberg of Cornell. Its first page carries the line "Preliminary Draft – Comments Welcome!" The work has not been through peer review.

"AI can make it easier for workers to enter an applicant pool without necessarily making it easier for firms to determine which applicants are good matches."

Ashlagi, Johari, Kleinberg and Murthy, arXiv:2609.30058, introduction

The argument is not that firms are busy because more people apply. Set the volume aside and a single application still carries and delivers less than it used to. When the same tool reads the same posting for everyone, the application of a strong match and the application of a mediocre one come out resembling each other. On the reading side, going through the text no longer tells them apart.

2

When Applications Blur, Firms Read Experience

The skeleton of the model is plain. An applicant either has visible experience or does not, and carries a match quality nobody can see. The application is a signal of that match quality with noise mixed in. Noise grows as AI spreads.

A firm acting rationally blends two things into a score. One is what this person's application says. The other is the average of the experience group this person belongs to. Noise sets the ratio of the blend. Trustworthy applications put the weight on content; blurry ones put it on the group average. Section 4 of the paper proves that single line. As AI saturation rises, the weight given to the content of an application falls, and the coarse marker of experience decides more of the score.

The same change therefore reaches people differently. The paper splits applicants four ways by experience and by match, then tracks which direction each group's average score moves. A good match was already sitting above the average of their own experience group. Cut the weight on individual information and the part above the average is what goes. A poor match was sitting below the average, so the identical change lifts their score, because less of the weakness shows.

Which way each group's average score moves as AI saturation rises no change ← score falls score rises → No experience · Good match falls the furthest Experience · Good match No experience · Poor match Experience · Poor match rises the furthest Bar lengths show only the order the paper proves among the four values. Actual magnitudes vary by market and the paper gives no numbers for them.
▲ Pebblous original diagram — drawn from the order of the inequalities in Proposition 5.1 of arXiv:2609.30058

The two groups that lose are the good matches, and the two that gain are the poor matches. At either end of that spread stands one person with experience and one without.

No prejudice is at work here. When reliable information runs short, leaning on whatever is left is what Bayes' rule instructs. What the paper aims at is the state of the information, not the attitude of the firm. Repairing it therefore calls for one more place where evidence can be produced, rather than an argument that changes anyone's mind, and that is the conclusion section 5 arrives at.

3

With No Experience, There Is Nothing to Fall Back On

Someone who has experience and fits the job also sees a lower score, only not by as much. The average of the experience group they sit in is high to begin with. As the weight on individual information shrinks, the firm's eye shifts to that group average, and the average already speaks for them to a degree. On the side without experience there is nothing to speak for them.

"In other words, they don't have pre-existing experience to compensate for the decreased weight given to their application materials."

arXiv:2609.30058, discussion of Proposition 5.1

If the score simply dropped and nothing followed, the story would stop here. But screening costs the firm something. Unlike flipping through a résumé, genuinely evaluating one person takes a human being's hours. So firms set a threshold and look only at the people above it. A score going down means falling under the line.

This is where the paper defines two failures. In the first, the firm screens nobody. When even the highest-scoring applicant comes in under the firm's threshold, the firm gives up on the hire altogether, though society would have counted that person as worth a look. The second failure is narrower. Every inexperienced, well-matched applicant who applied drops off the screening list, and at least one of them was socially worth evaluating.

The two thresholds sit apart because they count different shares. A firm counts only what the firm collects when a hire works out. What the person who lands the job collects never reaches the firm's ledger. Society's threshold therefore sits below the firm's, and people get caught between the two. The more AI blurs applications, the more people are pushed into that gap.

Where society's threshold and the firm's threshold split apart score → Society's threshold Firm's threshold the gap worth evaluating to society, but the firm never looks below threshold — nobody screens above threshold — firm screens The blurrier AI makes applications, the further right the firm's threshold moves, and the wider this gap grows.
▲ Pebblous original diagram — illustrating the two-threshold argument from section 4 of arXiv:2609.30058

In the limit where the noise grows very large, the failure stops being an exception and turns into the rule. Following a sequence in which the noise in applications grows without bound, the paper shows the probability of both failures going to one. At that point the firm's posterior converges back to the prior of the experience group. Reading an application changes nothing, so the firm looks only at people with experience or looks at nobody.

This blog has covered entry-level hiring twice, once on the bar that firms now set for juniors and once on the distance between a widely held belief that hiring has shrunk and the payroll data. Both are about what gets asked and how many get hired. This paper sits on another axis. Hold the requirements and the headcount where they are, change the basis on which applicants are told apart, and people still get pushed out.

It is worth saying plainly that this is a conclusion inside a model. Each of the two failures is a proposition with conditions attached, and the paper does not measure whether real labor markets fall inside those conditions. It handles no field data at all. That work belongs to other research.

4

In the Freelance Market It Already Happened

A record of the same mechanism observed in a live market came out a year earlier. It is "Signaling in the Age of AI: Evidence from Cover Letters," by Jingyi Cui, Gabriel Dias and Justin Ye of Yale's economics department. The theory paper cites the study in its own introduction, writing that AI increased tailoring and callbacks but reduced tailoring's predictive content, which shifted employers toward relying more heavily on past work history.

The setting is Freelancer.com. On April 19, 2023 the platform opened AI Bid Writer, a tool that reads a job description and drafts a cover letter in one click. The researchers collected 240 days of bids in two categories, PHP and internet marketing, running from January 19 to September 15, 2023. That comes to 5,499,707 bids, 264,082 applicants and 106,714 jobs, an average of 51.5 bids per job. Use was widespread: 62% of eligible users tried the tool at least once, and 17% of all cover letters were written with its help.

Start with the good news. Bids that actually used the tool were 3.56 percentage points more likely to draw a reply, which against the 7.02% average callback rate before the tool opened is a 51% increase. Measured by eligibility alone the figure is 0.43 points. It also says something about how badly cover letters had been tailored up to that point. The researchers attach two cautions of their own. The callback result is significant only at the 10% level and not at 5%, and the effect faded away after about two months.

After the tool opened, the correlation between how well a cover letter meshed with the posting and the chance of a reply fell 51%. Writing an application that fits the job used to predict success, and it has now lost half of that power. Once everyone can write to fit, having written to fit says nothing.

Firms did not leave that space empty. Over the same period, the correlation between callbacks and an applicant's reputation score rose 5%. The reputation score is the platform's own summary of the jobs a person has taken on there, and it is also the value that sets the default sort order of the bid list. The weight moved to a signal AI has trouble touching. In practice, this score sits in the place the theory paper calls a coarse observable.

After the AI bid-writer tool arrived, the signal behind callbacks changed Tailoring ↔ callback correlation −51% How well a bid fit the posting now says less Reputation ↔ callback correlation +5% The platform's own reputation score fills that space instead Both correlations come from 5.5M Freelancer.com bids, comparing 240 days before and after the tool.
▲ Pebblous original diagram — drawn from the correlation shift reported in arXiv:2509.25054

The shift did not reach far, however. The researchers report alongside it that the same reputation score gained no additional power to predict the final hiring decision. What changed is the stage of deciding whom to call in, not the stage of deciding whom to hire. The bottleneck the theory paper aims at is exactly that earlier stage.

Another study looks at the same platform from a different angle. Anaïs Galdin and Jesse Silbert built a structural model to simulate a world in which writing has lost its role as a signal completely. In that counterfactual, workers in the top 20% by ability are hired 19% less than they otherwise would be, and the bottom 20% are hired 14% more. When separating the strong from the weak gets hard, the market stops running on ability.

5

One Extra Step Changes the Outcome

The back half of the paper works out a remedy through the firm's own ledger rather than through ethics. Put one cheap intermediate evaluation ahead of the expensive full one. A work-sample task, a live skill test, a structured preliminary interview. The paper notes that internships, predoctoral positions and other temporary or probationary roles serve a related function, letting workers with limited prior experience demonstrate match quality by performing. It is a channel through which someone with a thin record can prove the fit by actually doing the work.

Why would a firm take this on by itself? Because acquiring information once more never cuts the firm's expected return. A bad result screens the person out, a good one sends them into the full evaluation, and either way the firm stands better than it does knowing nothing. So whenever the price of that information sits below its value, the firm runs the intermediate step. The paper states the conditions and shows that under them the firm's ex ante expected utility strictly increases. This is something a firm picks for itself, not something imposed on it.

The paper is equally precise about who gains. There is a band of applicants whose scores run too low for a full evaluation and too high to throw away. In hiring that ends at one stage, these people are simply rejected. Once an intermediate evaluation exists, a good result moves them up into the full review. And when AI saturation climbs high enough, the band fills up, until almost all of the inexperienced, well-matched applicants who applied belong to it.

Which candidates an intermediate step gives a second look One-stage hiring — only one threshold Applicant pool Full-evaluation threshold Inexperienced, well-matched all rejected With an intermediate evaluation Applicant pool Intermediate evaluation (cheap) Good result moves up to full evaluation A bad result still means rejection at this stage. But everyone gets at least one chance to prove match quality by actually doing the work.
▲ Pebblous original diagram — illustrating the multistage hiring argument from section 5 of arXiv:2609.30058

"Our results highlight that easier access to applying is not the same as access to credible evaluation."

arXiv:2609.30058, conclusion

Which way real firms are already moving is written into the Robert Half survey seen earlier. Asked how they are handling the surge in applications, 42% said they put more time into résumé review, 38% said they added interview rounds per candidate, and 32% said they rewrote their postings so that generic AI answers would not come back. The survey was not designed to test the model, and it comes from people with no connection to the researchers. That the second item has the same shape as the paper's intermediate evaluation is this article's reading.

Yet moving the same way is not automatically good news. An added interview round takes the applicant's time as well. The paper's condition is that the intermediate evaluation be cheap and, at the same time, sharp enough to separate people, and the stages that get added in practice do not always satisfy it. A stage that costs a great deal and separates nobody is just another threshold filtering out the inexperienced.

6

Why Pebblous Is Watching This Paper

The structure this paper draws is the structure we run into every week in data quality work. An application is observational data about a person, and hiring is a classification decision made from that data. When the data grows and the judgment gets harder, what grew is volume and what shrank is discriminative power. That is why, when we talk about AI-Ready Data, we ask about the conditions of collection before the amount collected.

The same thing plays out in labels. Better annotation tools pile labels up quickly, and when the labels that pile up do not separate from one another, a model trained on them escapes to whichever axis is easiest. It learns the imaging device instead of what is in the picture, the hospital instead of the patient's condition. A firm that reads experience in place of fit is doing the identical thing. Model and firm both grab whichever marker is still legible.

The remedy the paper puts forward is familiar from the data side too. Instead of refining an existing signal further, add a cheap verification step that produces new evidence. Our own data quality diagnosis runs on that principle. Rather than relabeling everything, we draw a small sample and split it once along the axis under suspicion. That one pass costs far less than rebuilding the whole set, and it settles what has to be fixed.

Job seekers get something out of this as well. None of these studies tells anyone to stop using AI. Cui, Dias and Ye's data in fact holds one observation running the other way: time spent editing AI-generated cover letter drafts is positively correlated with hiring success. An application sent out as drafted resembles everyone else's draft, and one edited at length resembles the person. Few people did the editing, though. More than 75% of AI-assisted cover letters went out within one minute of the click. What the paper was measuring was not the quality of the tool but how far its output ended up differing from everyone else's.

Thank you for reading this far. The model and the propositions quoted here can be checked in arXiv:2609.30058, and the Freelancer.com figures in arXiv:2509.25054. If your organization has started reading something other than the application, and if that judgment has held up better, we would be glad to hear about it.

R

References

Academic Papers

Industry Survey