Executive Summary

This article looks at a test that isolates one ability: can AI retrieve the papers that genuinely inspired a piece of research. The benchmark is ScholarCatalyst, posted to arXiv on October 1 by researchers at Stanford, Seoul National University, Carnegie Mellon, the University of Washington, the Allen Institute for AI, and MIT. What makes it unusual is who wrote the answer key. The 184 people who authored the work went back over 207 of their own projects, marked which prior studies actually helped, and wrote down why each one did.

The results cut against expectation. Searching 190,896 candidates for those papers, an agent that calls the retrieval tool directly landed at Recall@20 of 0.42, under the 0.48 of a single pass through the same retriever. Comprehension was never the limit. Supplying the agent with the source paper's full reference list lifts that same figure from 0.39 to 0.74.

Sections 1 through 4 follow what the paper measured and the cautions its authors attached. Section 5 moves to evaluation data design, and that move is this article's reading rather than a claim in the paper.

Key figures

Source: Kim et al., ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research, arXiv:2610.02202v1 (2026-10-01), body and tables

0.42

Recall@20 for the agent holding the search tool

One pass through that same retriever scores 0.48, higher than this

0.39→0.74

Shift when the reference list is handed over

Recall@20 on core research queries. Give it the candidates and the score nearly doubles

43.6%

Gold papers the source work never cited

Subfield queries. Follow citations alone and almost half of the inspiration stays invisible

43–60%

Overlap between co-author labels and the lead author's

A figure the authors measured and published themselves. Answer keys carry human disagreement too

1

Rewinding the question to before the work existed

The usual way to score a literature search tool is to feed it the title or abstract of a finished paper and check whether similar papers come back. That procedure takes the shape of completed research as its reference point. ScholarCatalyst pulls the reference point earlier. It reconstructs the question an author was holding while the project was still unfinished, then asks which prior papers helped at that moment, or would have helped had the author known about them.

Queries come in two layers. A core research query carries the central question of one project, and there are 207 of them. Subfield queries ask about the same project from a narrower angle, and there are 687. Together that makes 894. Core queries average 131.8 tokens and 3.19 gold papers each; subfield queries are shorter at 63.5 tokens and carry 6.09 gold papers. The haystack is 190,896 papers, built from the references the source papers cite plus roughly 181,000 computer science preprints posted to arXiv between 2020 and 2024.

Queries come in two layers How ScholarCatalyst's 894 queries are structured Core research query (CoreQ) 207 131.8 tokens avg · 3.19 gold papers avg One central question per project Longer queries, a narrower gold set Subfield query (SubQ) 687 63.5 tokens avg · 6.09 gold papers avg Same project, asked from a narrower angle Shorter queries, a wider gold set Together: 894 queries Haystack of 190,896 papers (cited work + 2020–2024 CS arXiv)
▲ Original Pebblous diagram | Source: Kim et al. (2026), arXiv:2610.02202v1, rebuilt from Section 3.1 and Table 2

A single table in the paper sets out the gap with earlier literature search benchmarks. Of SciFact, DORIS-MAE, ScholarQABench, LitSearch, and MIR, none satisfies three conditions at once: the query belongs to the period before publication, the labels come from the author, and relevance counts even where no citation documents it. ScholarCatalyst is the first to hold all three together. More than 100 of the source papers were orals, spotlights, or award winners at ICLR, ICML, NeurIPS, ACL, CVPR, CoLM, and CoRL in 2025 and 2026.

2

The answer key comes from the people who did the research

Building the key runs in two stages, and a machine handles the first. Given nothing but the arXiv identifier of a source paper, the pipeline walks the citation graph to gather candidates, and Gemini 3.1 Pro drafts the question that paper set out to answer. BM25 and Qwen3-Embedding-8B then sweep the corpus again to pull in work the paper never cited, and Gemini 3.6 Flash reranks everything down to ten candidates per query.

Stage two is what gives this benchmark its character. The lead or corresponding author opens each of those ten candidates in a web interface and marks whether the paper inspired the project, or would have been useful had they known about it at the time. Marking is not the end of it: the author writes out the reason in prose. The machine-drafted question is editable too. Of the 894 queries, 764 survived author review with no edits at all, a share of 85.5%.

The machine narrows the candidates, the authors settle the answers How the ScholarCatalyst answer key was built Stage 1 · Machine Starts from one arXiv ID Citation graph · query drafting BM25 + embedding + reranking 10 per query Cited papers and uncited papers arrive together Stage 2 · 184 authors Verify and rewrite candidates Mark what inspired the work Write down why it helped The finished answer key 207 projects · 894 queries (207 core + 687 subfield) · 190,896 papers 85.5% of the machine-written queries came through author review unchanged The authors edited or rewrote the rest
▲ Original Pebblous diagram | Source: Kim et al. (2026), arXiv:2610.02202v1, rebuilt from Sections 3.2 and 3.3

The citation rates show what this design changed. For core research queries, 95.5% of the gold papers do appear in the source paper's bibliography. Drop to subfield queries and that share falls to 56.4%, leaving 43.6% with no citation anywhere. Sitting in that remainder are papers the author had not found at the time but would have used, and papers they did read without ever finding a place to cite. Any key built out of citation links alone never sees them at all.

3

Topic similarity does not mark inspiration

The candidates an author rejected are useful material in their own right. Every one of them was ranked highly by the same retriever, so they are close in subject matter by construction. The authors measured the rejected papers against the accepted ones on lexical overlap and on embedding cosine similarity, and the rejected side came out at least as similar. Neither ruler separates the two groups.

So what were the authors responding to? A classification of 663 core inspiration cases into nine types supplies the answer. The reasons cited most often were adapting a technique, finding empirical support for a direction the author was already taking, generalizing something to a new setting, and reading a stated limitation that prompted a different approach. One paper fell into two or more types in 55.8% of cases. Similarity tells you whether two papers share a subject. These nine types describe what role one paper played in another's research. Different axis entirely.

Similarity does not separate them, role does Rejected vs. accepted papers · 663 core inspiration cases Similarity is nearly equal Lexical overlap Rejected Accepted Embedding cosine similarity Rejected Accepted Neither ruler tells them apart What separated them: role The four most common reasons Adapting a technique · Empirical support Generalizing to a new setting A limitation that prompted a new approach 55.8% fall into two or more types Similarity says two papers share a subject; role says what pushed the research forward Different axis entirely
▲ Original Pebblous diagram | Source: Kim et al. (2026), arXiv:2610.02202v1, rebuilt from Section 4.3

Citation records stand in for that role no better. Among gold papers that were cited, the ones carrying a citation-intent label meaning uses or extends amount to 46.2% for core research queries and 34.0% for subfield queries. A further 49.5% of gold papers came from outside the project's own topic area. Inspiration arrives from the field next door in nearly half of all cases.

One experiment handed Gemini 3.1 Pro the full text of a source paper together with its complete bibliography and asked it to name the papers that inspired the work. It averaged 32.2% of the gold papers for core research queries, and in 28.0% of cases it named none of the core inspirations at all. That is the score after reading the finished paper and seeing everything it cited.

4

Calling the search tool directly lowers the score

Two families of system went into the test. A standalone retriever throws one query and takes back the top documents. An agent holds that same retriever as a tool, calls it repeatedly, reads the intermediate results, and decides what to ask next. Common sense favors the second. The numbers went the other way.

R@20 in the table below is the share of author-approved gold papers that land in the top 20 of the results. The deep research agent splits a query into several sub-questions and searches repeatedly against them. The title-and-abstract grep agent skips embeddings altogether and scans the corpus titles and abstracts with regular expressions.

System Core research query R@20 Subfield query R@20
Embedding retrieval (best per model) 0.39 0.51
Deep research agent (o3) 0.33 0.38
Lexical retrieval (BM25) 0.23 0.33
Title-and-abstract grep agent (GPT-4.1) 0.06 0.09

The best embedding figures come from Qwen3-Embedding-4B on core research queries and Qwen3-Embedding-8B on subfield queries. Weighting the two query types by count, embedding retrieval reaches 0.48 and the GPT-4.1 agent calling that same retriever reaches 0.42. Source: arXiv:2610.02202v1, Table 3.

The bottom row exposes the nature of the bottleneck first. The paper separately counted how often each agent encountered a gold paper at any point during its search. For the grep agent that was 8%, against 46% for the agent calling the retriever. Nothing went wrong with the grep agent's ranking; it simply never met the right papers. And of the 46% the other agent did meet, only 41% made the final top 20. Either way, the ceiling on the score is set by the range of candidates that ever reached the agent.

Models trained specifically on scientific literature might seem better placed, but that is not how it turned out. SPECTER2 and OpenScholar trail general-purpose embedding models by at least 16 points on core research queries and at least 27 points on subfield queries. One result sits in the table under a separate marker: a deep research agent backed by Claude Fable 5.1 posted 0.51, and the authors ruled it out as a fair comparison because the source papers were published after the knowledge cutoff and may have entered training. Even on those terms it misses nearly as much as it finds.

The disadvantage stops being mysterious once you work out what an agent can actually do here. The only tool it can reach for is the same similarity-based retriever. However many times it rewrites the query and calls again, a paper that retriever never ranks highly never appears in front of the agent even once. The paper calls this bottleneck candidate coverage. Scaling the backbone does buy something: agent recall tracks the backbone's general intelligence score with a Spearman correlation of 0.67. Yet the strongest backbone merely draws level with the standalone retriever and stops there.

The shortfall is candidate coverage, not comprehension Core research query Recall@20 · same agent, different input With the retriever as a tool It rewrites the query and calls again Papers the retriever never ranks never come into view at all 0.39 With the reference list given The source paper's bibliography goes straight into the input The candidates are settled in advance 0.74 Same model, same tool, only the input changed The candidate list was the one thing missing Query rewriting produced no gain of that size, and HyDE cut recall instead
▲ Original Pebblous diagram | Source: Kim et al. (2026), arXiv:2610.02202v1, rebuilt from Tables 4 and 5

The decisive contrast changed one input and nothing else. Give the agent the full reference list of the source paper and Recall@20 on core research queries goes from 0.39 to 0.74. Same model, same tool. All it received was an indication of where to look. Prescriptions pointing the other way came up empty. Rewriting the query with an LLM or fanning it out into several produced no meaningful gain, and HyDE, which drafts a hypothetical answer document first and searches with that, cut recall sharply instead. Polishing the question does nothing for papers that sit outside the candidate pool.

There is a caveat the authors attached to that contrast. A source paper's reference list already contains 95% of the gold papers for core research queries, so the setup comes close to handing over the key, and 0.74 cannot be read as a reachable target. Under the same condition, subfield queries moved only from 0.47 to 0.57. The direction still holds. Supplying just the title and abstract of the source paper lifted core research queries by 0.11, which says the gain came from a narrower field of candidates rather than from more material to read.

5

Who writes an answer key, and on what grounds

The authors do not present their key as perfect. Their limitations section reports a preliminary study in which three co-authors of one paper labeled it independently. Of the gold papers those co-authors marked, between 43% and 60% overlapped with the lead author's labels, at precision between 84% and 88%. People who did the same research together agree about half the time on what counted as inspiration. The authors also note that a label is the author's own retrospective account, which a later reading could revise.

They decline to hand that figure to the machines as a defense. In the same passage the authors point out that even an overlap of 43% to 60% sits well above the best system's overall R@5 of 0.24. People agree with one another at a rate no system comes near. What the preliminary study told the authors, then, was not that the key wobbles but that there is a great deal of room left to climb.

The scope is bounded too. Coverage stops at computer science projects from 2025 and 2026, which no one would call representative of other fields, and the corpus holds 190,896 papers against the 225 million in Semantic Scholar, so the real problem dwarfs what this test measures. The authors place the task alongside what Swanson in 1986 called undiscovered public knowledge, meaning connections that already sit in the open literature with nobody joining them. Their outlook follows from there. No single researcher can track more than a slice of the field, so a system capable of expert-level judgment across many fields might unlock progress that was sitting in plain view all along. Researchers may be limited more by what they have seen than by how well they judge it, the authors add carefully.

Even so, three design principles come through clearly in this key. First, it changed who does the labeling. An external worker can judge whether two papers share a subject, but only the person who did the research knows what pushed it forward. Second, it attached a reason to every label. The nine-role classification was not imposed afterwards; it emerged from the rationales the authors wrote. Third, it detached the condition for a gold label from the record. Deciding that uncited work still counts when it inspired something rescued 43.6% of the subfield query gold set.

Anyone building evaluation data faces those three as questions. Who is writing the answer key for the work we assign to AI? Is that person qualified to judge, or merely conveniently placed to do so? And does the reason something counts as correct survive next to the label? More on designing evaluation data sits in why the thing that writes the work cannot grade it, and a case of asking what a benchmark score is actually attached to runs through the vision language models that made the same choice after the evidence changed.

Editor's Note

Where a label came from is the question Pebblous keeps returning to when we examine data quality. Who applied this label, and is there any record of why they applied it that way? ScholarCatalyst is a rare case that leaves an answer to both, and it published alongside them the fact that people overlap only 43% to 60% with each other. Good evaluation data does not mean flawless data; it means data that records how far it can be trusted. That connection is something we draw here, not something the paper argues.

Thank you for reading this far. The full text sits at arXiv:2610.02202, with the code and dataset released on GitHub and Hugging Face respectively. Every figure in this article was checked against the body and the tables. If your team has built evaluation data of its own, we would like to hear how you designed the part that keeps a reason beside each label.

Pebblous Data Communication Team
October 3, 2026

References

  • 1.Kim, S., Lee, Y., Liu, B., Ko, D., Shao, R., Kim, S., Neubig, G., Koh, P. W., Chowdhery, A., Asai, A., Khattab, O., Choi, Y., Kim, G., & Finn, C. (2026). ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research. arXiv:2610.02202
  • 2.Stanford IRIS Lab. (2026). ScholarCatalyst (code repository). GitHub
  • 3.ScholarCatalyst. (2026). ScholarCatalyst (dataset). Hugging Face Datasets