Executive Summary

Ethics review at AI conferences exists to turn research toward safer practice before it reaches publication. Researchers at MIT and Harvard asked whether it actually turns anything. They gathered 47,893 ICLR submissions and 185,194 public reviews, then followed papers that picked up an ethics flag and were rejected or withdrawn into the versions their authors later submitted elsewhere. The paper went up on arXiv on September 9, 2026.

Of the 446 resubmissions the team confirmed, 370 left the flagged methods or procedures in place. In 217 cases nothing observable changed in relation to the concern, and in another 153 the authors changed only how they described limitations and possible harms. Methods or procedures actually changed in 76 papers, or 17.0%. That 83% is not a share of every flagged paper. Its denominator is the 446 resubmissions the team could locate in public records out of 1,857 rejected or withdrawn submissions, and an author who abandoned a project after a flag never enters that denominator at all.

Sections 1, 2, 3 and 5 stay with what the paper reports. Section 4 reads the finding as a data governance problem. That reading is ours, not the paper's.

Key Numbers

Source: Nishi, Laprevotte, Bullock, and Andrew, Governing AI Research Through Peer Review, arXiv:2609.10740 (September 9, 2026)

83.0%

Resubmissions that kept the methods and procedures

370 of 446. In 217 nothing observable changed; in 153 only the framing did

76 papers

Resubmissions that changed a flagged method or procedure

17.0%. Leakage checks and safety evaluations added, consent or IRB review introduced, identifiers removed, data access or release changed

107 → 1,100

Ethics flags recorded at ICLR

From 107 in 2022 to 1,100 in 2026, summed across six ethics categories

3 / 25

Authors who answered the interview request

The qualitative sample ends here, and the paper lists the response rate as a limitation

1

Where the Flags Are Raised, and How 446 Papers Were Found

An ethics flag is a reviewer ticking the "Flag for ethics review" checkbox on the review form. NeurIPS, ICML and ICLR have spent several years formalizing research ethics through broader impact statements, submission checklists and flags of this kind. Under those requirements a reviewer can object not only to a technical claim but to how the authors collected their data, how they evaluated possible harms, how they released their artifacts, and how they anticipated misuse.

The paper lays the ground for that authority before it reports anything. Choices about data, system design and deployment distribute risk, encode values and reshape social relations. No researcher can foresee every downstream effect, yet Stilgoe et al. (2013), whom the paper cites, argue for anticipating plausible effects and responding to them, and Do et al. (2023) separate unintended consequences from unanticipated ones and place responsibility for the foreseeable kind with the researcher. Since major conferences decide which work receives attention and legitimacy, Hecht et al. (2021) hold that they share responsibility for what they accept and promote. That is where the study starts.

Some of this machinery already has a track record. NeurIPS established ethics guidelines in 2021, issued several conditional acceptances that cycle and rejected one paper on ethical grounds. The ICLR 2023 program chairs reported screening 190 papers that carried ethics flags. Earlier work, though, either examined a single review cycle or discussed ethical practice in general, and never asked whether ethics review shapes the research itself or in what way. That question is what this paper adds.

The team chose ICLR for its disclosure policy. Among the major AI conferences only ICLR releases every submission and review after decisions, while NeurIPS and ICML publish rejected papers only when authors opt in. From OpenReview the team collected 47,893 submissions and 185,194 public reviews covering 2021 through 2026. Their 2021 snapshot carries no public flags, so the counts begin in 2022, and they rise from 107 that year to 1,100 in 2026. That total sums six ethics categories.

Flags appeared on 2,498 submissions, and 1,857 of those were rejected or withdrawn. Authors routinely change a title, an abstract, a coauthor list or a scope, so matching titles alone cannot find the resubmissions. The team queried OpenAlex with title variants, author surnames and salient abstract terms, retrieved 6,618 candidates, and ranked them by title similarity, author and abstract overlap, and publication year to keep the top 601. A retrieval score cannot establish that two papers are the same project, so they read every pair by hand and retained 446. Code for building the corpus and running the retrieval sits in an anonymous repository.

Narrowing down to 446 resubmissions Five steps that narrow the denominator inside OpenReview's public record (bar length is log-scaled) ICLR submissions (2021–2026) 47,893 Submissions with an ethics flag 2,498 Of those, rejected or withdrawn 1,857 Top resubmission candidates 601 Resubmissions confirmed by hand 446 The team pulled 6,618 candidates, kept the 601 top matches, and read every pair to retain 446 The 83% figure uses these 446 as its denominator
▲ Pebblous original diagram | Each step's count is the value reported in Nishi et al. (2026) §3

So reading the 83% means reading its denominator alongside it. It is not 83% of the 2,498 flagged submissions, and it is not 83% of the 1,857 that were rejected or withdrawn. It is 83% of the 446 resubmissions that public records made findable. An author who took a flag and shelved the project never enters that set. The paper lists this as a limitation and notes that missing such cases would mean the study underestimates rather than overestimates how much review steers research.

2

The Prose Changed, the Methods Did Not

A difference between two versions does not by itself show that review steered the work, so the team fixed on one observable consequence: whether the methods and procedures that the ethics concern implicated had changed. For each pair they located the concern in the original review, compared the corresponding parts of the submission and the resubmission, and coded the resubmission into one of five categories. Leaving the concern unaddressed counts as no observable change. Changing limitations, discussion of harms or claims while leaving the implicated technical work alone counts as mainly rhetorical. Procedures cover consent, release, access, oversight and governance; methods cover data, models, evaluations and experimental design. Four researchers did the coding by hand and met regularly to calibrate their labels. The paper reports no inter-coder agreement statistic.

Change after the flag Count Share
No observable change 217 48.7%
Mainly rhetorical 153 34.3%
Procedures + methods 57 12.8%
Primarily procedures 14 3.1%
Primarily methods 5 1.1%
Total labeled 446 100%

The counts come from Table 1 of the paper. The shares are ours, computed against 446. The shares the paper states directly are the 83.0% covering the top two categories, 370 resubmissions, and the 17.0% covering the remaining 76.

In 217 resubmissions nothing observable changed in relation to the concern, and in 153 the authors changed only how they framed or discussed it. Among the remaining 76, authors added leakage checks, annotation workflows and safety evaluations, introduced consent procedures or review by institutional review boards, removed identifiers, and changed how data was accessed or released. For a process meant to steer research toward safer practice, the team wrote, substantive change in only 17.0% of cases is concerning.

None of this makes a prose revision worthless. The paper attaches a caveat: authors may narrow their claims, clarify limitations, document provenance or adopt a safer release plan, and revisions like those can improve a paper. They simply do not change the implicated methods or procedures, which is why the team did not count them as research changes.

A separate study has already measured what editing the sentences alone can do to a review outcome. In the experiment on 4,080 manuscripts rewritten only at the level of prose that we covered in August, the methods and the reported numbers stayed frozen while the framing of evidence and the novelty claims were reworked, and the LLM judges' scores diverged. If prose is what moves publication visibility, an author who invests in prose is not making a strange choice. Connecting the two studies this way is our reading, though, and neither one cites the other.

3

Concessions Made in Rebuttal Vanish With the Rejection

Much of the result comes down to when reviewers arrive. By the time a reviewer reads a submission, the authors have already picked the problem, collected the data, implemented the system and finished most of the experiments. During the discussion period reviewers can still update their scores in response to rebuttals and author comments, which gives authors a reason to acknowledge a concern and engage with it. A rejection decision removes that reason. The original reviewers no longer influence how the work travels afterward. Given how noisy reviewer judgments are known to be, there is little ground to expect a fresh panel at the next conference to raise the same concern on its own. The paper reads this structure as leaving room for concessions in the discussion period that never have to survive into the next venue.

A concession only carries weight during rebuttal From submission to resubmission, the window where a concession still counts, and the window where it does not Leverage ends here Submission Review & rebuttal Concession updates score Decision: Reject Resubmission Original reviewers have no influence A concession on the left counts toward the score; on the right, the same concession need not survive A pattern found in the close reading of 25 review histories
▲ Pebblous original diagram | This diagram illustrates the mechanism Nishi et al. (2026) describes for why concessions vanish after a rejection

The 25 cases the team read closely were not picked at random. They were chosen to span different relationships between the flag and the resubmitted project: 11 apparent research changes, 6 pre-existing practices, and 8 cases with no apparent change. By the nature of the concern they cover four strands, with 4 cases of harmful applications, 9 of fairness and discrimination, 6 of responsible research practice, and 6 of privacy, security and safety.

Those review histories show how the room gets used. Authors often answer ethics criticism at length during rebuttal and concede part of it, and those concessions are frequently absent from the later resubmission. In most of the cases only one of three to five reviewers raises the concern while the others discuss technical issues. In at least one case the reviewer who emphasized ethics scored the paper 6 while the reviewers who did not flag ethics scored it 4 to 5. Several rebuttals frame the concern as a dispute about disciplinary convention or required background rather than a valid reason to change the research.

The team calls this pattern filtering: adapting a paper for publication without treating criticism as grounds to redirect the research. They invited the authors of the 25 cases to interview, three responded, and after internal review they ran 30-minute conversations under standard consent and privacy safeguards starting April 15, 2026. To avoid planting an answer about why revisions happen, they kept the recruitment email neutral. They asked three things in each interview: what the authors changed, what they paid most attention to, and whether reviewer feedback affected decisions about the research or only about presentation.

Participant 1 (P1) described the resubmission as "mostly similar." The revisions centered on "adding another table" and "more metrics," and the issues at stake were "mostly metric based." P1 added that peer reviews are "mostly to filter out papers," on the grounds that many reviewers repeat objections to text the authors have already revised.

Participant 2 (P2) said reviewers were "very focused on a particular aspect, […] instead of seeing a larger picture." The authors "decided to just cut it out completely," rewrote the introduction for an audience in computer science, and recruited 5,000 participants for another annotation pass. P2 accepted several of the concerns as "fair points" and still called the central criticism "really about my presentation," and called peer review "really about an editorial process." P2 attributed the narrowing to reviewers in computer science, saying "the CS reviewers chose to filter it," and said reviewers "have so much power shaping what stories get seen."

Participant 3 (P3) rewrote the introduction and the abstract and left the methods and results essentially unchanged, and called review closer to a "coin flip" than a reason to change the research. The three accounts point the same way. Revising for publication and revising the direction of the research are separate acts for an author, and review moved the first one.

4

When the Warning Does Not Travel With the Artifact

From here on, this is our reading rather than the paper's.

This failure starts with where the warning gets filed. A flag looks like something attached to the paper, but it actually belongs to that review session. When the session closes with a rejection, the record survives and the warning stays locked inside it. The moment the paper moves to another conference or a journal, the warning does not move with it, and the new panel has no way to know the objection was ever raised. The artifact travels and the warning stays put, so travel amounts to erasure.

Anyone who has worked on data governance will find this familiar. A dataset review meeting raises a privacy concern, a bias, or a license restriction, and the conclusion ends up in the minutes and the ticket instead of the dataset card or the manifest. The moment that dataset moves to another pipeline, another team, or the training run for another model, the warning stays behind. Whoever receives it believes they have taken delivery of an artifact with nothing marked against it. When the audit trail and the artifact live in separate places, a risk label lasts no longer than the tenure of the people who remember the meeting.

Where you file a risk label decides how long it lives Two designs that apply equally to a paper's ethics flag and a dataset's risk label Filed in the review record Where the label lives That session's record When the artifact moves The label stays behind The recipient Cannot see the prior objection Attached to the artifact Where the label lives The paper's or dataset's own record When the artifact moves The label travels with it The recipient Can read the prior objection The 446 sit on the left. The paper's first recommendation asks for the right Either way a record exists. The difference is whether it follows the output
▲ Pebblous original diagram | The left panel is the state Nishi et al. (2026) observed, and the right panel is the first recommendation in §5 of the same paper. Carrying the reading over to datasets is ours, not the paper's

The paper's first recommendation aims at exactly this point. Requiring authors to disclose a prior ethics flag on resubmission means filing the warning against the artifact instead of the session. The same design shows up in training data work, where provenance and lineage get attached to the artifact itself, and we once wrote about keeping training data separated by source. There the question was whether you can say where your data came from when an audit asks. Here the question is whether the next user can read a concern that someone already raised. Provenance in one case and a warning in the other, and the design underneath is the same.

5

What the Paper Prescribes, and What It Cannot Say

The paper proposes three changes. The first is to carry ethics flags across resubmissions. Authors would confidentially disclose a prior flag when they resubmit, including which category it fell under and whether they changed the research or disputed the concern. The new venue could then ask an independent ethics reviewer to assess that concern without revealing the previous review, the identities, the scores, the venue or the manuscript to its own reviewers and area chairs. Unresolved concerns stay under independent scrutiny, and the new venue stays unbound by the prior verdict.

The second is early-stage ethics review. By the time a full paper reaches review, the data collection and the experiments are usually finished, so conferences could require an early-stage ethics pre-submission as a prerequisite for the full-paper submission. That earlier review would scrutinize data collection and experimental design, leaving broader critiques to the later stage. The third is reviewer fit. Criticism carries more credibility when it comes from a reviewer matched to the domain, and the evidence attached to that recommendation is P1 discounting a reviewer who repeated an objection to text already removed, and P2 pointing at reviewers in computer science.

What the study cannot say is equally clear. The corpus holds ICLR only, with no NeurIPS or ICML. The team expects the mechanisms they identify to extend past ICLR, given the similar review, rebuttal and decision cycle at top AI conferences, while noting that the prevalence may differ. The interviews are three people out of 25. And comparing two versions reveals what authors changed without revealing why they changed it. A method may have moved because the authors accepted the criticism, or because the project evolved on its own, or because they rejected the criticism and rewrote the paper for a different audience. That is why the team added the close reading of review histories and the interviews. The paper itself is a preprint posted to arXiv on September 9, 2026, and the text does not say whether a conference or journal has accepted it.

The paper's conclusion holds all the same. A substantial revision to a manuscript does not necessarily mean that ethics review redirected the underlying research. Concessions carry weight during rebuttal because reviewers still influence the decision; after a rejection the original reviewers lose that leverage and reviewer noise may keep the concern from resurfacing. So the paper insists on disclosure of prior flags upon resubmission, so that a rejection does not reset accountability for an unresolved ethics concern when the project moves to another venue.

The question this leaves for your own organization is not about review boards. It is about where you filed the risk label.

  • Do the risk labels on your datasets and models live in the artifact's own record, or only in the minutes of a review meeting?
  • When another team or another pipeline takes that artifact, can the recipient read the objection someone raised earlier?
  • Of the concerns raised in last quarter's data reviews, how many changed a pipeline or a procedure, and how many changed only a line of documentation?

Thank you for reading this far. The paper is at arXiv:2609.10740, and every figure and quotation here was checked against its text. If you have ever answered the third question, we would like to hear how you counted.

Pebblous Data Communication Team
September 12, 2026

R

References

Key Paper

Related Work