In early 2025, Sakana AI ran what was effectively an experiment in autonomous research. The company used its AI Scientist-v2 system to generate three research papers from scratch, covering topic selection, literature review, experiments, and writing, and submitted them to the "I Can't Believe It's Not Better" workshop at ICLR 2025. The workshop organizers knew AI-generated submissions might be in the pile. They didn't know which papers were which.

One of the three cleared peer review, with scores of 6, 7, and 6 from its three reviewers, above the workshop's acceptance threshold. The topic was appropriate for a conference named after scientific disappointments: the paper reported that a particular approach to improving neural network generalization didn't work as hoped.

Sakana then withdrew it before publication. The reason they gave was candid: "the AI and scientific communities have not yet decided whether we want to publish AI-generated manuscripts in the same venues." The paper had real problems too. Critics noted hallucinated references and duplicated figures. Generating it cost around $140 and took roughly 15 hours of compute time.

The ICBINB workshop is a lower-stakes venue than ICLR's main conference track, and its acceptance rates are higher. TechCrunch noted that reviewers there tend to be more junior than at the main track. But the organizers certified the review process as valid, and the experiment was conducted with their knowledge and consent from the start.

A technical report describing the AI Scientist-v2 system in full was posted to arXiv in April 2025, and the story has continued drawing attention, with Scientific American covering the broader implications this year.

The episode raises a question that has not been cleanly resolved: what does peer review actually certify? The reviewers found the paper plausible enough to accept. That says something about the AI system's capability, but it also says something about what a blind review process can and can't catch. A paper with hallucinated citations cleared evaluation by human experts, at least under these conditions.

Sakana did not argue this was a success. They framed it as a question. That framing is probably the most honest thing about the whole experiment.

Sources

  1. i. arxiv.org
  2. ii. sakana.ai
  3. iii. www.scientificamerican.com
  4. iv. techcrunch.com

Commentarii · 0

Add · a · Comment