When AI Writes the Papers and Agents Try to Read Them
ICML received nearly 24,000 submissions this year. When humans couldn't keep up, a massive community challenge turned coding agents loose to reproduce 2,200 of them. The results are a warning....

We are drowning in papers. If you need proof, look at ICML this year, where submissions nearly doubled to a staggering 23,918 manuscripts, driven almost entirely by the relentless efficiency of automated writing tools and AI-assisted research workflows.
Reviewers are drowning too. Most academic peer review is a volunteer gig run on borrowed time, which explains how a spotlight paper with unchecked proofs slides right through the cracks with glowing scores because nobody had a free weekend to verify the math.
So a fascinating question emerges: if AI generated this tidal wave of science, can we use AI to audit it? That is exactly what the recent ICML Open Reproductions challenge set out to test. Unleashing an army of coding agents across thousands of cloud jobs to see if peer review can finally scale.

The findings should terrify anyone who still trusts conference proceedings blindly. Out of more than two thousand papers examined by the community, only a fraction were fully verified, exposing a fragile ecosystem where publishing velocity has completely outstripped our ability to check if the work actually holds water.
We built our whole modern technological foundation on the promise of rigorous peer review. If that foundation turns into a rubber-stamped sludge of unverified agent-generated output, we aren't advancing human knowledge anymore – we're just running a very expensive, very sophisticated game of telephone.








