AI agents flunk real research: recursive self-improvement may be further off than promised

AI recursive self-improvement might not come so quickly after all (August 2026)

AI agents flunk real research: recursive self-improvement may be further off than promised

A new study led by Princeton researchers tested AI agents on open-ended AI research using a "shadow evaluation" method: agents had six days, $3,000 in API credits, and a GPU budget to produce a paper worthy of NeurIPS 2026. Both papers were rejected. The agents handled all the engineering—reviewing literature, running hundreds of experiments—but lacked the judgment and creativity to make novel contributions. The authors say this gap suggests hyped timelines for recursive self-improvement may be running ahead of the evidence.

There’s a certain absence of valuable, intuitive creativity in today’s AI systems, and though they’re extraordinarily capable engineers they seem to have a certain property of rote, formulaic thinking that might prevent them [from] being good researchers.

More from this day

2026-09-13