One Run Is Not an Idea: The Implementation Lottery in Automated Research
4.00T1 sourcearXiv cs.MA
Source record
Published by arXiv cs.MA (T1 source). The original is at https://arxiv.org/abs/2607.26587.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryThe paper identifies the 'implementation lottery' in automated research: treating a single experimental run as evidence for an underlying idea conflates implementation variance with idea-level claims. The proposed Idea Reliability Audit samples multiple implementations per idea, finding winner-reversal rates of 25.6–43.6% across 312 assignments on 13 tabular tasks and two coding-agent setups.
Why it mattersA concrete diagnostic protocol for anyone running AI-agent research pipelines. It quantifies how often a winning idea flips when re-implemented, pushing toward multi-implementation evidence before any idea-level branching or transfer.
Cited by
No citations on record.
