Auto Research for Materials: Auditable AI-Scientist Workflows with Held-Out Transfer
3.80T1 sourcearXiv cs.MA
Source record
Published by arXiv cs.MA (T1 source). The original is at https://arxiv.org/abs/2607.17100.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA study of an AI research agent for materials science that searches separately across features, models, representations, and training data, evaluating 701 changes across 10 Matbench endpoints. Selected code is frozen and tested once on untouched holdout data; 9 of 10 choices remain the best single intervention, with 26.3% mean improvement from combining feature and model changes.
Why it mattersProposes a held-out evaluation protocol for closed-loop AI scientists, addressing whether agent-selected modelling changes generalize beyond the development loop and whether their code is reusable across tasks.
Cited by
No citations on record.
