At Equal Inference Cost, Multi-Agent Structure Does Not Beat a Single Frozen Agent
4.20T1 sourcearXiv cs.MA
Source record
Published by arXiv cs.MA (T1 source). The original is at https://arxiv.org/abs/2609.04217.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA controlled study finds multi-agent LLM pipelines (Planner-Executor-Critic) provide no statistically significant improvement over a single evolved agent when total inference cost is held equal. On ALFWorld, the team scored 0.769 vs single agent 0.754 (p=0.80) despite 1.8x more evaluation calls; on WebShop the team trended worse.
Why it mattersA cost-controlled benchmark that challenges the common assumption that multi-agent topologies improve outcomes - the gains come from one executor, not the team structure.
Cited by
No citations on record.
