AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs
4.40T1 sourcearXiv cs.MA
Source record
Published by arXiv cs.MA (T1 source). The original is at https://arxiv.org/abs/2609.31590.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryAgentWorld is a benchmark of 100 human-annotated tasks in an MMORPG sandbox evaluating long-horizon collaboration among 3–20 LLM-based agents with asymmetric roles across 50+ interaction rounds. It introduces a Causal Collaboration Effectiveness metric; tests on four top models show best achieves 52% task success with systematic failure modes including communication breakdowns and role confusion.
Why it mattersA rare empirical look at what actually breaks in multi-agent workflows beyond 20 steps. The 52% ceiling and specific failure modes matter for anyone deploying coordinated agent teams.
Cited by
No citations on record.
