Where Do Multi-Agent Systems Fail? Evidence-Grounded Diagnosis of Collective Mechanisms
3.80T1 sourcearXiv cs.MA
Source record
Published by arXiv cs.MA (T1 source). The original is at https://arxiv.org/abs/2609.38761.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryProposes diagnostic contracts that judge whether specific collective mechanisms in multi-agent systems were violated, based on execution records rather than final outputs. Replay experiments show broken mechanisms often still yield correct answers, and LLM diagnosers extract more from internal records than public outputs, but transfer poorly across systems.
Why it mattersConcrete empirical evidence that a correct answer does not confirm internal mechanism integrity, with a tested diagnostic framework and an explicit caveat on cross-system transferability for builders of multi-agent workflows.
Cited by
No citations on record.
