ToMAS: A Pilot Failure-Grounded Theory-of-Mind Benchmark from Multi-Agent LLM Failures
2.80T1 sourcearXiv cs.MA
Source record
Published by arXiv cs.MA (T1 source). The original is at https://arxiv.org/abs/2609.16986.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryToMAS converts 242 diagnosed multi-agent LLM coordination failures into 39 clean Theory-of-Mind benchmark items, with 94.4% annotator agreement. A GRPO training pilot on Qwen2.5-1.5B failed to show learning effects because LoRA updates were numerically negligible (max ΔW ≈ 7e-6).
Why it mattersUnusually candid about its own experimental null result; the paper's value is the conversion pipeline and the two matched-domain evaluation requirements it surfaces for future work.
Cited by
No citations on record.
