Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?
3.60T1 sourcearXiv cs.MA
Source record
Published by arXiv cs.MA (T1 source). The original is at https://arxiv.org/abs/2607.09996.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryWho&When Pro is a benchmark of 12,326 failed agent trajectories across 26 benchmarks and 3 modalities, designed to evaluate LLM-based automated failure attribution in agentic systems and reveal systematic patterns in attribution accuracy.
Why it mattersOffers a controlled benchmark and empirical findings on LLM failure attribution in agent workflows, relevant to anyone building or debugging multi-step agent systems.
Cited by
No citations on record.
