AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use
4.00T1 sourcearXiv cs.MA
Source record
Published by arXiv cs.MA (T1 source). The original is at https://arxiv.org/abs/2607.20536.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryAppWorld-UL is a benchmark of 516 tasks across 9 simulated apps (Amazon, Spotify, etc.) evaluating tool-use agents' ability to interact with users via clarifications, confirmations, and infeasibility notices. Claude Opus 4.7 achieves only 48.6% overall success, dropping to 21.3% on compositional tasks.
Why it mattersQuantifies the user-interaction gap in tool-use agents with a concrete benchmark and a frontier-model score, giving builders a baseline reference for interactive agent design.
Cited by
No citations on record.
