ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software
4.20T1 sourcearXiv cs.MA
Source record
Published by arXiv cs.MA (T1 source). The original is at https://arxiv.org/abs/2609.17885.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryERPBench is a benchmark for evaluating screenshot-based computer-use agents on a live, reproducible Enterprise Resource Planning system. It scores tasks against ground-truth database values and includes a harness that gates agent actions behind human approval. Testing six agents shows strong general GUI performance does not transfer to enterprise reliability, with some agents saving forms in 85% of runs but writing correct values in as few as 3%.
Why it mattersQuantifies a concrete gap between surface-level agent success and correct data persistence in enterprise software, with a reusable benchmark and human-approval harness.
Cited by
No citations on record.
