WebUIProof: Benchmarking WebUI Code Generators with UI-Agent Execution Harness
3.40T1 sourcearXiv cs.MA
Source record
Published by arXiv cs.MA (T1 source). The original is at https://arxiv.org/abs/2610.02617.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryWebUIProof is an execution-oriented benchmark for WebUI code generation that uses a UI-agent harness with a plan–act–observe loop to run dense, executable interaction tests in headless browsers across general WebUIs and 3D simulations. Evaluations on 8 commercial LLMs reveal frequent interaction failures; RL training on these signals improves compact models.
Why it mattersOffers a concrete, reproducible harness for testing generated UIs at the interaction level rather than only build success, plus usable training signals for small open models.
Cited by
No citations on record.
