Living-Harness Is an Interactive-Agent Evolver
3.40T1 sourcearXiv cs.MA
Source record
Published by arXiv cs.MA (T1 source). The original is at https://arxiv.org/abs/2607.26598.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryLiving-Harness is a self-evolving agent framework that converts each completed trajectory and evaluator signals into posterior evidence for bounded harness updates. It writes episodic memory and a state graph, then retrieves updated harness state to guide future interactions while keeping base tools and context frozen. Tests on eight interactive environments from tau^2-Bench and MultiWOZ-2.4 show 10.07 and 9.91 percentage point gains in average Pass@1 over the strongest interactive baseline.
Why it mattersConcrete pattern for letting agent harnesses accumulate procedural repairs from past failures, with measured gains on standard interactive-agent benchmarks. Worth reading if you build long-running agents.
Cited by
No citations on record.
