Harness-RL: Black-Box Reinforcement Learning with Action-Args Decoupling for Central-Agent Multi-Agent Harnesses
3.60T1 sourcearXiv cs.MA
Source record
Published by arXiv cs.MA (T1 source). The original is at https://arxiv.org/abs/2608.29641.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryHarness-RL is a reinforcement learning framework for training central-agent multi-agent LLM harnesses on long-horizon tasks. It introduces Conflict-Aware Policy Optimization (CAPO) that routes gradients separately for action and args tokens, combined with black-box trajectory construction from interface call records. Across seven QA and agentic retrieval benchmarks, authors report F1 scores of 42.93 (Qwen2.5-1.5B) and 47.79 (Qwen2.5-3B).
Why it mattersProposes a training method that decouples action and args optimization for central-agent harnesses, with code and benchmark numbers. Worth checking for teams scaling multi-agent LLM systems.
Cited by
No citations on record.
