MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for Parameter-Efficient Reasoning in Compact Models
3.20T1 sourcearXiv cs.MA
Source record
Published by arXiv cs.MA (T1 source). The original is at https://arxiv.org/abs/2607.18006.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryMADA-RL is a post-training framework that uses LoRA adapters to specialize compact models (≤4B parameters) into generator and critic roles for multi-agent debate. Its counterfactual critic advantage trains critics to improve over generator consensus, raising DeepSeek-R1-Distill-Qwen-1.5B accuracy from 39.9% to 41.9% on five math reasoning benchmarks with 16x fewer trainable parameters.
Why it mattersConcrete credit-assignment formulation for multi-agent debate, with a controlled ablation isolating the source of the critic's gains over static baselines.
Cited by
No citations on record.
