MASPRM: Multi-Agent System Process Reward Model
3.80T1 sourcearXiv cs.MA
Source record
Published by arXiv cs.MA (T1 source). The original is at https://arxiv.org/abs/2510.24803.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryMASPRM is a process reward model that scores intermediate agent messages in multi-agent systems to guide step-level beam search and Monte Carlo Tree Search at inference time. Trained from MCTS rollouts with only terminal rewards, it improves over size-matched outcome reward models by 2–14.5 points across GSM8K, MATH, MMLU, and LogiQA.
Why it mattersConcrete training recipe and benchmarks for a relatively underexplored problem: credit assignment across agents in multi-agent inference. Code is released. Relevant to anyone building or evaluating orchestrated LLM agent pipelines.
Cited by
No citations on record.
