What Confidence Routing Is Actually Doing: Auditing Routing, Calibration, and Commitment in Multi-Agent Deliberation
4.40T1 sourcearXiv cs.MA
Source record
Published by arXiv cs.MA (T1 source). The original is at https://arxiv.org/abs/2609.27822.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryAn audit of confidence-routed multi-agent broadcast protocols across 4,181 olympiad-math traces and a 2-by-2 model/benchmark grid finds that routing discrimination, probability calibration, and public commitment are three distinct properties that must be measured separately before raw-confidence routing is deployed, as fixed routers and cross-fitted calibration do not recover missing discrimination in the Gemma cells.
Why it mattersConcrete trace-level evidence with specific numbers that the popular 'highest-confidence agent speaks next' pattern is more fragile across models and tasks than its adoption suggests.
Cited by
No citations on record.
