Reviewer Capability Governs Rejection Targeting, Not Repair Skill: Evidence from LLM Execute-Review-Revise Pipelines
4.20T1 sourcearXiv cs.MA
Source record
Published by arXiv cs.MA (T1 source). The original is at https://arxiv.org/abs/2609.04270.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryControlled pilot on 100 olympiad math problems examining how reviewer capability in execute-review-revise LLM pipelines affects rejection targeting versus repair. A cross-family mid-tier reviewer raised accuracy 12 points (52% to 64%); same-model self-review achieved 0.85 error-detection recall but produced no net gain due to over-rejection and revision inertia.
Why it mattersQuantified, counterintuitive guidance on reviewer model selection in multi-agent setups. Shows self-review is not optimal and identifies capability floors where reviewer roles add cost without effect.
Cited by
No citations on record.
