Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents
3.60T1 sourcearXiv cs.MA
Source record
Published by arXiv cs.MA (T1 source). The original is at https://arxiv.org/abs/2607.12790.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryAuthors propose Double Ratchet, a system that co-evolves evaluation metrics with skill loops for self-improving LLM agents when no reliable automatic verifier exists. The method evolves compositions of drawback detectors under an evolutionary lifecycle, retaining 88–110% of held-out lift achieved with ground-truth metrics across MBPP+, Spider 2.0-Snow, and reference-free report generation.
Why it mattersTargets a real bottleneck in self-evolving agents: the unstated assumption that a reliable evaluation metric already exists. Offers a concrete co-evolution design and quantifies how much lift is lost without ground truth.
Cited by
No citations on record.
