Evaluation of Multi-Turn Consistency in LLM Agents: Survival Analysis and Failure-Rationale Taxonomy
3.80T1 sourcearXiv cs.MA
Source record
Published by arXiv cs.MA (T1 source). The original is at https://arxiv.org/abs/2609.29508.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryResearchers evaluated multi-turn consistency in LLM agents using a 20-step delayed-gratification task across 84,540 trajectories and 8 model families. Using survival analysis and a seven-category failure-rationale taxonomy, they found longer deliberation correlates with more intra-rationale contradiction, and failure profiles shift systematically with time and social context.
Why it mattersRigorous evaluation method for agent reliability over multi-turn interactions, with a counterintuitive finding: longer deliberation text correlates with more contradictory rationales. Useful for anyone designing or auditing multi-step agent workflows.
Cited by
No citations on record.
