How to Dogfood Your AI Chat Agent: A Three-Layer Evaluation Framework with Goal-Directed NPC Simulation
4.20T1 sourcearXiv cs.HC
Source record
Published by arXiv cs.HC (T1 source). The original is at https://arxiv.org/abs/2608.09939.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA three-layer dogfooding framework for evaluating production LLM chat agents: canonical question-bank testing, random-walk multi-turn evaluation, and a goal-directed NPC simulator with a ten-category failure taxonomy. A three-month case study (257 runs, 108 scenarios) shows layer correlations are weak or negative, and the NPC layer achieves 77% goal completion at $0.17 per run, enabling automated PROMOTE/HOLD/ROLLBACK CI/CD decisions. Authors release prompt templates and a Python replicability guide.
Why it mattersReleased artifacts (prompts, failure taxonomy, Python guide) plus the counter-intuitive finding that canonical response-level scores do not predict multi-turn goal success — directly useful for any team shipping chat agents.
Cited by
No citations on record.
