Auditing Belief-Conditioned LLM Agents in Hidden-Information Social Deduction Games
3.80T1 sourcearXiv cs.MA
Source record
Published by arXiv cs.MA (T1 source). The original is at https://arxiv.org/abs/2607.10814.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryPaper presents an audit framework for LLM agents in 9-player Werewolf with code-level information isolation. It tracks external belief states, logs belief-action deviations, and runs 1,080 frozen games across six conditions. Active-belief associates with better good-side outcomes (0.205 to 0.390 win rate, McNemar p<0.001), but mechanism remains unresolved given low action-belief consistency (~0.21).
Why it mattersUseful for builders of multi-agent LLM systems needing auditable belief tracking, structured evidence logs, and offline improvement loops in hidden-information settings.
Cited by
No citations on record.
