Training Small LLMs as Spatial Multi-Agent Policies
3.60T1 sourcearXiv cs.MA
Source record
Published by arXiv cs.MA (T1 source). The original is at https://arxiv.org/abs/2608.01425.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryPaper training small frozen LLMs as multi-agent policies in spatial cooperative games using symbolic option libraries (synthesized from game source code with mechanically-derived feasibility guards) and per-agent LoRA adapters trained via a per-agent variant of multi-agent GRPO. Behavioral audits show reward gains can mask lack of cooperation, requiring behavior-based evaluation alongside reward.
Why it mattersDocuments that reward curves in LLM multi-agent training can hide one agent idling while the other does the work—useful caveat for anyone building or evaluating agent teams.
Cited by
No citations on record.
