From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL
3.40T1 sourcearXiv cs.MA
Source record
Published by arXiv cs.MA (T1 source). The original is at https://arxiv.org/abs/2608.13787.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummarySocialRL is a training recipe for social reasoning applied to a 4B model across six negotiation domains (Deal-or-No-Deal, CaSiNo, Craigslist, Job Interview, Calendar, Marketplace). The 4B matches or exceeds GPT-5 on held-out scenarios, and cross-domain transfer follows structural similarity. Cascade RL and multi-teacher distillation consolidate domain specialists into one model that beats GPT-4.1/5.1/5.2 average utility.
Why it mattersConcrete recipe and empirical transfer map for training small models to negotiate on a user's behalf, with measured gains over frontier models across six domains.
Cited by
No citations on record.
