A Pinch of SFT, A Dash of RL: When Reinforcement Learning Helps Long-Horizon Advertising Agents
3.60T1 sourcearXiv cs.MA
Source record
Published by arXiv cs.MA (T1 source). The original is at https://arxiv.org/abs/2609.22194.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryPaper on balancing SFT and RL for enterprise analytics agents. Proposes a diagnostic routing framework categorizing features into Imitation, Lift, and Discovery regimes to select training strategy. On GPT-OSS 120B, targeted SFT+RL improved 7/8 advertiser skills, reduced standard leakage from 11.8% to 2.9%, and used 43% less RL compute than uniform application.
Why it mattersConcrete diagnostic for deciding when RL helps versus perturbs SFT-calibrated agent skills, with measured compute savings and leakage drops. Useful for teams training long-horizon tool-use agents.
Cited by
No citations on record.
