This week, the focus is on an optimization technique for search agents: on-policy self-distillation (OPSD). 18 papers discuss it, up from 6 four weeks ago, and the term appears in three of the most dynamic topics at the moment - Information Retrieval and Search Behavior (88 papers, +60 in one month), Multi-Agent Systems and Negotiation (104 papers, +45), and language-model agents (18 papers, +11).
The titles show an exploration of the method's limits:
- Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents
- Latent On-Policy Self-Distillation
- Adaptive Supervised Anchoring for On-Policy Self-Distillation
Two variants are emerging: the use of privileged information (23 papers, +13) to guide distillation, and token-level supervision (20 papers, +11) to refine the learning signal. These approaches aim to stabilize agent training in environments where rewards are sparse or noisy, such as information retrieval or negotiation between agents.
