AI Resonance · 阅读最新日报 · AI 入门推荐

2026-10-02 · America/Los_Angeles · 论文 · #8

On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics

On-policy learning has been argued to reduce catastrophic forgetting, produce sparser parameter updates, and improve generalisation. However, existing comparisons between supervised fine-tuning and reinforcement learning vary many factors simultaneously, making the contribution of rollout policy difficult to isolate. We study the effect of rollout policy in a controlled strong-to-weak distillation setting, by independently varying rollout policy, token-level KL direction, and learning rate across the Llama3 and Qwen2.5 model families and reasoning tasks spanning scientific, medical, and…

On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics

热度 54.9 / 100;排名与评分保留该期记录。