AI Resonance · 阅读最新日报 · AI 入门推荐

2026-09-18 · America/Los_Angeles · 论文 · #13

What Does Privileged Information Add to On-Policy Self-Distillation?

On-policy self-distillation (OPSD) lets a language model learn from a frozen copy of itself that sees an answer or a worked solution. Giving the teacher this extra information seems to offer the student more to learn, but how much does it add beyond distillation itself? To isolate that contribution, we construct AMPLE-Math, a reusable suite of 5,319 mathematical problems with six reasoning views that share the same answer, and compare each view with matched reference-free distillation. With a thinking-enabled teacher supervising direct-response rollouts, reference-free distillation accounts…

What Does Privileged Information Add to On-Policy Self-Distillation?

热度 43.9 / 100;排名与评分保留该期记录。