AI Resonance · 阅读最新日报 · AI 入门推荐

2026-09-21 · America/Los_Angeles · 论文 · #15

Calibrating Teacher--Student Discrepancy for On-Policy Distillation

On-policy distillation (OPD) improves reasoning models by learning the token-level discrepancy between a stronger teacher and an on-policy student. However, this discrepancy does not purely reflect the capability gap between the teacher and the student: it also contains deviations arising from the teacher itself, which are consequently mixed into the observed teacher--student discrepancy and indiscriminately learned by standard OPD during training. This issue is further exacerbated by privileged OPD, where privileged information induces larger teacher-side likelihood shifts, thereby…

Calibrating Teacher--Student Discrepancy for On-Policy Distillation

热度 35.0 / 100;排名与评分保留该期记录。