2026-09-30 · America/Los_Angeles · 论文 · #17
SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation
On-policy distillation (OPD) reduces train-test state mismatch by training a student on its own generated trajectories, but weak students may visit teacher-misaligned prefixes where supervision is less representative. We introduce SAKI (Supervision Allocation with KL-constrained Interpolation), which combines a KL-constrained teacher-guided rollout with maximal coupling and reuses realized accept/correction events to route token-level supervision. Accepted positions retain sampled-token reverse-KL supervision, while correction positions receive direct supervision on the teacher's…
热度 53.4 / 100;排名与评分保留该期记录。