AI Resonance · 阅读最新日报 · AI 入门推荐

2026-09-29 · America/Los_Angeles · 论文 · #5

Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation

Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) merges them by letting the specialists teach one student: the student answers each prompt, and the specialist for that prompt's domain gives feedback on every token. This routing decides which specialist teaches, but not how strongly its feedback moves the shared student. In Qwen3.5 models at three sizes, we find that MOPD's student…

Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation

热度 59.8 / 100;排名与评分保留该期记录。