2026-09-29 · America/Los_Angeles · 论文 · #3
Post-Training Leaves Behavioral Shadows on Unrelated Decisions
We find that language models can transfer capabilities through task-unrelated text. Post-training typically improves language models using task-specific data. Prior work on subliminal learning shows that information about these updates can pass through unrelated generations, but has largely focused on traits or preferences using extensive teacher outputs. We introduce Active Taskless Distillation (ATD), which achieves capability transfer using only a single word from the teacher per prompt. ATD probes the behavioral shadow of post-training by selecting prompts where the teacher and student's…
热度 60.2 / 100;排名与评分保留该期记录。