AI Resonance · 阅读最新日报 · AI 入门推荐

2026-09-29 · America/Los_Angeles · 论文 · #9

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be learned jointly, over the whole loop. Supervised fine-tuning (SFT) on reflection trajectories gives a cold start but does not find the high-success repair paths, and naive RL that optimizes only the renderer or only one head leaves most of the gain untapped. We introduce…

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

热度 52.7 / 100;排名与评分保留该期记录。