2026-09-29 · America/Los_Angeles · 论文 · #18
REALM: A Coarse-to-Fine Generative Framework for Embodied Reactive Listening
Generating responsive listener facial motion is an important task for embodied conversational AI. Two modeling challenges are central: accounting for the timing of speaker cues while maintaining continuity with the listener's ongoing motion, and capturing locally variable facial events alongside the overall motion trajectory. Listener responses may follow preceding cues with a temporal lag, while brief expressions and blinks introduce variation that is difficult to predict deterministically. These challenges motivate a framework that combines history-aware temporal alignment with stochastic…
热度 47.0 / 100;排名与评分保留该期记录。