AI Resonance · 阅读最新日报 · AI 入门推荐

2026-09-20 · America/Los_Angeles · 论文 · #13

Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation

Conversational speech depends on dialogue context and the listener's immediately preceding behavior. We propose ReACT-TTS, a two-stage framework that uses a one-second pre-response listener facial sequence to plan the next utterance's emotion and prosody before speech realization. On a strict dyadic MELD protocol, Temporal conditioning yields higher mean macro-F1 and VAD concordance than Text-only across ten seeds, while accuracy remains essentially unchanged. Ablations show that temporal modeling performs best among the visual variants and that an explicit early-to-late difference is…

Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation

热度 14.0 / 100;排名与评分保留该期记录。