AI Resonance · 阅读最新日报 · AI 入门推荐

2026-09-21 · America/Los_Angeles · 论文 · #5

OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue

We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In OmniVChat, omni models directly and simultaneously receive audio and video from a user and return text. The user's query is embedded in the audio and video, without a separate text question, external captioning, or speech recognition. Direct audio-visual input reduces external latency and computation while preserving perceptual cues. However, research on OmniVChat faces two constraints: data availability and evaluation. Recordings of people using their own devices are…

OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue

热度 52.4 / 100;排名与评分保留该期记录。