2026-09-23 · America/Los_Angeles · 论文 · #3
Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings
In this report, we introduce Ovis-Embedding, a state-of-the-art omni-modal embedding family built on native integration of text, image, video, and audio. Instead of assembling separate modality towers, Ovis-Embedding uses a shared multimodal backbone to encode different modalities in a common representation space. Specifically, we make three key advances: (1) native omni-modal initialization: we adopt a pretrained Qwen-omni model as the embedding backbone and adapt it through contrastive training with low-rank initialization; (2) data-centric omni-modal training: we construct a broad,…
热度 52.9 / 100;排名与评分保留该期记录。