2026-10-01 · America/Los_Angeles · 论文 · #6
Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering
Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close this gap by injecting 3D awareness into MLLMs, either by boosting fine-grained pixel-level cross-view correspondence or by fusing features from 3D geometry foundation models, yet a substantial gap to human reasoning persists. In this work, we revisit human spatial reasoning,…
热度 56.1 / 100;排名与评分保留该期记录。