AI Resonance · 阅读最新日报 · AI 入门推荐

2026-10-01 · America/Los_Angeles · 论文 · #1

Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence

Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words. However, unlike textual CoT, latent reasoning is not directly observable, making it difficult to supervise what latent tokens learn. In this work, we first conduct a thorough analysis of latent-token behavior and identify a latent evidence-credit gap: latent tokens respond only weakly to image perturbations that alter the correct answer. We hypothesize that this issue stems from the lack of…

Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence

热度 60.8 / 100;排名与评分保留该期记录。