AI Resonance · 阅读最新日报 · AI 入门推荐

2026-09-20 · America/Los_Angeles · 论文 · #4

AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents

Comprehensive video understanding is crucial for advancing artificial intelligence toward the intricate dynamics of the physical world. While recent advances in Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in video understanding, existing benchmarks remain confined to simple scene-level queries or global summaries that require only single-step inference. Real-world video understanding involves more challenging tasks that require multi-hop multimodal reasoning, and there is a critical absence of video benchmarks equipped to rigorously evaluate these…

AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents

热度 14.8 / 100;排名与评分保留该期记录。