AI Resonance · 阅读最新日报 · AI 入门推荐

2026-09-18 · America/Los_Angeles · 论文 · #14

Region-Level Policy Optimization for Fine-grained MLLM Perception

Fine-grained visual perception in MLLMs is commonly improved by raising the resolution, but the added visual tokens inflate vision-encoding and language-model prefilling costs. We show that the two operations underlying fine-grained perception, localizing the region of interest (RoI) and recognizing its content, have different resolution requirements. In a controlled diagnostic, localization tolerates roughly 3 to 4 times stronger token compression than recognition, which motivates localizing from a coarse view and concentrating resolution on the selected evidence. Decoding coordinates with…

Region-Level Policy Optimization for Fine-grained MLLM Perception

热度 43.5 / 100;排名与评分保留该期记录。