AI Resonance · 阅读最新日报 · AI 入门推荐

2026-09-20 · America/Los_Angeles · 论文 · #15

Visual Navigation Transformer with Pose Attention

Learned navigation policies typically consume observations as a temporally ordered history, with positional encodings tying each observation to when it was seen, making it difficult to reuse experience from earlier traversals of an environment. Systems that do reuse such experience usually construct an explicit representation, such as a map or a topological graph, and plan on it. We propose VNT-PA (Visual Navigation Transformer with Pose Attention), a transformer planner whose context is a set of depth keyframes indexed by camera pose. With camera poses as positional encoding, attention…

Visual Navigation Transformer with Pose Attention

热度 12.6 / 100;排名与评分保留该期记录。