AI Resonance · 阅读最新日报 · AI 入门推荐

2026-09-30 · America/Los_Angeles · 论文 · #11

VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models

Spoken conversational systems must recover information from prior interactions (i.e., memory), yet relevant information in speech extends beyond what was said to who said it, how it was spoken, and what was audible, information that exists only in the audio signal and cannot be recovered from a transcript. Beyond what to remember, memory also demands diverse operations: retrieving a single fact, integrating evidence across turns, tracking an evolving state. Real interactions further unfold across sessions, meaning information accumulates across distinct episodes rather than a single…

VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models

热度 56.5 / 100;排名与评分保留该期记录。