AI Resonance · 阅读最新日报 · AI 入门推荐

2026-09-29 · America/Los_Angeles · 论文 · #2

MassAlloc Attention: Let Attention Allocate Its Own Compute

FullAttn often assigns negligible normalized mass to much of the causal score space, yet dense kernels execute the complete post-score path after forming each QK tile. We introduce MALA, a fused attention primitive that preserves score access to every legal causal interaction and uses normalized contribution to allocate post-score computation. Forward uses its evolving online-softmax normalizer, while backward reuses the finalized normalizer to derive nested retained support using only standard attention state. A common tolerance governs training and inference, allowing for adaptive…

MassAlloc Attention: Let Attention Allocate Its Own Compute

热度 62.8 / 100;排名与评分保留该期记录。