2026-09-29 · America/Los_Angeles · 论文 · #2
MassAlloc Attention: Let Attention Allocate Its Own Compute
FullAttn often assigns negligible normalized mass to much of the causal score space, yet dense kernels execute the complete post-score path after forming each QK tile. We introduce MALA, a fused attention primitive that preserves score access to every legal causal interaction and uses normalized contribution to allocate post-score computation. Forward uses its evolving online-softmax normalizer, while backward reuses the finalized normalizer to derive nested retained support using only standard attention state. A common tolerance governs training and inference, allowing for adaptive…
热度 62.8 / 100;排名与评分保留该期记录。