AI Resonance · 阅读最新日报 · AI 入门推荐

2026-09-28 · America/Los_Angeles · 论文 · #5

Block Sparse Attention with Log-Linear Complexity

Scaling language models to long contexts is limited by the quadratic cost of self-attention. Block sparse attention offers an efficient alternative, but selecting the retained blocks remains a bottleneck. Conventional block selection requires scoring all query-block pairs and therefore remains quadratic in sequence length. To address this issue, we propose PISA, a block-sparse attention mechanism that employs a pyramid Top-K selection strategy. The main idea is to gradually narrow down the candidates across different levels, making it more efficient to find the most relevant keys.…

Block Sparse Attention with Log-Linear Complexity

热度 42.1 / 100;排名与评分保留该期记录。