AI Resonance · 阅读最新日报 · AI 入门推荐

2026-09-28 · America/Los_Angeles · 论文 · #2

Disaggregated Quantization: Specializing LLM Prefill and Decode

Prefill and decode reward different approaches to quantization: low-precision arithmetic accelerates prompt processing, while compact weights reduce memory traffic during generation. We propose "disaggregated quantization" (DQ), which specializes computation formats, weights and storage placement to both of these phases. On Qwen 3 and Gemma 3, removing activation quantization specifically on decode improves accuracy on decode-heavy tasks without increasing inference cost. Training separate compute-native prefill weights accelerates prompt processing relative to weight-only inference while…

Disaggregated Quantization: Specializing LLM Prefill and Decode

热度 55.2 / 100;排名与评分保留该期记录。