2026-09-28 · America/Los_Angeles · 论文 · #2
Disaggregated Quantization: Specializing LLM Prefill and Decode
Prefill and decode reward different approaches to quantization: low-precision arithmetic accelerates prompt processing, while compact weights reduce memory traffic during generation. We propose "disaggregated quantization" (DQ), which specializes computation formats, weights and storage placement to both of these phases. On Qwen 3 and Gemma 3, removing activation quantization specifically on decode improves accuracy on decode-heavy tasks without increasing inference cost. Training separate compute-native prefill weights accelerates prompt processing relative to weight-only inference while…
热度 55.2 / 100;排名与评分保留该期记录。