AI Resonance · 阅读最新日报 · AI 入门推荐

2026-09-24 · America/Los_Angeles · 社区动态 · #11

Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second

A while ago I posted 15 tok/s output and 100-120 tok/s prompt processing with the IQ3_XXS quant on a 12GB RTX 5070 using llama.cpp. Since then I built my own inference engine for this one model and this kind of PC. The same IQ3_XXS now runs at ~65 tok/s output and ~430 tok/s prompt processing, and the 2-bit quants run faster still using RCO-GSQ quantization. Using: 64GB DDR5 (5600) 12GB RTX 5070 SFF (Gigabyte) Ryzen 5 7600 CPU Windows Output (tokens/s) on 128K context: Q2_0 (equivalent to unsloth Q3): 65.1 IQ2_XS (equivalent to unsloth Q4): 52.0 IQ3_XXS (equivalent to unsloth Q5): 44.8…

Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second

热度 46.5 / 100;排名与评分保留该期记录。