2026-10-03 · America/Los_Angeles · 社区动态 · #13
Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
I wanted to see how far I could push a fairly ordinary laptop with a huge MoE model. Turns out, Qwen3.8 Flash Next 176B can run on: RTX 3080 Laptop — 16GB VRAM 32GB system RAM SSD No 128GB/256GB RAM workstation and no multi-GPU setup. I’m running it with TensorSharp, my open-source local LLM inference engine: TensorSharp on GitHub The interesting part for me wasn't simply getting a 176B model to load. I wanted to make a model much larger than both available VRAM and RAM actually usable. The approach is basically: Quantization + MoE-aware unified scheduling across cache, VRAM, system RAM, and…
热度 45.3 / 100;排名与评分保留该期记录。