2026-09-26 · America/Los_Angeles · 社区动态 · #7
Splash 1.1.0 released, GGUF quants support, MLX import and more
On my M5 Pro 64GB I can comfortably work in an agentic setup with the Qwen3.8 27B model in good quality (Unsloth UD-Q4_K_XL) at a decent speed of 50 t/s. Splash combines optimized kernels, excellent speculative decoding, a well-implemented prefix cache, and mixed-weight support in a single program. To me, this is a breakthrough in local inference on Apple Silicon. https://github.com/incoai/splash/releases/tag/1.1.0
热度 47.3 / 100;排名与评分保留该期记录。