2026-09-30 · America/Los_Angeles · 社区动态 · #20
LessThink-Qwen3-4B: the same model, with far less thinking [P]
I post-trained Qwen3-4B to spend 44% fewer tokens on reasoning, keeping its knowledge and answer style. The whole pipeline ran on one GPU. folks, you can check it out on: https://5ivatej.com/lessthink/
热度 45.3 / 100;排名与评分保留该期记录。