2026-09-25 · America/Los_Angeles · 社区动态 · #2
Qwengram-0.8B: I transferred Qwen3.8 Flash-Next’s n-gram memory into Qwen3.5-0.8B — 5.05% lower validation perplexity
I’ve been experimenting with whether Qwen3.8-Flash-Next’s pretrained PLE n-gram memory can improve a much smaller Qwen3.5-0.8B model. I trained the 0.8B setup with limited resources, mostly using free Kaggle notebook GPUs. The setup keeps both the Qwen3.5-0.8B backbone and the roughly 51B-parameter PLE memory frozen. A small R=1 reader is trained at decoder layers 3 and 9, with a token-dependent linear gate controlling the later injection. There is no backbone fine-tuning . The current balanced checkpoint uses a reader trained for 15M tokens and a separately calibrated dynamic gate. On the…
热度 48.0 / 100;排名与评分保留该期记录。