AI Resonance · 阅读最新日报 · AI 入门推荐

2026-09-25 · America/Los_Angeles · 论文 · #9

Rufus-Air: An Open LLM Post-Training Recipe

Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. We document the data, reward design, infrastructure, stage order, and stagewise results needed to reproduce the recipe. Stages progress from basic to advanced capabilities and from hard, verifiable rewards to softer judge-based signals. Training builds on open-source components and public data, much of it used as released, without new human…

Rufus-Air: An Open LLM Post-Training Recipe

热度 38.4 / 100;排名与评分保留该期记录。