AI Resonance · 阅读最新日报 · AI 入门推荐

2026-09-26 · America/Los_Angeles · 官方发布 · #14

Reproducing OLMo 3 7B Pre-training in MaxText: case study of large scale training on TPUs

The MaxText team successfully reproduced AI2’s OLMo 3 7B language model from scratch on Google Cloud TPUs using JAX/XLA, precisely matching the original PyTorch-on-GPU reference across pre-training and mid-training stages on all held-out evaluations. The implementation achieved up to 57.4% Model Flops Utilization (MFU) and demonstrated robust infrastructure portability by surviving mid-run cluster resizes and cross-generation TPU shifts without requiring recipe alterations. Crucially, the exercise proved the necessity of comprehensive held-out validation by catching a silent data-loader…

Reproducing OLMo 3 7B Pre-training in MaxText: case study of large scale training on TPUs

热度 32.0 / 100;排名与评分保留该期记录。