---
format: "aidr-story-markdown/v1"
id: "377aabd58d63ffd47e9220d3075971083718c0e68f2d1a9d85d7ca1e8ed06b3a"
canonical_url: "https://aidr.today/377aabd5?lang=en"
title: "Reproducing OLMo 3 7B Pre-training in MaxText: case study of large scale training on TPUs"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-24T16:14:05.000Z"
category: "Research"
topics: ["research","training","tpu","jax","olmo","reproducibility"]
source_urls: ["https://developers.googleblog.com/reproducing-olmo-3-7b-pre-training-in-maxtext-case-study-of-large-scale-training-on-tpus/"]
summary: "Learn how MaxText reproduced AI2’s OLMo 3 7B on Google Cloud TPUs, matching PyTorch GPU benchmarks across pre-training with up to 57.4% MFU."
---

# Reproducing OLMo 3 7B Pre\-training in MaxText: case study of large scale training on TPUs

> [Open the canonical story](<https://aidr.today/377aabd5?lang=en>)

**Published:** 2026-09-24T16:14:05.000Z
**Category:** Research
**Topics:** research, training, tpu, jax, olmo, reproducibility

## Summary

Learn how MaxText reproduced AI2’s OLMo 3 7B on Google Cloud TPUs, matching PyTorch GPU benchmarks across pre\-training with up to 57\.4% MFU\.

## Sources

- [Story source](<https://developers.googleblog.com/reproducing-olmo-3-7b-pre-training-in-maxtext-case-study-of-large-scale-training-on-tpus/>)

