---
format: "aidr-story-markdown/v1"
id: "316351daeab76486e9c53c4f5aa7330200b1daca08d25c1a5683ce883e074aa9"
canonical_url: "https://aidr.today/316351da?lang=en"
title: "Allen AI Releases Open Training Stack for 1 Trillion Parameter MoE Models"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-10-01T20:39:38.000Z"
category: "Releases"
topics: ["open-source"]
source_urls: ["https://huggingnews.com/ai/allen-ai-releases-open-training-stack-for-1-trillion-parameter-moe-model-4ffb6b8e","https://marketbrief.now/ai/allen-ai-releases-open-training-stack-for-1-trillion-parameter-moe-model-4ffb6b8e"]
summary: "← Back to live feed · 1 stories across 1 day Olmo-core 3 serves as the structural foundation for the next generation of AI models developed by Allen AI. The newly released training infrastructure for mixture-of-experts (MoE) architectures enables these systems to scale up to 1 trillion parameters, providing an open-source framework for creating massive model architectures. The system's primary technical shift replaces FSDP weight gather and resharding with DDP and GPU-resident experts. This update allows the expert pool to increase from 8 to 128, growing from 4.6B to 47B total parameters with approximately 3.2B active. By utilizing rowwise expert parallelism and grouped GEMM, the stack increases performance to 2.7x tokens/s/GPU compared to the previous version while maintaining throughput loss below 5%."
---

# Allen AI Releases Open Training Stack for 1 Trillion Parameter MoE Models

> [Open the canonical story](<https://aidr.today/316351da?lang=en>)

**Published:** 2026-10-01T20:39:38.000Z
**Category:** Releases
**Topics:** open\-source

## Summary

← Back to live feed · 1 stories across 1 day Olmo\-core 3 serves as the structural foundation for the next generation of AI models developed by Allen AI\. The newly released training infrastructure for mixture\-of\-experts \(MoE\) architectures enables these systems to scale up to 1 trillion parameters, providing an open\-source framework for creating massive model architectures\. The system's primary technical shift replaces FSDP weight gather and resharding with DDP and GPU\-resident experts\. This update allows the expert pool to increase from 8 to 128, growing from 4\.6B to 47B total parameters with approximately 3\.2B active\. By utilizing rowwise expert parallelism and grouped GEMM, the stack increases performance to 2\.7x tokens/s/GPU compared to the previous version while maintaining throughput loss below 5%\.

## Sources

- [Story source](<https://huggingnews.com/ai/allen-ai-releases-open-training-stack-for-1-trillion-parameter-moe-model-4ffb6b8e>)
- [Story source](<https://marketbrief.now/ai/allen-ai-releases-open-training-stack-for-1-trillion-parameter-moe-model-4ffb6b8e>)

