Thursday, Oct 1, 2026

← Back to live feed · 1 stories across 1 day Olmo-core 3 serves as the structural foundation for the next generation of AI models developed by Allen AI. The newly released training infrastructure for mixture-of-experts (MoE) architectures enables these systems to scale up to 1 trillion parameters, providing an open-source framework for creating massive model architectures. The system's primary technical shift replaces FSDP weight gather and resharding with DDP and GPU-resident experts. This update allows the expert pool to increase from 8 to 128, growing from 4.6B to 47B total parameters with approximately 3.2B active. By utilizing rowwise expert parallelism and grouped GEMM, the stack increases performance to 2.7x tokens/s/GPU compared to the previous version while maintaining throughput loss below 5%.

Sign in to suggest edits

Key sources

  1. SOURCEmarketbrief.now
Markdown