← Back to live feed1 story
Friday, Oct 2, 2026
A new high-parameter neural network is currently executing a public training trajectory to test the predictability of AI scaling laws. Marin has spent 43 days training the model on 18T tokens, employing a publicly pre-registered scaling ladder to forecast the evaluation loss throughout the process.
The Mixture of Experts model has 535B parameters and is tracking within 0.3% of its loss forecast despite a 300x extrapolation, with the run now halfway complete. The project's scale increased from an original plan of 360B parameters after developer ravwojdyla used custom kernels to improve hardware efficiency and token throughput.
Sign in to suggest edits
Key sources
- SOURCE@classiclarryd“Marin kicked off the largest live-streamed pretraining run in history: a 535B MoE trained on 18T tokens”x.com
- SOURCEmarketbrief.now