---
format: "aidr-story-markdown/v1"
id: "4f1019dc7acfbda693684200b4b22c7595bbeae9829f23c573e9a9fe275eeb6f"
canonical_url: "https://aidr.today/4f1019dc?lang=en"
title: "Marin’s 535B Model Tracks loss within 0.3% in Largest Live Run"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-10-02T04:24:46.000Z"
category: "Models"
topics: ["marin","mixture-of-experts","scaling-laws","open-source","training","research"]
source_urls: ["https://huggingnews.com/ai/marins-535b-model-tracks-loss-within-03percent-in-largest-live-run-f207ebe1","https://x.com/classiclarryd/status/2105794059424215317","https://marketbrief.now/ai/marins-535b-model-tracks-loss-within-03percent-in-largest-live-run-f207ebe1"]
summary: "A new high-parameter neural network is currently executing a public training trajectory to test the predictability of AI scaling laws. Marin has spent 43 days training the model on 18T tokens, employing a publicly pre-registered scaling ladder to forecast the evaluation loss throughout the process. The Mixture of Experts model has 535B parameters and is tracking within 0.3% of its loss forecast despite a 300x extrapolation, with the run now halfway complete. The project's scale increased from an original plan of 360B parameters after developer ravwojdyla used custom kernels to improve hardware efficiency and token throughput."
---

# Marin’s 535B Model Tracks loss within 0\.3% in Largest Live Run

> [Open the canonical story](<https://aidr.today/4f1019dc?lang=en>)

**Published:** 2026-10-02T04:24:46.000Z
**Category:** Models
**Topics:** marin, mixture\-of\-experts, scaling\-laws, open\-source, training, research

## Summary

A new high\-parameter neural network is currently executing a public training trajectory to test the predictability of AI scaling laws\. Marin has spent 43 days training the model on 18T tokens, employing a publicly pre\-registered scaling ladder to forecast the evaluation loss throughout the process\. The Mixture of Experts model has 535B parameters and is tracking within 0\.3% of its loss forecast despite a 300x extrapolation, with the run now halfway complete\. The project's scale increased from an original plan of 360B parameters after developer ravwojdyla used custom kernels to improve hardware efficiency and token throughput\.

## Sources

- [Story source](<https://huggingnews.com/ai/marins-535b-model-tracks-loss-within-03percent-in-largest-live-run-f207ebe1>)
- [Story source](<https://x.com/classiclarryd/status/2105794059424215317>)
- [Story source](<https://marketbrief.now/ai/marins-535b-model-tracks-loss-within-03percent-in-largest-live-run-f207ebe1>)

