---
format: "aidr-story-markdown/v1"
id: "b233f46ac2f35680b5bd12720015aba2540e61e16868638cd4588858cf7cd676"
canonical_url: "https://aidr.today/b233f46a?lang=en"
title: "SMELT MoE Model Cuts Training FLOPs 18% at 54B Parameter Scale"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-02T08:57:34.000Z"
category: "Research"
topics: ["llm","training","moe"]
source_urls: ["https://huggingnews.com/ai/smelt-moe-model-cuts-training-flops-18percent-at-54b-parameter-scale-cf8bf4f7","https://x.com/iScienceLuvr/status/2095026196698345481","https://x.com/teortaxesTex/status/2095035600969556194"]
summary: "The Sparse MoE Transformer, or SMELT, loops the middle half of its layers twice to accelerate loss reduction during training. This architecture reduces training FLOPs by 6.8% to 18.0% on the compute-optimal frontier compared to unlooped baselines while maintaining parity for per-token FLOPs, total non-embedding parameters, and KV cache. The recipe enables a more efficient path to converge on model weights by repeating computations in the network's center. Researchers scaled SMELT across four configurations, the largest of which reached 54B non-embedding parameters using Chinchilla-style scaling laws. The block-level looping approach differs from an IQuest Research method known as Loop the Loopies, which scaled to 20B A2B by repeating individual layers rather than entire blocks."
---

# SMELT MoE Model Cuts Training FLOPs 18% at 54B Parameter Scale

> [Open the canonical story](<https://aidr.today/b233f46a?lang=en>)

**Published:** 2026-09-02T08:57:34.000Z
**Category:** Research
**Topics:** llm, training, moe

## Summary

The Sparse MoE Transformer, or SMELT, loops the middle half of its layers twice to accelerate loss reduction during training\. This architecture reduces training FLOPs by 6\.8% to 18\.0% on the compute\-optimal frontier compared to unlooped baselines while maintaining parity for per\-token FLOPs, total non\-embedding parameters, and KV cache\. The recipe enables a more efficient path to converge on model weights by repeating computations in the network's center\. Researchers scaled SMELT across four configurations, the largest of which reached 54B non\-embedding parameters using Chinchilla\-style scaling laws\. The block\-level looping approach differs from an IQuest Research method known as Loop the Loopies, which scaled to 20B A2B by repeating individual layers rather than entire blocks\.

## Sources

- [Story source](<https://huggingnews.com/ai/smelt-moe-model-cuts-training-flops-18percent-at-54b-parameter-scale-cf8bf4f7>)
- [Story source](<https://x.com/iScienceLuvr/status/2095026196698345481>)
- [Supporting source](<https://x.com/teortaxesTex/status/2095035600969556194>)

