The Sparse MoE Transformer, or SMELT, loops the middle half of its layers twice to accelerate loss reduction during training. This architecture reduces training FLOPs by 6.8% to 18.0% on the compute-optimal frontier compared to unlooped baselines while maintaining parity for per-token FLOPs, total non-embedding parameters, and KV cache. The recipe enables a more efficient path to converge on model weights by repeating computations in the network's center.
Researchers scaled SMELT across four configurations, the largest of which reached 54B non-embedding parameters using Chinchilla-style scaling laws. The block-level looping approach differs from an IQuest Research method known as Loop the Loopies, which scaled to 20B A2B by repeating individual layers rather than entire blocks.