The Sparse MoE Transformer, or SMELT, loops the middle half of its layers twice to accelerate loss reduction during training. This architecture reduces training FLOPs by 6.8% to 18.0% on the compute-optimal frontier compared to unlooped baselines while maintaining parity for per-token FLOPs, total non-embedding parameters, and KV cache. The recipe enables a more efficient path to converge on model weights by repeating computations in the network's center.

Researchers scaled SMELT across four configurations, the largest of which reached 54B non-embedding parameters using Chinchilla-style scaling laws. The block-level looping approach differs from an IQuest Research method known as Loop the Loopies, which scaled to 20B A2B by repeating individual layers rather than entire blocks.

Sign in to suggest edits

Key sources

  1. SOURCE@iscienceluvr“saving 6.8–18.0% of training FLOPs on the compute-optimal frontier”x.com
  2. SUPPORT@teortaxestex“Loopies repeats individual layers; SMELT repeats on block level”x.com
Markdown