Skip to content

Loopie: MoE Models with Looping Architecture Bridge Efficiency Gap

In brief: Looped Transformers with MoE architecture achieve better results than larger vanilla models at constant training resources.

Loopie is a new series of Mixture-of-Experts Transformers that achieve better results than baseline models through repetition (looping) rather than parameter scaling at the same computational power. The two variants with 20B and 6B parameters demonstrate gold medal performance on mathematics olympiads.

Loopie consists of two Mixture-of-Experts models: a 20B-parameter model with 2B active parameters and a 6B-parameter model with 0.6B active parameters. The central technical advance addresses a long-standing problem of looped Transformer architectures: when training compute is increased N-fold, scaling parameter count N-fold has previously outperformed looping a smaller model N times.

Ablation studies, including direct comparisons with a 30B vanilla model with 3B active parameters, demonstrate that Loopie substantially outperforms this baseline under identical compute budget. This means for engineering teams: with the same training compute, a more effective architecture can be achieved rather than simply scaling model size.

Loopie is equipped with a novel post-training pipeline that enables strong reasoning capabilities. In the International Mathematics Olympiads (IMO) and Physics Olympiads (IPhO) 2025, Loopie achieved gold medal performance without external tools. This suggests improved capabilities in structured logical reasoning that go beyond standard language model training.


Source: arxiv.org · Published 16 July 2026
Lumi AI News — AI-assisted curation in accordance with Art. 50 EU AI Act. Paraphrase and classification via Lumi News Pipeline v1.7.3.

Share on: