Kimi K3 as an open frontier model with native vision, million-token context, and 2.5× better scaling efficiency compared to K2, with all weights released.
A 35B agent model with horizon scaling and multi-teacher distillation achieves comparable performance to trillion-parameter models on long-horizon benchmarks.
Aligning router rows with the principal singular directions of their associated expert matrices improves the efficiency and stability of Mixture-of-Experts models.