In brief: Kimi K3 as an open frontier model with native vision, million-token context, and 2.5× better scaling efficiency compared to K2, with all weights released.
Zhao Hui (Kimi manufacturer) presents Kimi K3, an open-source language model with 2.8 trillion parameters and native capabilities for image processing. The mixture-of-experts system activates 104 billion parameters and achieves a million-token context window.
Kimi K3 is a mixture-of-experts model with 2.8 trillion parameters, of which 104 billion are activated per inference step. The system uses Stable LatentMoE to select 16 out of 896 routed experts per token. With a 1-million-token context window and integrated visual capabilities, it addresses tasks with long-range dependencies across multiple steps.
Technically, Kimi K3 is based on two architecture elements: Kimi Delta Attention improves information flow across sequence length and model depth, while Attention Residuals provide additional stability. Scaling efficiency improved compared to K2 by approximately a factor of 2.5. Post-training comprises reinforcement learning across three domains (general, agentic tasks, code) at multiple effort levels, enabling compositional generalization and robust multi-step execution.
The infrastructure realized several innovations: algorithm-system co-design for Kimi Delta Attention, perfectly balanced expert-parallel training with efficient memory management, million-token agentic reinforcement learning with persistent rollout and sandbox states, and specialized deployment optimizations.
In benchmarks, Kimi K3 demonstrates frontier performance on code, agent, knowledge, reasoning, and vision tasks. Overall performance positioned behind Claude Fable 5 and GPT-5.6 Sol, but consistently outperforms other evaluated open and proprietary models. The complete model weights have been released.
Source: arxiv.org · Published July 26, 2026
Lumi AI News — AI-assisted curation in accordance with Article 50 EU AI Act. Paraphrase and classification by Lumi News Pipeline v1.7.3.