arXiv 2607.24653

Kimi K3: Open Frontier Intelligence

By Kimi Team, Tongtong Bai, et al.

Published 2026-07-27

Discussion

Read the public discussion and references gathered around this paper.

We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. Together with Stable LatentMoE, which effectively activates 16 of 896 routed experts per token, and refined train…

View the original paper on arXiv