MoE scale description
According to the model card, Kimi K2 is a mixture-of-experts model with 1T total parameters, 32B activated parameters, 61 layers, and 128K context length.
Open Source Model Profile · moonshotai
Kimi-K2-Instruct is a trillion-scale MoE text model from moonshotai. Its model card documents 32B activated parameters and chat plus agentic tool use.
Kimi-K2-Instruct is published by moonshotai as a text-generation model. Safetensors metadata reports 1,026,408,235,864 parameters, with DeepseekV3ForCausalLM and model type kimi_k2. According to the model card, it is the post-trained Kimi K2 variant for chat and agentic use, described as a mixture-of-experts model.
According to the model card, Kimi K2 is a mixture-of-experts model with 1T total parameters, 32B activated parameters, 61 layers, and 128K context length.
The model card says the 1T-parameter MoE was pre-trained on 15.5T tokens and describes MuonClip optimization for stability at scale.
The model card documents tool-calling usage with an OpenAI-compatible client and says the inference engine must support Kimi-K2 native tool-parsing logic.
Source: moonshotai/Kimi-K2-Instruct
Captured: Unknown. Processed: 2026-09-07T19:34:52.408136+00:00.
📰 Tech Blog | 📄 Paper 0. Changelog 2025.8.11 Messages with name field are now supported. We’ve also moved the chat template to a standalone file for easier viewing. 2025.7.18 We further modified our chat template to improve its robustness. The default system prompt has also been updated. 2025.7.15 We have updated our tokenizer implementation. Now special tokens like [EOS] can be encoded to their token ids. We fixed a bug in the chat template that was breaking multi-turn tool calls. 1. Model Introduction Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model with 32 billion activated parameters and 1 trillion total p…
F001F002F003F004F005F006F007F010F015F016F017F018F019F020F021F022F023F024F025F026F027F028F029F032