671B MoE scale
According to the model card, DeepSeek-V3 totals 671B parameters with 37B activated per token, plus a 14B multi-token prediction module.
Open Source Model Profile · deepseek-ai
DeepSeek-V3 is a 671B-total Mixture-of-Experts text-generation model from deepseek-ai. According to the model card, 37B parameters activate per token with MLA and DeepSeekMoE design.
DeepSeek-V3 is published by deepseek-ai as a text-generation model. The captured configuration identifies DeepseekV3ForCausalLM with model type deepseek_v3, and Safetensors metadata reports 684,531,386,000 parameters. According to the model card, it is a 671B Mixture-of-Experts model with 37B parameters activated per token, built on Multi-head Latent Attention and DeepSeekMoE.
According to the model card, DeepSeek-V3 totals 671B parameters with 37B activated per token, plus a 14B multi-token prediction module.
According to the model card, the model combines Multi-head Latent Attention with DeepSeekMoE and an auxiliary-loss-free load-balancing strategy.
According to the model card, training used an FP8 mixed-precision framework and only FP8 weights are provided, with a script for BF16 conversion.
According to the model card, base and chat evaluation tables compare DeepSeek-V3 against DeepSeek-V2, Qwen2.5 72B, LLaMA3.1 405B, Claude-3.5-Sonnet, and GPT-4o.
According to the model card, SGLang, LMDeploy, TensorRT-LLM, and vLLM paths cover FP8 and BF16 inference on NVIDIA and AMD GPUs.
Source: deepseek-ai/DeepSeek-V3
Captured: Unknown. Processed: 2026-09-07T19:34:43.528650+00:00.
Paper Link 👁️ 1. Introduction We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, which were thoroughly validated in DeepSeek-V2. Furthermore, DeepSeek-V3 pioneers an auxiliary-loss-free strategy for load balancing and sets a multi-token prediction training objective for stronger performance. We pre-train DeepSeek-V3 on 14.8 trillion diverse and high-quality tokens, followed by Supervised Fine-Tuning and Rein…
F001F002F003F004F005F006F009F010F011F012F013F017F020F025F027F028F029F030F031F034F041