Skip to content

EthenEthenEthen

Open Source Model Profile · deepseek-ai

DeepSeek-V3

DeepSeek-V3 is a 671B-total Mixture-of-Experts text-generation model from deepseek-ai. According to the model card, 37B parameters activate per token with MLA and DeepSeekMoE design.

Publisher
deepseek-ai
Task
text-generation
Model type
deepseek_v3
License
Unknown
Library
transformers
Publication status
Accepted · not indexed

Model overview

DeepSeek-V3 is published by deepseek-ai as a text-generation model. The captured configuration identifies DeepseekV3ForCausalLM with model type deepseek_v3, and Safetensors metadata reports 684,531,386,000 parameters. According to the model card, it is a 671B Mixture-of-Experts model with 37B parameters activated per token, built on Multi-head Latent Attention and DeepSeekMoE.

Recorded capabilities

671B MoE scale

According to the model card, DeepSeek-V3 totals 671B parameters with 37B activated per token, plus a 14B multi-token prediction module.

MLA plus DeepSeekMoE

According to the model card, the model combines Multi-head Latent Attention with DeepSeekMoE and an auxiliary-loss-free load-balancing strategy.

FP8 training and weights

According to the model card, training used an FP8 mixed-precision framework and only FP8 weights are provided, with a script for BF16 conversion.

Published benchmark tables

According to the model card, base and chat evaluation tables compare DeepSeek-V3 against DeepSeek-V2, Qwen2.5 72B, LLaMA3.1 405B, Claude-3.5-Sonnet, and GPT-4o.

Multi-framework deployment

According to the model card, SGLang, LMDeploy, TensorRT-LLM, and vLLM paths cover FP8 and BF16 inference on NVIDIA and AMD GPUs.

Use cases in the source record

  • Conversational text-generation through the snapshot-listed providers, subject to refreshed availability.
  • Self-hosted FP8 or BF16 inference experiments using the card's documented SGLang, LMDeploy, TensorRT-LLM, or vLLM paths.
  • Benchmark-comparison research that reuses the card's published MMLU, DROP, HumanEval, LiveCodeBench, and math evaluation tables.

Limitations and unknowns

  • No license value was extracted from this record.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.
  • Benchmark figures are publisher-reported model-card tables and were not independently verified by Ethen.
  • According to the model card, the publisher lists 128K context length; Ethen has not independently verified serving behavior at that length.

Source and provenance

Source: deepseek-ai/DeepSeek-V3

Captured: Unknown. Processed: 2026-09-07T19:34:43.528650+00:00.

Paper Link 👁️ 1. Introduction We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, which were thoroughly validated in DeepSeek-V2. Furthermore, DeepSeek-V3 pioneers an auxiliary-loss-free strategy for load balancing and sets a multi-token prediction training objective for stronger performance. We pre-train DeepSeek-V3 on 14.8 trillion diverse and high-quality tokens, followed by Supervised Fine-Tuning and Rein…

F001F002F003F004F005F006F009F010F011F012F013F017F020F025F027F028F029F030F031F034F041