Skip to content

EthenEthenEthen

Open Source Model Profile · deepseek-ai

DeepSeek-V4-Flash-0731

DeepSeek-V4-Flash-0731 is a 304B-parameter DeepSeek-V4 text-generation release from deepseek-ai. According to the model card, it is the official Flash release with an attached speculative-decoding module.

Publisher
deepseek-ai
Task
text-generation
Model type
deepseek_v4
License
mit
Library
transformers
Publication status
Accepted · not indexed

Model overview

DeepSeek-V4-Flash-0731 is published by deepseek-ai as a text-generation model. The captured configuration identifies DeepseekV4ForCausalLM with model type deepseek_v4, and Safetensors metadata reports 304,180,418,494 parameters. According to the model card, it is the official DeepSeek-V4-Flash release superseding the preview version, with enhanced agentic capabilities and MIT-licensed weights.

Recorded capabilities

Official Flash release with DSpark module

The model card presents DeepSeek-V4-Flash-0731 as the official release superseding the preview version, with the same structure as the DSpark variant including an attached speculative-decoding module.

Publisher-reported agent benchmarks

According to the model card, publisher-reported figures include 82.7 on Terminal Bench 2.1, 76.7 on Cybergym, 54.4 on DeepSWE, and 70.3 on Toolathlon-Verified, with a stated comparison against the Pro preview and proprietary models.

Documented vLLM and sampling setup

The model card documents vLLM serving with a dspark speculative-config flag and recommends temperature 1.0 with top_p 0.95 for agentic scenarios and up to 384K output tokens at high and max reasoning effort.

DeepSeek-V4 architecture

Captured configuration records DeepseekV4ForCausalLM with model type deepseek_v4 and a transformers library tag for text generation.

Use cases in the source record

  • Agent-oriented text-generation experiments using the release the model card describes as having enhanced agentic capabilities and publisher-reported agent benchmarks.
  • Local or vLLM-based deployment trials following the card's documented DSpark speculative-config flag, sampling values, and output-length guidance.

Limitations and unknowns

  • No context-window value was extracted from this record.
  • Provider state is historical snapshot data and should be refreshed before being presented as current availability.
  • Benchmark figures come from the publisher model card and have not been independently measured or verified by Ethen.
  • No pricing, VRAM, latency, or supported-language values were extracted from this record.

Source and provenance

Source: deepseek-ai/DeepSeek-V4-Flash-0731

Captured: Unknown. Processed: 2026-09-07T19:34:43.593067+00:00.

DeepSeek-V4-Flash-0731 Technical Report 👁️ Introduction DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash , superseding the preview version, with substantially enhanced agentic capabilities. It has the same model structure as DeepSeek-V4-Flash-DSpark , i.e. it comes with a speculative decoding module attached. DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) on benchmarks listed below despite its far smaller activated parameter count, and is broadly competitive with the strongest proprietary models available. Benchmark DeepSeek-V4-Flash-0731 DeepSeek-V4-Flash (Preview) DeepSeek-V4-Pro (Preview) GLM-5.2…

F001F002F003F004F005F006F007F008F010F011F012F013F014F015F017F018