Skip to content

EthenEthenEthen

Open Source Model Profile · deepseek-ai

DeepSeek-R1-0528

DeepSeek-R1-0528 is a 684.53B-parameter deepseek_v3 reasoning release from deepseek-ai. Its card documents deeper reasoning traces and stronger AIME performance.

Publisher
deepseek-ai
Task
text-generation
Model type
deepseek_v3
License
mit
Library
transformers
Publication status
Accepted · not indexed

Model overview

DeepSeek-R1-0528 is published by deepseek-ai as a deepseek_v3 text-generation release. The captured configuration identifies DeepseekV3ForCausalLM and Safetensors metadata reports 684,531,386,000 parameters, or about 684.53B. The model card presents it as a reasoning-focused minor upgrade with stronger mathematics, programming, and general-logic performance.

Recorded capabilities

Reasoning-depth upgrade

The model card describes improved reasoning depth through added compute and post-training optimization, with reduced hallucination, enhanced function calling, and revised vibe-coding experience.

Publisher-reported AIME improvement

According to the model card, AIME 2025 accuracy increased from 70% to 87.5% alongside deeper thinking traces, evaluated with 64K generation length and repeated sampling.

Distilled variant and system prompt

The model card also documents DeepSeek-R1-0528-Qwen3-8B distilled from its chain-of-thought and notes that system prompts are now supported.

Use cases in the source record

  • Complex reasoning workflows in mathematics, programming, and general logic, which the model card names as improved areas.
  • Function-calling and system-prompt-guided conversational deployments following the publisher's documented usage changes.

Limitations and unknowns

  • Reasoning, benchmark, and comparison statements are publisher claims, not Ethen-measured results, and should not be generalized beyond the reported tests.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.
  • No VRAM, quantization, pricing, context-window, or latency figures were extracted beyond the 64K generation-length evaluation setting.

Source and provenance

Source: deepseek-ai/DeepSeek-R1-0528

Captured: Unknown. Processed: 2026-09-07T19:34:43.379163+00:00.

DeepSeek-R1-0528 Paper Link 👁️ 1. Introduction The DeepSeek R1 model has undergone a minor version upgrade, with the current version being DeepSeek-R1-0528. In the latest update, DeepSeek R1 has significantly improved its depth of reasoning and inference capabilities by leveraging increased computational resources and introducing algorithmic optimization mechanisms during post-training. The model has demonstrated outstanding performance across various benchmark evaluations, including mathematics, programming, and general logic. Its overall performance is now approaching that of leading models, such as O3 and Gemini 2.5 Pro. Compare…

F001F002F003F004F005F006F007F009F010F011F012F013F014F015F016F018