DeepSeek-R1

DeepSeek-R1-0528-Qwen3-8B

by DeepSeek · Current open-weight model; downloadable checkpoint

An MIT-licensed 8-billion-parameter reasoning model distilled from DeepSeek-R1-0528 into Qwen3-8B, designed for local mathematics, coding, logical reasoning, and experimentation.

Text Reasoning Coding
DeepSeek-R1-0528-Qwen3-8B is a compact reasoning model from DeepSeek that focuses on mathematics, coding, and difficult multi-step problems. It uses the Qwen3-8B architecture but was post-trained with reasoning traces distilled from DeepSeek-R1-0528. The result is a downloadable checkpoint that can be run with local inference tools instead of requiring a DeepSeek-hosted API.
Outputs

What DeepSeek-R1-0528-Qwen3-8B can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Fine-tuning
Model profile

Performance characteristics

8/10 Reasoning
7/10 Coding
7/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family DeepSeek-R1
Model type Reasoning
Context window 131K tokens
Maximum output 64K tokens
Release date 2025-05-28
Status Current open-weight model; downloadable checkpoint
Knowledge cutoff notes

No authoritative knowledge-cutoff date was published for this exact distilled checkpoint in the reviewed first-party model documentation.

Model notes

This is an open-weight distilled model, not the full DeepSeek-R1-0528 checkpoint. DeepSeek distilled reasoning traces from DeepSeek-R1-0528 into Qwen3-8B Base. The model card reports 8B parameters, bfloat16 weights, Qwen3 architecture, and a 131,072-token configured context window. DeepSeek's benchmark methodology used a 64,000-token maximum generation length. The checkpoint is MIT-licensed and supports commercial use and distillation. No official DeepSeek hosted API price is specified for this exact checkpoint. Function calling and web-search instructions documented in the model card describe the DeepSeek-R1 product environment and should not be assumed to be native capabilities of this standalone checkpoint. Editorial scores are comparative estimates, not vendor ratings.

Model guide

DeepSeek-R1-0528-Qwen3-8B: Open-Weight Reasoning for Local AI

DeepSeek-R1-0528-Qwen3-8B is an 8-billion-parameter open-weight reasoning model distilled from DeepSeek-R1-0528 into the Qwen3-8B architecture. Released under the MIT license, it is designed for local mathematics, coding, logical reasoning, experimentation, and smaller deployments that need stronger reasoning than a typical model of its size.

What is DeepSeek-R1-0528-Qwen3-8B?

DeepSeek-R1-0528-Qwen3-8B is an open-weight language model released by DeepSeek in May 2025. It contains approximately 8 billion parameters and is intended to generate text with an emphasis on deliberate, multi-step reasoning. In practical terms, it is aimed at tasks such as solving mathematics problems, writing and analyzing code, and working through logical questions rather than simply producing short conversational replies.

The model was created by distilling reasoning traces from DeepSeek-R1-0528 into Qwen3-8B Base. Distillation is a training approach in which a smaller model learns patterns from the outputs or reasoning behavior of a larger or more capable source model. DeepSeek-R1-0528-Qwen3-8B therefore combines the Qwen3-8B model architecture with post-training intended to improve reasoning quality.

This is a standalone downloadable checkpoint, not the full DeepSeek-R1-0528 model and not a provider-managed hosted API product. The official checkpoint is published through DeepSeek's Hugging Face organization and is licensed under the MIT license, which permits local use, research, commercial use, and further distillation or fine-tuning subject to the license terms.

Where it fits in DeepSeek's lineup

DeepSeek-R1-0528-Qwen3-8B belongs to the DeepSeek-R1 reasoning family, but its defining positioning is its relatively small size. Instead of using the complete DeepSeek-R1-0528 checkpoint, DeepSeek transferred reasoning behavior into an 8-billion-parameter Qwen3-based model. That makes it relevant to users who want a more accessible local model and can accept lower overall capacity than a much larger reasoning system.

The model's architecture is described as identical to Qwen3-8B, while its tokenizer configuration and configuration files come from the DeepSeek checkpoint. This distinction matters when deploying it: users should use the files supplied with the DeepSeek-R1-0528-Qwen3-8B repository rather than replacing them with configuration files from the original Qwen3 model.

Its value is therefore not that it provides the largest possible model capacity. Its value is the reported reasoning improvement relative to its size. DeepSeek reported that it exceeded the original Qwen3-8B by 10 percentage points on AIME 2024 and matched the reported AIME 2024 performance of Qwen3-235B-Thinking in its published comparison. Those are provider-reported benchmark results under a stated evaluation setup, not a guarantee of performance on every prompt or application.

Architecture and verified limits

The official configuration identifies a Qwen3-based transformer with approximately 8 billion parameters. It specifies 36 transformer layers, a 4,096-dimensional hidden state, grouped-query attention, and bfloat16 weights. Grouped-query attention is an attention design that can reduce memory and processing requirements compared with some conventional attention arrangements, although actual performance depends on the inference engine, hardware, prompt length, and quantization.

The configured context window is 131,072 tokens. A context window is the amount of text the model can consider across the prompt and generated conversation. This is a configuration-level limit, not a promise that every local runtime will support the full length efficiently. Memory use and response speed can become significant concerns as prompts approach the maximum.

The researched deployment information lists a maximum output length of 64,000 tokens for the published benchmark methodology. This should be distinguished from the 131,072-token context configuration. A deployment must account for both the input and output together, and a particular runtime may impose a lower effective limit.

SpecificationVerified information
ProviderDeepSeek
Release dateMay 28, 2025
Model sizeApproximately 8 billion parameters
ArchitectureQwen3-based transformer
Configured context length131,072 tokens
Published maximum generation setting64,000 tokens
LicenseMIT
Native outputText

Reasoning and benchmark performance

Reasoning is the model's primary specialization. It is designed to spend more generation effort working through a problem before presenting an answer, which can help with mathematical derivations, code-related analysis, and logic-heavy tasks. This behavior can also produce longer responses and higher latency than a model optimized mainly for quick chat or short completion.

DeepSeek's published comparison reported the following results for DeepSeek-R1-0528-Qwen3-8B: 86.0 on AIME 2024, 76.3 on AIME 2025, 61.5 on HMMT February 2025, 61.1 on GPQA Diamond, and 60.5 on LiveCodeBench. The figures are useful for understanding the intended capability profile, particularly its emphasis on mathematics and coding. They should not be treated as universal accuracy scores, because real-world results vary with prompt design, sampling settings, evaluation method, and task difficulty.

For coding, the model is best understood as a reasoning-oriented coding assistant rather than a complete software engineering platform. It can be useful for explaining code, proposing implementations, debugging examples, and reasoning about algorithms. The supplied research does not establish native execution, repository access, automated testing, or guaranteed function calling, so those capabilities should come from the surrounding application rather than being assumed to be built into the checkpoint.

Input, output, and tool support

DeepSeek-R1-0528-Qwen3-8B is a text-generation model. It accepts text prompts and produces text output. It does not natively generate images, audio, video, embeddings, music, or speech. It is consequently a poor fit when a project requires a single model to directly understand or produce multiple media types.

The checkpoint can be used with Transformers, vLLM, SGLang, Docker Model Runner, and compatible quantization tools. The official model card also provides examples for OpenAI-compatible local servers based on vLLM and SGLang. These integrations can make the model easier to connect to existing applications, but an OpenAI-compatible interface is an interface convention, not evidence that DeepSeek-R1-0528-Qwen3-8B is a hosted OpenAI-style service.

Tool use and function calling should be treated cautiously. The researched model documentation does not verify guaranteed native tool calling for this standalone checkpoint. References to function calling or web search in DeepSeek documentation may describe the broader DeepSeek-R1 product environment rather than an inherent capability of this downloadable model. Applications can wrap the model with tools, but the developer must implement and test that orchestration.

Deployment, cost, and speed trade-offs

There is no official DeepSeek per-token price for this exact checkpoint in the supplied research. Because it is open weight, users normally download and run it on their own hardware or access it through an independent inference provider. The total cost therefore depends on hardware ownership, rental rates, quantization, runtime efficiency, concurrency, and the provider selected for hosted inference.

Running the model locally can provide control over data, deployment configuration, and ongoing usage costs. It also shifts the operational burden to the user. The required memory depends on the bfloat16 or quantized format, runtime overhead, context length, and batch size; the supplied research does not specify a single hardware requirement. Quantized derivatives may reduce memory use, but they are community-produced files rather than the canonical DeepSeek checkpoint. Their accuracy, format, compatibility, and licensing should be checked independently.

Reasoning can make the model slower than a similarly sized general-purpose text model, especially on difficult prompts that produce long reasoning traces. In exchange, it may offer stronger performance on mathematics, code, and logic than its parameter count would otherwise suggest. The appropriate comparison is therefore not simply model size: users should weigh reasoning quality against response latency, available memory, throughput, and operational cost.

Main strengths and limitations

Strengths

  • Reasoning-focused training: Distillation from DeepSeek-R1-0528 gives the model a specific emphasis on multi-step problem solving.
  • Compact deployment target: At approximately 8 billion parameters, it is intended to be more approachable for local and smaller-scale deployments than much larger reasoning models.
  • Useful benchmark profile: DeepSeek's reported scores show particular strength in mathematical reasoning, general expert-level question answering, and coding evaluations.
  • Open licensing: The MIT license supports research, commercial use, and further model work within the applicable license terms.
  • Long configured context: The 131,072-token setting can support large prompts when the selected hardware and runtime can handle them.

Limitations

  • Text only: It cannot natively generate or process image, audio, or video modalities according to the supplied model data.
  • Variable local requirements: There is no single guaranteed hardware profile, and long contexts or long reasoning traces can increase memory use and latency.
  • No official checkpoint pricing: Users must calculate local infrastructure costs or evaluate a third-party inference service.
  • Not the full DeepSeek-R1-0528 model: Its smaller size makes it more deployable, but it should not be assumed to have the same broad capacity as the source model.
  • Unverified native tool calling: Tool orchestration, web search, and function execution should be supplied by the surrounding application unless separately tested.

When to choose DeepSeek-R1-0528-Qwen3-8B

This model is a strong candidate when the main requirement is local or self-managed reasoning at a relatively modest model scale. Suitable examples include solving and explaining mathematics problems, generating code snippets, reviewing algorithms, exploring research ideas, building experimental assistants, and deploying a text-only reasoning model where an external hosted API is undesirable.

It is particularly appealing when licensing flexibility and control over the checkpoint matter. A team can inspect the model files, select a compatible runtime, apply an appropriate quantization method, and decide whether to run on owned or rented hardware. The trade-off is that the team also becomes responsible for installation, performance tuning, scaling, monitoring, and application-level safety checks.

Another option may be more appropriate when low latency and high throughput are more important than extended reasoning, when the application requires native multimodal input or output, or when a managed service with documented tool calling and service-level commitments is required. A larger reasoning model may be preferable for tasks that exceed the smaller checkpoint's capacity, while a smaller non-reasoning model may be preferable for simple classification, short drafting, or high-volume responses where long reasoning traces add cost without much benefit.

Availability and final assessment

DeepSeek-R1-0528-Qwen3-8B is available as an open-weight checkpoint from DeepSeek's Hugging Face organization. Its combination of an 8-billion-parameter footprint, MIT licensing, Qwen3 architecture, and distilled reasoning training makes it a practical model to evaluate for local mathematics, coding, and logic workloads.

Its most important distinction is the balance it attempts to strike: stronger reasoning behavior than a conventional model of similar size, without requiring the full deployment profile of a much larger model. That balance comes with clear trade-offs. Responses may be slower or longer, the checkpoint does not provide native non-text modalities, and there is no official hosted price for this exact model. For users who can manage local inference and primarily need text-based reasoning, those trade-offs may be worthwhile; for managed, multimodal, or guaranteed tool-enabled applications, another type of model or service is likely a better fit.


Answers to Frequently Asked Questions

What is DeepSeek-R1-0528-Qwen3-8B?
DeepSeek-R1-0528-Qwen3-8B is an open-weight, approximately 8-billion-parameter language model released by DeepSeek on May 28, 2025. It is based on the Qwen3-8B architecture and was trained through distillation from DeepSeek-R1-0528 to improve multi-step reasoning for mathematics, coding, and logic tasks.
What license does DeepSeek-R1-0528-Qwen3-8B use?
The official DeepSeek-R1-0528-Qwen3-8B checkpoint is licensed under the MIT license. This generally permits local use, research, commercial use, and further distillation or fine-tuning, subject to the license terms.
What are the context length and main specifications of DeepSeek-R1-0528-Qwen3-8B?
The model has approximately 8 billion parameters, 36 transformer layers, a 4,096-dimensional hidden state, grouped-query attention, and bfloat16 weights. Its configured context length is 131,072 tokens, while the published benchmark methodology used a maximum generation setting of 64,000 tokens. Actual limits depend on the runtime and hardware.
Can DeepSeek-R1-0528-Qwen3-8B be run locally?
Yes. It is a downloadable open-weight checkpoint that can be deployed with compatible tools such as Transformers, vLLM, SGLang, Docker Model Runner, and quantization utilities. Memory use, speed, and supported context length depend on the model format, hardware, runtime, prompt length, and batch size.
What are the main limitations of DeepSeek-R1-0528-Qwen3-8B?
DeepSeek-R1-0528-Qwen3-8B is a text-only model and does not natively process or generate images, audio, or video. It may produce longer and slower responses because of its reasoning-focused behavior, has no official per-token price for this exact checkpoint, and does not have verified native tool calling or guaranteed function execution.


Sources 5
Provider

About DeepSeek