DeepSeek-R1

DeepSeek-R1-Zero

by DeepSeek · Open-weight and downloadable; experimental/legacy research model; not a current first-party hosted API model

DeepSeek-R1-Zero is a 671B mixture-of-experts reasoning model trained directly with large-scale reinforcement learning on a DeepSeek-V3 base model. It demonstrated emergent self-verification and long-chain reasoning, but also showed repetition, readability, and language-mixing problems that motivated the later DeepSeek-R1 training approach.

Text Reasoning Coding
DeepSeek-R1-Zero is DeepSeek's experimental first-generation reasoning model and the foundation for the later DeepSeek-R1 release. It is notable for showing that large-scale reinforcement learning applied directly to a base model can produce self-verification, reflection, and extended chain-of-thought behavior without a supervised fine-tuning stage.
Outputs

What DeepSeek-R1-Zero can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Fine-tuning
Model profile

Performance characteristics

9/10 Reasoning
8/10 Coding
3/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family DeepSeek-R1
Model type Reasoning
Context window 131K tokens
Maximum output 33K tokens
Knowledge cutoff July 2024
Release date 2025-01-20
Status Open-weight and downloadable; experimental/legacy research model; not a current first-party hosted API model
Knowledge cutoff notes

July 2024 is reported by independent model metadata for the DeepSeek-R1 model line. DeepSeek's primary R1-Zero model card and release materials do not explicitly publish a separate knowledge-cutoff date for the exact R1-Zero checkpoint, so this value should be treated as an informed but not directly first-party-verified estimate.

Model notes

DeepSeek-R1-Zero is the experimental predecessor to DeepSeek-R1. It applies large-scale reinforcement learning directly to a DeepSeek-V3-Base-derived model without supervised fine-tuning as a preliminary stage. DeepSeek reports emergent self-verification, reflection, and long chain-of-thought behavior, but also identifies endless repetition, poor readability, and language mixing as limitations. The model has 671B total parameters and approximately 37B activated parameters per token. The published R1 model table gives a 128K context length, and DeepSeek's evaluation configuration uses a maximum generation length of 32,768 tokens. The full checkpoint is distributed under the MIT License and is extremely demanding to deploy. The official DeepSeek API historically exposed DeepSeek-R1 through the deepseek-reasoner identifier, not DeepSeek-R1-Zero; the legacy deepseek-reasoner API name was scheduled for discontinuation on July 24, 2026. Pricing is therefore not applicable to the exact downloadable R1-Zero checkpoint. Streaming is available when the checkpoint is served through compatible inference servers such as vLLM or SGLang, rather than being an intrinsic model output feature. Fine-tuning is technically possible because the weights are openly licensed, but full-model fine-tuning requires substantial infrastructure.

Model guide

DeepSeek-R1-Zero: The Open-Weight Reinforcement-Learning Reasoning Model

DeepSeek-R1-Zero is a 671-billion-parameter mixture-of-experts reasoning model from DeepSeek that demonstrated emergent reasoning behaviors through large-scale reinforcement learning applied directly to a base model without supervised fine-tuning.

What is DeepSeek-R1-Zero?

DeepSeek-R1-Zero is an open-weight reasoning model released by DeepSeek on January 20, 2025. It is designed primarily for mathematical reasoning, coding, research, and experimentation with models that generate extended reasoning traces before producing an answer.

The model is the experimental predecessor to DeepSeek-R1. Its importance is not simply its size, but its training approach: DeepSeek applied large-scale reinforcement learning directly to a base model instead of first performing supervised fine-tuning with human-written reasoning examples. The result demonstrated that some reasoning behaviors could emerge from reinforcement learning at scale.

DeepSeek-R1-Zero is not the same thing as the current DeepSeek consumer assistant or a generally available hosted API model. It is a downloadable checkpoint intended for model research and self-managed deployment. The model remains useful for understanding the R1 family, but its experimental behavior makes it less suitable than a polished instruction-tuned model for ordinary chat applications.

Architecture and training approach

DeepSeek-R1-Zero is based on a DeepSeek-V3-Base-derived model and uses a mixture-of-experts architecture. It contains 671 billion total parameters, while approximately 37 billion parameters are activated for each token. In a mixture-of-experts model, different portions of the network handle different tokens, so the total parameter count is much larger than the number used for any individual token. This reduces per-token computation compared with a dense model of the same total size, although the complete checkpoint still has very substantial memory and infrastructure requirements.

The defining feature is direct reinforcement learning from the base model. Rather than using supervised fine-tuning as a preliminary step, the training process rewarded behavior associated with solving reasoning tasks. DeepSeek reports that this produced self-verification, reflection, and long chains of thought. These behaviors are particularly relevant to tasks where the model benefits from checking intermediate steps, such as mathematics, algorithms, and code reasoning.

DeepSeek-R1-Zero also illustrates the limits of this approach. The model can generate very long or repetitive reasoning traces, mix languages, and produce responses that are difficult to read. DeepSeek therefore used a more structured cold-start and multi-stage training process for DeepSeek-R1, which was intended to provide more consistent and user-friendly behavior.

Capabilities and supported inputs

DeepSeek-R1-Zero is a text-only model. Its supported input and output modality is text; it does not provide verified image, audio, or video input or output. It is therefore not suitable for visual document analysis, image understanding, speech workflows, or media generation.

Its main capability is extended text reasoning. The model is especially relevant to:

  • Mathematical and symbolic problem solving
  • Algorithm design and code reasoning
  • Evaluation of open-weight reasoning systems
  • Research into reinforcement-learning-based language models
  • Model distillation and derivative-model development

The model's reasoning behavior should not be confused with guaranteed correctness. A long chain of thought can contain errors, and the model's reported self-verification behavior does not eliminate the need to check important answers. For production systems, generated code and mathematical results should be tested independently.

Context window and output limits

The published R1 model information gives DeepSeek-R1-Zero a context length of 131,072 tokens, commonly described as a 128K-token context window. This is the maximum amount of input and conversation material that a compatible deployment can process according to the supplied model specifications. Actual usable capacity can depend on the serving system, memory limits, prompt structure, and generation settings.

DeepSeek's evaluation configuration for the R1 series uses a maximum generation length of 32,768 tokens. That limit is important for this model because its reasoning traces can be substantially longer than those of a conventional instruction-following model. A large output allowance can help with difficult reasoning tasks, but it can also increase latency, memory use, and cost when the model is deployed on rented or shared infrastructure.

SpecificationDeepSeek-R1-Zero
ProviderDeepSeek
Release dateJanuary 20, 2025
ArchitectureMixture of experts
Total parameters671 billion
Activated parameters per tokenApproximately 37 billion
Context length131,072 tokens, approximately 128K
Maximum evaluated generation length32,768 tokens
Input and outputText only
LicenseMIT

Reasoning, coding, and tool support

Reasoning is the model's central purpose. Compared with a conventional chat model, DeepSeek-R1-Zero is intended to spend more computation producing intermediate reasoning before presenting an answer. That makes it a research subject for tasks that benefit from decomposition, checking, and persistence rather than immediate short responses.

It is also appropriate for coding experiments. The model can be used to explore algorithmic solutions, generate code, explain implementation decisions, and test reasoning-oriented programming workflows. However, the supplied specifications do not identify built-in tool calling, function calling, web search, code execution, or an agent action interface for the checkpoint. A deployment can potentially be integrated with external tools, but that would be a property of the surrounding application or serving framework, not a verified intrinsic capability of DeepSeek-R1-Zero.

The model does not have a verified native structured-output or JSON-mode feature in the supplied research. Applications that need reliable machine-readable output should validate and, if necessary, repair the model's responses rather than assuming strict schema compliance.

Main strengths and limitations

Strengths

  • Research significance: It provides a concrete example of reasoning behavior emerging from reinforcement learning applied directly to a base model.
  • Open availability: The weights are downloadable, and the model is distributed under the MIT License.
  • Large reasoning capacity: Its long context and high generation allowance support extended problem-solving experiments.
  • Strong fit for technical research: Mathematics, coding, model evaluation, distillation, and reinforcement-learning studies are natural use cases.
  • Deployment flexibility: Researchers can inspect, serve, adapt, or evaluate the checkpoint using compatible open-model infrastructure rather than relying on a first-party hosted endpoint.

Limitations

  • Experimental response quality: DeepSeek reports endless repetition, poor readability, and language mixing in R1-Zero outputs.
  • High deployment demands: A 671-billion-parameter checkpoint requires distributed inference infrastructure and substantial memory, even though only about 37 billion parameters are activated per token.
  • No first-party R1-Zero API: DeepSeek's historical hosted reasoning API used the deepseek-reasoner identifier for DeepSeek-R1, not for this exact checkpoint.
  • No native multimodal support: It is not a vision, audio, or video model.
  • Potentially high latency: Long reasoning traces and large outputs can make responses slower than smaller or more instruction-tuned models.
  • Output verification remains necessary: Reinforcement-learned reasoning does not guarantee factual, mathematical, or programming correctness.

Availability and pricing

DeepSeek-R1-Zero is available as open weights through the DeepSeek-R1 repository and the deepseek-ai/DeepSeek-R1-Zero repository on Hugging Face. It can be deployed with compatible open-model serving systems such as vLLM or SGLang, subject to the infrastructure and configuration requirements of the specific deployment.

There is no verified provider price for the exact downloadable R1-Zero checkpoint. It is not presented in the supplied research as a current first-party hosted API product with its own per-token tariff. The cost of using it therefore depends on hardware ownership, cloud GPU rental, hosting, engineering, and inference efficiency. Open weights remove a model-access fee, but they do not make large-scale inference free.

DeepSeek's historical API documentation distinguishes this checkpoint from DeepSeek-R1 served through deepseek-reasoner. The legacy API identifier was scheduled for discontinuation on July 24, 2026. That API history should not be interpreted as evidence that DeepSeek-R1-Zero itself has a current hosted API endpoint.

Deployment considerations

Serving the full model is a substantial systems project. Operators need distributed inference infrastructure, enough memory for the checkpoint and runtime overhead, and serving software that supports the model architecture. The research identifies vLLM and SGLang as examples of compatible serving paths, but exact hardware requirements depend on quantization, parallelism, batch size, context length, and deployment settings.

Streaming is possible when the checkpoint is served through a compatible inference server, but streaming is not an intrinsic model capability in the same way as its learned text-generation behavior. Similarly, fine-tuning is technically possible because the weights are openly licensed, but adapting a model of this size requires considerable compute and engineering resources. Smaller distilled or successor models may be more practical when the goal is experimentation on limited hardware.

When to choose DeepSeek-R1-Zero

Choose DeepSeek-R1-Zero when the priority is studying open-weight reasoning rather than obtaining the smoothest conversational experience. It is a strong candidate for researchers investigating how reinforcement learning affects reasoning, teams evaluating long-form mathematical or coding behavior, and developers experimenting with distillation or custom serving.

Its open-weight status is also useful when the ability to inspect and operate a model independently matters more than turnkey access. The MIT License provides broad reuse rights, although deployment, modification, and redistribution should still be reviewed in the context of the applicable project and dependencies.

Another option is more appropriate when you need low-latency production chat, reliable instruction following, native multimodal input, built-in tools, predictable structured output, or a managed API. DeepSeek-R1 is the relevant sibling for understanding the later, more polished training direction, while a smaller model or hosted service may offer a better speed-and-cost trade-off for routine workloads. DeepSeek-R1-Zero is best viewed as a research checkpoint and open-model foundation, not as a drop-in replacement for a mature assistant.

Bottom line

DeepSeek-R1-Zero is important because it demonstrated a distinctive training result: a very large base model developed recognizable reasoning behaviors through reinforcement learning without supervised fine-tuning as the initial stage. Its 671-billion-parameter mixture-of-experts architecture, long context, and open MIT-licensed weights make it valuable for research and large-scale experimentation.

Those same characteristics limit its practicality. The full model is expensive to deploy, responses may be repetitive or difficult to read, and there is no verified current first-party API for the exact checkpoint. For researchers studying reasoning or open-weight model development, it is a significant model. For everyday users and production applications, its successor direction or a smaller, hosted, instruction-tuned alternative is likely to be more appropriate.


Answers to Frequently Asked Questions

Is DeepSeek-R1-Zero available through a hosted API?
There is no verified current first-party hosted API for the exact DeepSeek-R1-Zero checkpoint. The historical deepseek-reasoner API identifier referred to DeepSeek-R1 rather than R1-Zero, so using R1-Zero generally requires self-managed or third-party deployment.
How can DeepSeek-R1-Zero be deployed?
The model is available as open weights through the DeepSeek-R1 repository and the deepseek-ai/DeepSeek-R1-Zero repository on Hugging Face. It can be served with compatible systems such as vLLM or SGLang, but deploying the full checkpoint requires distributed inference infrastructure and substantial memory.
How many parameters does DeepSeek-R1-Zero have?
DeepSeek-R1-Zero uses a mixture-of-experts architecture with 671 billion total parameters. Approximately 37 billion parameters are activated for each token, which reduces per-token computation compared with a dense model of the same total size.
What are the main limitations of DeepSeek-R1-Zero?
DeepSeek-R1-Zero can produce very long or repetitive reasoning traces, mix languages, and generate responses that are difficult to read. It also requires substantial deployment infrastructure, has no verified native multimodal or tool-calling support, and does not guarantee correct mathematical, factual, or programming results.
What is DeepSeek-R1-Zero?
DeepSeek-R1-Zero is an open-weight reasoning model released by DeepSeek on January 20, 2025. It is designed for mathematical reasoning, coding, research, and experiments involving extended reasoning traces generated through reinforcement learning.


Sources 6
Provider

About DeepSeek