DeepSeek-R1

DeepSeek-R1

by DeepSeek · Open-weight checkpoint available; original hosted API identity superseded and scheduled for discontinuation

DeepSeek-R1 is DeepSeek's original January 2025 open-weight reasoning model. Its 671-billion-parameter mixture-of-experts architecture, approximately 37-billion active parameters, 128K context window, and MIT license make it relevant to mathematical reasoning, coding, research, and self-hosted deployments. The model is text-only, computationally demanding, and no longer represented by a current first-party price for the exact checkpoint.

Text Reasoning Coding
DeepSeek-R1 is DeepSeek's original large-scale open reasoning model. Rather than optimizing mainly for short, fast answers, it is intended to spend more computation working through difficult problems before returning a response. That makes it especially relevant to mathematics, coding, algorithm design, debugging, and other tasks where a carefully developed solution matters more than minimum latency. The model was released as downloadable weights under the MIT license, alongside DeepSeek-R1-Zero and smaller distilled models. The original checkpoint remains technically significant, but it should be distinguished from later hosted revisions such as DeepSeek-R1-0528 and from newer DeepSeek model families. Current access to the exact January 2025 checkpoint is more likely to involve self-hosting or third-party inference than DeepSeek's current default API catalog.
Outputs

What DeepSeek-R1 can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming
Model profile

Performance characteristics

9/10 Reasoning
8/10 Coding
4/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family DeepSeek-R1
Model type Reasoning
Context window 128K tokens
Maximum output 33K tokens
Release date 2025-01-20
Status Open-weight checkpoint available; original hosted API identity superseded and scheduled for discontinuation
Deprecation date 2026-04-24
Shutdown date 2026-07-24
Knowledge cutoff notes

DeepSeek's official model card and repository document the model architecture, context length, training approach, and release information but do not state an authoritative knowledge-cutoff date for the exact original DeepSeek-R1 checkpoint.

Model notes

DeepSeek-R1 is the original January 2025 671B-total-parameter, approximately 37B-active mixture-of-experts checkpoint based on DeepSeek-V3-Base. The official model card lists a 128K context length and uses 32,768 tokens as the maximum generation length in its evaluation configuration. The model and weights are MIT licensed. DeepSeek originally exposed the model through the deepseek-reasoner API identity with historical pricing of $0.14 per million cached input tokens, $0.55 per million uncached input tokens, and $2.19 per million output tokens. That API identity was later upgraded to newer model revisions and is scheduled for discontinuation as part of DeepSeek's transition to the V4 family, so those historical prices do not represent current access to the exact original checkpoint. The original model is text-only. The reasoning, coding, speed, and cost values are editorial comparative scores rather than vendor-provided ratings. No authoritative first-party knowledge-cutoff date was found for the exact original checkpoint.

Model guide

DeepSeek-R1: Open-Weight Reasoning for Mathematics, Coding, and Research

DeepSeek-R1 is a 671-billion-parameter open-weight mixture-of-experts reasoning model released by DeepSeek on January 20, 2025. It is designed for difficult mathematics, programming, logical reasoning, and technical analysis, with approximately 37 billion parameters activated per inference step. The released checkpoint has a 128K-token context window, a 32,768-token maximum generation configuration, text-only input and output, and an MIT license. Its original hosted API identity has been superseded by newer DeepSeek revisions, so the exact checkpoint is now primarily relevant for self-hosting, research, and third-party inference.

What is DeepSeek-R1?

DeepSeek-R1 is an open-weight reasoning model provided by DeepSeek and released on January 20, 2025. Open-weight means that the model parameters can be downloaded and run or adapted under the applicable license terms, rather than being available only through a provider-operated interface.

The released checkpoint is a mixture-of-experts, or MoE, model with 671 billion total parameters. Approximately 37 billion parameters are activated for each inference step, so the total model size and the computation used for an individual token are different measures. In practical terms, the full checkpoint is still extremely large to operate, even though each step uses only a portion of the available parameters.

DeepSeek based R1 on DeepSeek-V3-Base and trained it with a process combining supervised data, reinforcement learning, and preference-alignment stages. The training approach was intended to produce longer, more deliberate reasoning traces for difficult problems.

Reasoning purpose and training approach

DeepSeek-R1 is aimed at tasks that benefit from multi-step problem solving. It can work through mathematical derivations, programming problems, logical questions, and technical analysis before presenting a final answer. This style can make the path toward a conclusion more inspectable than a short answer from a conventional language model, although a visible reasoning trace is not a guarantee that every intermediate step is correct.

DeepSeek released R1 after DeepSeek-R1-Zero, an earlier experiment that relied more heavily on reinforcement learning. According to DeepSeek's release materials, R1 added cold-start data and additional alignment work to address issues observed in R1-Zero, including repetition, readability problems, and language mixing. The result was positioned as a more usable general reasoning model.

Reasoning has a practical cost. Longer internal or visible solution paths can increase latency and token consumption. A simple classification, short rewrite, or routine chat response may therefore be better served by a smaller or faster non-reasoning model, while a difficult proof, debugging task, or algorithmic question can justify the additional computation.

Verified technical specifications

SpecificationDeepSeek-R1
ProviderDeepSeek
Release dateJanuary 20, 2025
ArchitectureMixture of experts
Total parameters671 billion
Activated parametersApproximately 37 billion per inference step
Context length128,000 tokens in the released checkpoint
Maximum generation32,768 tokens in the released model evaluation configuration
InputText
OutputText
LicenseMIT
Knowledge cutoffNot officially published for the exact original checkpoint

The 128K context window describes how much text the released checkpoint can handle in a request, subject to the limits of the serving software and available hardware. The 32,768-token figure is the maximum generation length documented in the model's evaluation configuration; a particular deployment may impose a lower limit.

Mathematics, coding, and complex tasks

DeepSeek-R1's main purpose is deliberate problem solving. It is a suitable candidate for mathematical exercises, symbolic or logical analysis, algorithm development, code generation, debugging, and technical research. It can also help examine a long specification or reason across a large body of text when the deployment preserves enough of its context window.

For coding, the model is most useful when the task requires explanation as well as code: designing an algorithm, finding a likely cause of a bug, comparing implementation strategies, or translating requirements into a structured solution. It should still be tested with a compiler, test suite, or human review. DeepSeek-R1 is a language model, not an independently verified software execution environment.

The available research supports text reasoning and text generation, not native image, audio, or video processing. The original checkpoint therefore cannot directly inspect a screenshot, listen to a recording, or analyze a video unless another system first converts that material into text.

Availability and current API position

DeepSeek-R1 weights are available through DeepSeek's official repository and Hugging Face model page. The MIT license supports commercial use, modification, and derivative works, subject to the terms of that license and any obligations associated with a particular deployment.

The original hosted API identity was deepseek-reasoner. DeepSeek subsequently upgraded that identity to newer R1 and V3.1 revisions. Its API change log also announces the discontinuation of the legacy identity as part of the transition to the V4 model family, with a listed deprecation date of April 24, 2026 and shutdown date of July 24, 2026. Those dates describe the hosted API lifecycle, not the disappearance of the downloadable open-weight checkpoint.

As a result, the exact original January 2025 model should not be treated as a currently listed default model in DeepSeek's hosted catalog. Developers who need this precise checkpoint should evaluate self-hosting, a compatible inference framework, or a third-party service that explicitly identifies the original DeepSeek-R1 weights. A provider's newer R1 revision may be more convenient, but it is not automatically the same model.

Pricing and cost trade-offs

There is no verified current first-party hosted price for the exact original checkpoint in DeepSeek's current pricing catalog. Historical pricing for the original deepseek-reasoner API identity was reported as $0.14 per million cached input tokens, $0.55 per million uncached input tokens, and $2.19 per million output tokens. These figures are historical and should not be used as current pricing for the original checkpoint.

For self-hosting, the main cost is infrastructure rather than a per-token subscription. The full 671-billion-parameter checkpoint requires substantial memory, compute capacity, deployment engineering, and ongoing operations. Smaller distilled R1 variants can make the reasoning approach more accessible, but they are separate model checkpoints and should not be represented as the full DeepSeek-R1 model.

There is also a usage trade-off within each request. Long reasoning outputs consume more tokens and may take longer than short responses. DeepSeek-R1 can therefore offer attractive capability per token in difficult tasks while being a poor fit for high-volume, latency-sensitive workloads where a smaller model is sufficient.

Tools, modalities, and output control

The original DeepSeek-R1 record identifies text as both its input and output modality. It does not natively support image, audio, or video input, and it does not produce images, audio, video, or speech. It also has no verified built-in web search or current-information access. Applications that need current information must provide retrieval or another external data source.

Tool or function calling is not identified as a supported capability for this original checkpoint. Similarly, structured output is not listed as a native feature. Developers can build wrappers that ask for a particular format, but prompt instructions are not equivalent to a provider-guaranteed JSON or schema mode.

The model record marks streaming as supported in the historical API context. For self-hosted use, actual streaming behavior depends on the inference server and serving framework selected. The same distinction applies to batching, caching, quantization, and hardware acceleration: these are deployment properties rather than universal properties of the model weights.

Main strengths and limitations

Strengths

  • Strong fit for extended mathematical, logical, programming, and technical reasoning.
  • Open weights allow research, commercial deployment, modification, and experimentation under the MIT license.
  • A 128K context window supports long technical prompts and documents in the released checkpoint.
  • Approximately 37 billion active parameters per step provide an MoE efficiency advantage relative to the total parameter count, although the full model remains demanding to serve.
  • Smaller distilled variants make related R1-style reasoning available to deployments that cannot operate the full checkpoint.

Limitations

  • The full checkpoint is very large and is not a practical small-device model.
  • Extended reasoning can increase response time, token use, and operating cost.
  • The original model is text-only and lacks native image, audio, and video understanding.
  • There is no verified built-in web search, so answers requiring current information need external retrieval.
  • Native tool calling and guaranteed structured-output modes are not documented for the original checkpoint.
  • The exact model's official knowledge-cutoff date has not been published.
  • The original checkpoint is easy to confuse with later hosted R1 revisions and newer DeepSeek model families.

When to choose DeepSeek-R1

Choose DeepSeek-R1 when you specifically want an open-weight reasoning model for mathematics, coding, technical investigation, or research and have the infrastructure to serve a very large checkpoint. It is also a reasonable choice for teams that value the MIT license and the ability to inspect, modify, or independently deploy the model rather than relying entirely on a closed hosted service.

Consider a smaller distilled reasoning model when the same general problem-solving style is needed with lower hardware requirements and latency. Consider a newer hosted model when managed API access, current pricing, built-in operational features, or a supported lifecycle matter more than using the exact original checkpoint. A multimodal model is more appropriate for images, audio, or video, while a retrieval-enabled system is preferable when answers must reflect current web information.

DeepSeek-R1 is therefore best understood as a substantial open reasoning checkpoint rather than a universal assistant. Its central advantage is the combination of difficult-task reasoning, downloadable weights, and a permissive license. Its central cost is the infrastructure and latency required to make that capability practical.


Answers to Frequently Asked Questions

What are the main limitations of DeepSeek-R1?
The full 671-billion-parameter checkpoint requires substantial infrastructure and can produce higher latency and token costs because of its extended reasoning. The original model is text-only, has no verified built-in web search or native multimodal input, and does not document guaranteed tool calling or structured-output modes.
Is DeepSeek-R1 available through an API, and what happened to deepseek-reasoner?
The original DeepSeek-R1 weights are available through DeepSeek's official repository and Hugging Face. The original hosted API identity was deepseek-reasoner, but DeepSeek later moved to newer R1 and V3.1 revisions and announced the eventual discontinuation of the legacy identity. Developers who need the exact January 2025 checkpoint should consider self-hosting or a service that explicitly identifies those weights.
Can DeepSeek-R1 be used for coding and mathematical problem solving?
Yes. DeepSeek-R1 is intended for mathematical derivations, programming, debugging, algorithm design, logical analysis, and other complex technical tasks. Generated code should still be checked with a compiler, test suite, or human review because the model does not independently verify software execution.
What is DeepSeek-R1?
DeepSeek-R1 is an open-weight reasoning model released by DeepSeek on January 20, 2025. It is designed for multi-step mathematics, coding, logical analysis, and technical research, and its downloadable weights can be run or adapted under the MIT license.
What are the main technical specifications of DeepSeek-R1?
DeepSeek-R1 uses a mixture-of-experts architecture with 671 billion total parameters and approximately 37 billion activated parameters per inference step. The released checkpoint supports a 128,000-token context window and up to 32,768 generated tokens in the documented evaluation configuration.


Sources 5
Provider

About DeepSeek