DeepSeek V4

DeepSeek-V4-Pro-Base

by DeepSeek · Current open-weight base checkpoint; downloadable under the MIT License

DeepSeek-V4-Pro-Base is the downloadable base checkpoint of DeepSeek-V4-Pro. It uses a Mixture-of-Experts architecture with 1.6 trillion total parameters, 49 billion activated per token, FP8 mixed precision, and a one-million-token context window. It is a foundation model rather than an instruction-tuned API chatbot, and its official repository provides the weights under the MIT License.

Text Reasoning Coding
DeepSeek-V4-Pro-Base is a downloadable foundation model for teams that want to build on a large language model rather than use a finished conversational assistant. The checkpoint combines 1.6 trillion total parameters with 49 billion active parameters per token, uses FP8 mixed precision, and provides a one-million-token context window. Its scale makes it suitable for serious research and large deployments, but its approximately 1.61 TB of model files and lack of hosted base-checkpoint pricing make it impractical for casual users or small local installations.
Outputs

What DeepSeek-V4-Pro-Base can produce

Text
Inputs

What it can understand

Text
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
4/10 Speed
3/10 Cost efficiency
Specifications

Technical details

Model family DeepSeek V4
Model type General Purpose
Context window 1.05M tokens
Release date 2026-04-24
Status Current open-weight base checkpoint; downloadable under the MIT License
Knowledge cutoff notes

No authoritative model-specific knowledge-cutoff date was found in the official model card, technical documentation, or repository configuration.

Model notes

This is the base, non-instruction-tuned checkpoint rather than the separately distributed DeepSeek-V4-Pro instruction-tuned model. The official DeepSeek V4 repository lists 1.6T total parameters, 49B activated parameters, a 1M-token context window, and FP8 mixed precision for the base checkpoint. The Hugging Face repository contains approximately 1.61 TB of model files and indicates that the checkpoint is not deployed by an inference provider. The official API pricing page applies to the hosted deepseek-v4-pro model version, not to DeepSeek-V4-Pro-Base, so no hosted input or output price is assigned here. Editorial capability scores reflect the model's published base-model benchmark results and its very large deployment requirements; they are not vendor-provided ratings. The checkpoint uses the DeepSeek V4 causal-language-model architecture and has no native image, audio, or video input or output documented.

Model guide

DeepSeek-V4-Pro-Base: A 1.6T Open-Weight Foundation Model for Custom Training

DeepSeek-V4-Pro-Base is the base, non-instruction-tuned open-weight checkpoint of DeepSeek's 1.6-trillion-parameter Mixture-of-Experts V4 Pro model. It activates 49 billion parameters per token, supports a one-million-token context window, and is designed for continued pretraining, fine-tuning, evaluation, and custom inference rather than turnkey chat.

DeepSeek-V4-Pro-Base is DeepSeek's downloadable base checkpoint for the V4 Pro model family. It is a large causal language model with a Mixture-of-Experts (MoE) architecture: the model contains 1.6 trillion total parameters, while approximately 49 billion are activated for each token. In practical terms, this allows the model to have very large overall capacity without using every parameter for every word or code token.

The important distinction is that this is a base model, not an instruction-tuned chat model. It is intended to be adapted, evaluated, or deployed by researchers and engineering teams that need control over training and inference. Someone looking for a ready-to-use conversational endpoint should not treat the Base checkpoint as a drop-in replacement for a hosted instruction-following model.

What DeepSeek-V4-Pro-Base is

DeepSeek-V4-Pro-Base is provided by DeepSeek as an open-weight model under the MIT License, according to the supplied official model materials. “Open-weight” means that the model files are available for download and use under the stated license; it does not mean that the model is small, inexpensive to run, or automatically configured for a particular application.

The checkpoint uses the DeepSeek V4 causal-language-model architecture. A causal language model generates text by predicting the next token from the preceding context. Because this checkpoint is not instruction-tuned, it should be viewed as a foundation for further work rather than as a polished assistant with guaranteed conversational behavior, built-in safety behavior, or a provider-managed application layer.

The official release materials identify DeepSeek-V4-Pro-Base as part of the DeepSeek V4 release announced on April 24, 2026. The model repository is hosted on Hugging Face under DeepSeek's official account and contains approximately 1.61 TB of model files. The repository indicates that the checkpoint is not deployed by an inference provider.

Verified specifications and context capacity

SpecificationDeepSeek-V4-Pro-Base
Model familyDeepSeek V4
Model typeGeneral-purpose causal language model
ArchitectureMixture of Experts
Total parameters1.6 trillion
Activated parameters per token49 billion
Context window1,048,576 tokens, or one million tokens
PrecisionFP8 mixed precision
LicenseMIT License, according to the official model materials
Download sizeApproximately 1.61 TB in the Hugging Face repository

The one-million-token context window is one of the checkpoint's most significant specifications. It indicates the maximum context length documented for the model configuration, allowing a compatible deployment to process extremely large collections of text, source code, or other tokenized inputs in a single context. Actual usable capacity can still depend on the inference software, memory configuration, batching strategy, and deployment hardware.

The supplied research does not identify a separate maximum output-token limit for this checkpoint. The one-million-token context value should therefore not be interpreted as a guaranteed one-million-token response limit. It describes the model's configured context capacity, which includes the input and any generated output handled by a deployment.

Modalities and supported outputs

DeepSeek-V4-Pro-Base is documented as a text model. It accepts text input and produces text output. There is no documented native image, audio, or video input, and no documented image, audio, video, music, speech, embedding, or other non-text output for this checkpoint.

This limitation matters when comparing the Base checkpoint with multimodal services in the broader DeepSeek ecosystem. The consumer DeepSeek product and some other model-family services may offer visual understanding or file-oriented features, but those capabilities should not be attributed to DeepSeek-V4-Pro-Base without model-specific documentation. This checkpoint is a text-generation foundation model.

Reasoning and coding capabilities

DeepSeek-V4-Pro-Base is positioned for general language-model research, including reasoning and code-related evaluation. Its large parameter count, long context window, and foundation-model status make it relevant to experiments involving long documents, large codebases, continued pretraining, and domain adaptation.

The supplied editorial dataset assigns a reasoning score of 8 out of 10 and a coding score of 8 out of 10. These are editorial evaluations, not scores published by DeepSeek and not a substitute for a benchmark result. They indicate the checkpoint's expected research relevance in reasoning and programming workloads, while recognizing that a base model may require prompting, fine-tuning, or additional post-training before it behaves like a production coding assistant.

There is no verified native tool-use or function-calling interface listed for this checkpoint. Tool orchestration would therefore need to be implemented by the surrounding application or supported by a separately adapted model and serving stack. Likewise, structured-output and JSON-mode support are not documented as native capabilities in the supplied research.

Pricing and access

There is no verified hosted input or output price for DeepSeek-V4-Pro-Base. The official DeepSeek pricing documentation applies to the hosted deepseek-v4-pro model version, which is distinct from this downloadable base checkpoint. Those hosted prices should not be transferred to the Base model.

Instead, users access the checkpoint by downloading and operating the model files through a compatible inference environment. The approximately 1.61 TB repository size is a major practical cost factor before accounting for accelerators, memory, storage, networking, electricity, engineering time, and maintenance. The model's FP8 mixed-precision configuration may help reduce memory and computation requirements compared with an equivalent full-precision deployment, but the supplied research does not provide a minimum hardware configuration or a guaranteed operating cost.

This makes the model's economic profile different from a small hosted API model. There is no per-token price assigned to the Base checkpoint, but self-hosting at this scale can require substantial capital and operational resources. A team should estimate total infrastructure cost rather than assuming that an open MIT-licensed checkpoint is inexpensive to run.

Main strengths

  • Very large model capacity: The 1.6-trillion-parameter MoE design provides substantial overall model capacity while activating 49 billion parameters per token.
  • Long-context research potential: The documented one-million-token context window is useful for experiments involving large documents, repositories, and extended sequences.
  • Adaptability: As a base checkpoint, it can serve as a starting point for continued pretraining, supervised fine-tuning, evaluation, or specialized deployment.
  • Open-weight availability: The official materials identify the model as available under the MIT License, giving developers more control than a closed hosted endpoint, subject to the license and applicable obligations.
  • Text and code focus: The checkpoint is suited to teams evaluating large-scale language generation, reasoning, and programming behavior without requiring native image or audio processing.

Main limitations and trade-offs

  • Not a turnkey assistant: The checkpoint is non-instruction-tuned, so it is not the most convenient choice for direct chat or an immediately polished user-facing application.
  • Large deployment footprint: Approximately 1.61 TB of model files creates demanding storage, transfer, memory, and serving requirements.
  • No documented hosted Base pricing: Users cannot use the hosted deepseek-v4-pro price as a verified price for this checkpoint.
  • No documented native multimodality: Image, audio, and video inputs and outputs are not supported in the supplied model-specific documentation.
  • No verified native tools: Tool use, function calling, structured output, streaming, fine-tuning service access, caching, and batch API support are not documented as built-in features of this repository.
  • Unspecified output limit: The research confirms the context window but does not provide a separate maximum output-token value.
  • Speed and cost concerns: The editorial speed score is 4 out of 10 and cost score is 3 out of 10. These are subjective editorial assessments reflecting the model's scale and deployment demands, not provider-published ratings.

When to choose DeepSeek-V4-Pro-Base

Choose DeepSeek-V4-Pro-Base when your project needs a large downloadable foundation model and your team can operate the required infrastructure. It is a good candidate for continued pretraining on a specialized corpus, fine-tuning experiments, long-context evaluation, research into MoE models, and custom inference where control over the weights and serving stack is more important than immediate convenience.

It can also make sense when you need to study or adapt a very large text model without depending entirely on a closed API. The MIT License and open-weight distribution may support workflows that require local control, reproducibility, or custom model modification, although legal and operational review remains the user's responsibility.

A hosted instruction-tuned model is likely more appropriate for a customer-facing chatbot, a quick prototype, or an application that needs published token pricing and managed infrastructure. A smaller model may be preferable when response speed, low hardware cost, or easy local deployment matters more than maximum model capacity. A multimodal model should be selected for image, audio, or video workloads, because those capabilities are not documented for DeepSeek-V4-Pro-Base.

Bottom line

DeepSeek-V4-Pro-Base is best understood as a large research and engineering asset, not as a ready-made chat product. Its defining characteristics are the 1.6-trillion-parameter MoE architecture, 49 billion active parameters per token, one-million-token context window, FP8 mixed precision, and downloadable open-weight distribution.

Those specifications make it interesting for teams building custom language-model systems, but they also define its limits. There is no verified hosted Base-model price, no documented native multimodal or tool interface, and no supplied maximum output-token limit. For organizations with the infrastructure and expertise to adapt a foundation model, it offers considerable flexibility. For users who primarily want fast, inexpensive, managed conversation, a smaller or instruction-tuned hosted alternative is likely the more practical choice.


Answers to Frequently Asked Questions

Does DeepSeek-V4-Pro-Base support multimodal inputs, tool use, or hosted API pricing?
The checkpoint is documented as a text-only model and has no verified native image, audio, or video capabilities. Native tool use, function calling, and structured-output modes are also not documented. There is no verified hosted input or output price for the Base checkpoint; users generally need to download and operate it in their own compatible inference environment.
How large is the DeepSeek-V4-Pro-Base download and what hardware does it require?
The Hugging Face repository contains approximately 1.61 TB of model files. The model uses FP8 mixed precision, which may reduce memory and compute requirements compared with full precision, but the supplied documentation does not specify a minimum hardware configuration or guaranteed operating cost.
What is the context window of DeepSeek-V4-Pro-Base?
The documented context window is 1,048,576 tokens, or approximately one million tokens. Actual usable capacity depends on the inference software, hardware memory, batching strategy, and deployment configuration. This value does not guarantee a one-million-token output limit.
What is DeepSeek-V4-Pro-Base?
DeepSeek-V4-Pro-Base is a downloadable, open-weight foundation model from the DeepSeek V4 family. It is a causal language model with a Mixture-of-Experts architecture, 1.6 trillion total parameters, and approximately 49 billion activated parameters per token.
Is DeepSeek-V4-Pro-Base an instruction-tuned chat model?
No. DeepSeek-V4-Pro-Base is a base model rather than an instruction-tuned assistant. It is intended for continued pretraining, fine-tuning, evaluation, and custom deployment, so it may require additional adaptation before being used as a polished conversational chatbot.


Sources 6
Provider

About DeepSeek