DeepSeek V4

DeepSeek-V4-Pro

by DeepSeek · Deprecated for independent serving; deepseek-v4-pro API requests are routed to DeepSeek-V4.1-Flash until V4.1-Pro launches

DeepSeek-V4-Pro is a text-based MoE model designed for complex reasoning, coding, long-context analysis, and agent workflows. It has a 1-million-token context, configurable thinking effort, a 384,000-token maximum output, tool calling, JSON output, streaming, and prompt caching. Its main qualification is that DeepSeek announced routing deepseek-v4-pro API requests to V4.1-Flash until V4.1-Pro launches.

Text Reasoning Coding
DeepSeek-V4-Pro was designed for demanding language-model workloads rather than quick, low-cost chat alone. Its unusually large 1-million-token context window makes it suitable for large codebases, long documents, and extended agent tasks, while configurable thinking modes are intended to provide more deliberate reasoning when a problem requires it. The model also supports coding, tool use, streaming, JSON output, and multiple API formats. The most important practical qualification is its current serving status: although the deepseek-v4-pro identifier remains available, DeepSeek announced that requests using it would be routed to DeepSeek-V4.1-Flash until V4.1-Pro launches.
Outputs

What DeepSeek-V4-Pro can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Streaming JSON mode Structured output Prompt caching
Model profile

Performance characteristics

9/10 Reasoning
9/10 Coding
6/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family DeepSeek V4
Model type General Purpose
Context window 1M tokens
Maximum output 384K tokens
Release date 2026-04-24
Status Deprecated for independent serving; deepseek-v4-pro API requests are routed to DeepSeek-V4.1-Flash until V4.1-Pro launches
Deprecation date 2026-09-14
Knowledge cutoff notes

No exact knowledge cutoff for DeepSeek-V4-Pro was found in the reviewed official model card, API documentation, release announcement, or pricing documentation.

Model notes

DeepSeek-V4-Pro was preview-released on April 24, 2026 and reached general availability on August 13, 2026. The model card specifies 1.6T total parameters with 49B active per token, text-only modality, a 1M context window, and MIT-licensed open-source distribution for released assets. The API supports non-thinking and thinking modes, configurable thinking effort, tool calls, JSON output, streaming, OpenAI Responses API compatibility, Anthropic-compatible access, and prompt caching. On September 10, 2026, DeepSeek announced that all deepseek-v4-pro requests would route to DeepSeek-V4.1-Flash from September 14, 2026 until V4.1-Pro launches. The API change log also states that V4-Pro API services would continue after September 14, creating a distinction between identifier availability and independent V4-Pro model serving. The exact knowledge cutoff was not identified in the reviewed official documentation. Editorial scores are comparative estimates, not vendor ratings.

Cost

Model pricing

Input Official V4-Pro list price: US$0.66 per 1M cache-miss input tokens off-peak and US$1.32 peak; US$0.022 per 1M cache-hit input tokens off-peak and US$0.044 peak. DeepSeek announced that routed requests use V4.1-Flash rates.
Output Official V4-Pro list price: US$1.98 per 1M output tokens off-peak and US$3.96 peak. DeepSeek announced that routed requests use V4.1-Flash rates.
Model guide

DeepSeek-V4-Pro: A 1M-Context Reasoning Model with a Major Routing Caveat

DeepSeek-V4-Pro is DeepSeek’s high-capability V4 language model for complex reasoning, coding, long-context analysis, and agent workflows. It offers a 1-million-token context window, configurable thinking modes, tool calling, JSON output, streaming, prompt caching, and OpenAI-compatible interfaces. However, DeepSeek announced that requests using the deepseek-v4-pro identifier would be routed to DeepSeek-V4.1-Flash from September 14, 2026 until V4.1-Pro launches, so the identifier may no longer represent independently served V4-Pro weights.

What is DeepSeek-V4-Pro?

DeepSeek-V4-Pro is the larger, higher-capability model in DeepSeek’s V4 family. DeepSeek positions it as a general-purpose language model for advanced reasoning, software development, long-context processing, question answering, and agent workflows. An agent workflow is a process in which the model plans steps, calls tools, examines results, and continues working toward a goal rather than producing a single response.

The model was preview-released on April 24, 2026 and reached general availability on August 13, 2026. According to DeepSeek’s model documentation, it is a Mixture-of-Experts, or MoE, model with 1.6 trillion total parameters and approximately 49 billion active parameters per token. In practical terms, MoE architecture allows the system to contain a very large total model while activating only a subset of its components for each piece of text.

DeepSeek describes the V4 architecture as using compressed attention, sparse attention, manifold-constrained hyper-connections, and the Muon optimizer. These are technical design choices intended to improve efficiency, long-context handling, or training stability. They are provider-documented architecture details rather than independent evidence that every workload will be faster or more accurate.

Where it fits in DeepSeek’s lineup—and its current status

Within the V4 family, V4-Pro was introduced as the more capable option alongside DeepSeek-V4-Flash. Its intended role is to handle harder reasoning, coding, and tool-mediated tasks, while a Flash-class model is generally better suited to workloads that prioritize speed or lower cost.

There is an important difference between the model’s original specification and what the API may serve today. On September 10, 2026, DeepSeek announced that requests using the deepseek-v4-pro identifier would be routed to DeepSeek-V4.1-Flash starting at 04:00 UTC on September 14, 2026, until V4.1-Pro launches. DeepSeek’s API change log also states that API services for V4-Pro would continue after that date, with billing unchanged. Therefore, the identifier may remain accepted while the underlying served model is not an independently served V4-Pro checkpoint.

This matters for evaluations, reproducibility, and production systems. If an application depends on the original V4-Pro behavior, parameterization, or latency profile, developers should verify the currently served model with DeepSeek rather than assuming that the identifier guarantees the original weights.

Context window and reasoning modes

DeepSeek-V4-Pro has a documented context window of 1,000,000 tokens. A token is a unit of text used by the model; it may represent a word, part of a word, punctuation, or another text fragment. A one-million-token context is large enough for substantial repositories, lengthy technical archives, or long-running conversations, although the usable amount will depend on the application’s input, output, and prompt structure.

The documented maximum output length is 384,000 tokens. This is an upper API limit, not a guarantee that every request can produce an output of that size efficiently or economically.

The model supports non-thinking responses and thinking modes. Non-thinking mode is intended for tasks where speed and directness matter. Thinking modes allow more deliberate internal reasoning, with documented effort choices including low, high, and maximum effort where supported by the API. Higher effort can be appropriate for difficult mathematics, complex code changes, multi-step planning, or tool-based investigations, but it may increase response time and token consumption.

DeepSeek’s general-availability announcement highlighted improvements on agent-oriented evaluations including Terminal-Bench 2.1, NL2Repo, CyberGym, DeepSWE, Toolathlon-Verified, and AutomationBench. These are provider-reported claims and should not be treated as a guarantee of performance on a particular application. The supplied documentation does not specify an exact knowledge cutoff for V4-Pro.

Coding, tools, and agent use

V4-Pro is primarily aimed at text-based software and reasoning tasks. Its long context can be useful when an application must supply several related files, documentation, test output, issue descriptions, and prior decisions in one request. The model is also positioned for coding-agent workflows, where it can plan changes, call tools, inspect returned information, and revise its response.

The API supports tool calls, allowing an application to expose functions such as database lookups, search, file operations, or deployment checks. The model does not independently gain access to those systems; the surrounding application must define the tools, execute calls, and return results.

Supported interfaces include DeepSeek’s OpenAI-compatible Chat Completions API, the OpenAI Responses API format, and Anthropic-compatible access. Streaming is available for applications that want partial output as it is generated. DeepSeek also documented Codex integration for the model. The Responses API is stateless, meaning that an application must provide the relevant conversation history when making later requests.

JSON output is supported for applications that need machine-readable responses. This is useful for extraction, classification, structured planning, and workflow automation, but valid JSON alone does not guarantee that the output will match every application-specific schema without validation.

Supported modalities and important limitations

The reviewed V4-Pro model card identifies text as the supported modality. The model therefore should be evaluated as a text-input, text-output model rather than as a native vision, audio, or video system. DeepSeek’s broader product materials describe multimodal capabilities in other services, but those capabilities should not automatically be attributed to V4-Pro.

CapabilityDocumented status
Text inputSupported
Text outputSupported
Image inputNot listed for V4-Pro
Audio or video inputNot listed for V4-Pro
Image, audio, or video generationNot supported as a documented output modality
Tool callingSupported
StreamingSupported
JSON outputSupported
Fine-tuningNot verified in the supplied research

Other limitations are operational rather than purely technical. The routing change can make behavior, latency, and pricing differ from expectations based on the original V4-Pro documentation. Peak and off-peak pricing also differ, and prompt-cache hits have separate rates. Users requiring stable model weights, consistent benchmark conditions, or predictable deployment behavior should confirm the serving arrangement before committing to the model.

DeepSeek-V4-Pro pricing

DeepSeek’s published V4-Pro list price is calculated per one million tokens and varies according to time of day and whether an input is found in the prompt cache.

Token categoryOff-peak pricePeak price
Cache-hit inputUS$0.022 per 1M tokensUS$0.044 per 1M tokens
Cache-miss inputUS$0.66 per 1M tokensUS$1.32 per 1M tokens
OutputUS$1.98 per 1M tokensUS$3.96 per 1M tokens

These are the documented V4-Pro list prices. DeepSeek’s routing announcement says that requests sent through the deepseek-v4-pro identifier are billed at V4.1-Flash rates after the routing change, so users should check the current pricing page and effective model before estimating production costs. Cache-hit pricing can substantially reduce the cost of repeatedly sending unchanged prompt material, while generated output remains much more expensive than cached input.

Capability, speed, and cost trade-offs

V4-Pro’s intended trade-off is depth and capacity in exchange for more computation and potentially higher latency than a smaller or Flash-oriented model. Its 1-million-token context and maximum reasoning effort are most valuable when the task genuinely requires them. For a short factual answer, simple extraction, or high-volume low-latency interaction, a smaller model may be more economical and responsive.

The model’s published pricing is competitive for a high-capability model, especially when prompt caching is effective. However, cost comparisons must account for peak versus off-peak rates, cache misses, generated output, and the possibility that the identifier is currently routed to V4.1-Flash. The editorial assessment supplied for this page rates reasoning and coding highly, speed as moderate, and cost competitiveness positively; these are comparative editorial judgments, not DeepSeek’s official scores.

When to choose DeepSeek-V4-Pro

Choose V4-Pro when the original model’s documented role matches the workload and you can verify the current serving status. It is a strong candidate for:

  • Large codebase analysis and repository-scale software work.
  • Complex coding tasks that benefit from deliberate reasoning and tool calls.
  • Long documents, technical archives, and extended context-dependent questions.
  • Agent workflows involving planning, function calls, test results, and iterative correction.
  • Structured extraction or automation using JSON responses.
  • Applications that can benefit from prompt caching and competitive per-token pricing.

Another option may be more appropriate when the application needs native image, audio, or video understanding; very low latency; a guaranteed independently served V4-Pro checkpoint; or a model selected specifically for high-volume simple requests. A Flash-class model may be preferable for speed and cost, while a native multimodal model is a better fit for image or other non-text inputs.

Overall, DeepSeek-V4-Pro remains technically notable for its very large context window, configurable reasoning, coding orientation, and tool support. Its most consequential practical issue is not a missing feature but model identity: the deepseek-v4-pro name may now route to another model. Verify the effective model, price, and behavior before using the identifier as a fixed production dependency.


Answers to Frequently Asked Questions

What is the context window and maximum output length of DeepSeek-V4-Pro?
DeepSeek-V4-Pro has a documented context window of 1,000,000 tokens and a maximum output length of 384,000 tokens. These are API limits rather than guarantees that every request will use the full capacity efficiently or economically.
Does the deepseek-v4-pro API identifier still serve the original V4-Pro model?
Not necessarily. DeepSeek announced that requests using the deepseek-v4-pro identifier would be routed to DeepSeek-V4.1-Flash starting September 14, 2026, until V4.1-Pro launches. Developers should verify the effective served model with DeepSeek because the identifier may remain accepted without guaranteeing the original V4-Pro weights, behavior, or latency.
What are the main use cases for DeepSeek-V4-Pro?
The model is suited to large codebase analysis, complex software development, long technical documents, multi-step reasoning, tool-based agent workflows, structured extraction, and JSON-based automation. Its long context is useful for combining source files, documentation, test results, and issue descriptions in a single request.
How much does DeepSeek-V4-Pro cost?
The documented V4-Pro list price ranges from US$0.022 to US$0.044 per one million cache-hit input tokens, US$0.66 to US$1.32 per one million cache-miss input tokens, and US$1.98 to US$3.96 per one million output tokens, depending on off-peak or peak pricing. After the routing change, requests using the deepseek-v4-pro identifier are billed at V4.1-Flash rates, so users should check the current pricing page and effective model.
What is DeepSeek-V4-Pro?
DeepSeek-V4-Pro is a general-purpose Mixture-of-Experts language model designed for advanced reasoning, software development, long-context processing, question answering, and agent workflows. It is documented as having 1.6 trillion total parameters, approximately 49 billion active parameters per token, a one-million-token context window, tool calling, streaming, and JSON output support.


Sources 8
Provider

About DeepSeek