What is DeepSeek-V4-Pro?
DeepSeek-V4-Pro is the larger, higher-capability model in DeepSeek’s V4 family. DeepSeek positions it as a general-purpose language model for advanced reasoning, software development, long-context processing, question answering, and agent workflows. An agent workflow is a process in which the model plans steps, calls tools, examines results, and continues working toward a goal rather than producing a single response.
The model was preview-released on April 24, 2026 and reached general availability on August 13, 2026. According to DeepSeek’s model documentation, it is a Mixture-of-Experts, or MoE, model with 1.6 trillion total parameters and approximately 49 billion active parameters per token. In practical terms, MoE architecture allows the system to contain a very large total model while activating only a subset of its components for each piece of text.
DeepSeek describes the V4 architecture as using compressed attention, sparse attention, manifold-constrained hyper-connections, and the Muon optimizer. These are technical design choices intended to improve efficiency, long-context handling, or training stability. They are provider-documented architecture details rather than independent evidence that every workload will be faster or more accurate.
Where it fits in DeepSeek’s lineup—and its current status
Within the V4 family, V4-Pro was introduced as the more capable option alongside DeepSeek-V4-Flash. Its intended role is to handle harder reasoning, coding, and tool-mediated tasks, while a Flash-class model is generally better suited to workloads that prioritize speed or lower cost.
There is an important difference between the model’s original specification and what the API may serve today. On September 10, 2026, DeepSeek announced that requests using the deepseek-v4-pro identifier would be routed to DeepSeek-V4.1-Flash starting at 04:00 UTC on September 14, 2026, until V4.1-Pro launches. DeepSeek’s API change log also states that API services for V4-Pro would continue after that date, with billing unchanged. Therefore, the identifier may remain accepted while the underlying served model is not an independently served V4-Pro checkpoint.
This matters for evaluations, reproducibility, and production systems. If an application depends on the original V4-Pro behavior, parameterization, or latency profile, developers should verify the currently served model with DeepSeek rather than assuming that the identifier guarantees the original weights.
Context window and reasoning modes
DeepSeek-V4-Pro has a documented context window of 1,000,000 tokens. A token is a unit of text used by the model; it may represent a word, part of a word, punctuation, or another text fragment. A one-million-token context is large enough for substantial repositories, lengthy technical archives, or long-running conversations, although the usable amount will depend on the application’s input, output, and prompt structure.
The documented maximum output length is 384,000 tokens. This is an upper API limit, not a guarantee that every request can produce an output of that size efficiently or economically.
The model supports non-thinking responses and thinking modes. Non-thinking mode is intended for tasks where speed and directness matter. Thinking modes allow more deliberate internal reasoning, with documented effort choices including low, high, and maximum effort where supported by the API. Higher effort can be appropriate for difficult mathematics, complex code changes, multi-step planning, or tool-based investigations, but it may increase response time and token consumption.
DeepSeek’s general-availability announcement highlighted improvements on agent-oriented evaluations including Terminal-Bench 2.1, NL2Repo, CyberGym, DeepSWE, Toolathlon-Verified, and AutomationBench. These are provider-reported claims and should not be treated as a guarantee of performance on a particular application. The supplied documentation does not specify an exact knowledge cutoff for V4-Pro.
Coding, tools, and agent use
V4-Pro is primarily aimed at text-based software and reasoning tasks. Its long context can be useful when an application must supply several related files, documentation, test output, issue descriptions, and prior decisions in one request. The model is also positioned for coding-agent workflows, where it can plan changes, call tools, inspect returned information, and revise its response.
The API supports tool calls, allowing an application to expose functions such as database lookups, search, file operations, or deployment checks. The model does not independently gain access to those systems; the surrounding application must define the tools, execute calls, and return results.
Supported interfaces include DeepSeek’s OpenAI-compatible Chat Completions API, the OpenAI Responses API format, and Anthropic-compatible access. Streaming is available for applications that want partial output as it is generated. DeepSeek also documented Codex integration for the model. The Responses API is stateless, meaning that an application must provide the relevant conversation history when making later requests.
JSON output is supported for applications that need machine-readable responses. This is useful for extraction, classification, structured planning, and workflow automation, but valid JSON alone does not guarantee that the output will match every application-specific schema without validation.
Supported modalities and important limitations
The reviewed V4-Pro model card identifies text as the supported modality. The model therefore should be evaluated as a text-input, text-output model rather than as a native vision, audio, or video system. DeepSeek’s broader product materials describe multimodal capabilities in other services, but those capabilities should not automatically be attributed to V4-Pro.
| Capability | Documented status |
|---|---|
| Text input | Supported |
| Text output | Supported |
| Image input | Not listed for V4-Pro |
| Audio or video input | Not listed for V4-Pro |
| Image, audio, or video generation | Not supported as a documented output modality |
| Tool calling | Supported |
| Streaming | Supported |
| JSON output | Supported |
| Fine-tuning | Not verified in the supplied research |
Other limitations are operational rather than purely technical. The routing change can make behavior, latency, and pricing differ from expectations based on the original V4-Pro documentation. Peak and off-peak pricing also differ, and prompt-cache hits have separate rates. Users requiring stable model weights, consistent benchmark conditions, or predictable deployment behavior should confirm the serving arrangement before committing to the model.
DeepSeek-V4-Pro pricing
DeepSeek’s published V4-Pro list price is calculated per one million tokens and varies according to time of day and whether an input is found in the prompt cache.
| Token category | Off-peak price | Peak price |
|---|---|---|
| Cache-hit input | US$0.022 per 1M tokens | US$0.044 per 1M tokens |
| Cache-miss input | US$0.66 per 1M tokens | US$1.32 per 1M tokens |
| Output | US$1.98 per 1M tokens | US$3.96 per 1M tokens |
These are the documented V4-Pro list prices. DeepSeek’s routing announcement says that requests sent through the deepseek-v4-pro identifier are billed at V4.1-Flash rates after the routing change, so users should check the current pricing page and effective model before estimating production costs. Cache-hit pricing can substantially reduce the cost of repeatedly sending unchanged prompt material, while generated output remains much more expensive than cached input.
Capability, speed, and cost trade-offs
V4-Pro’s intended trade-off is depth and capacity in exchange for more computation and potentially higher latency than a smaller or Flash-oriented model. Its 1-million-token context and maximum reasoning effort are most valuable when the task genuinely requires them. For a short factual answer, simple extraction, or high-volume low-latency interaction, a smaller model may be more economical and responsive.
The model’s published pricing is competitive for a high-capability model, especially when prompt caching is effective. However, cost comparisons must account for peak versus off-peak rates, cache misses, generated output, and the possibility that the identifier is currently routed to V4.1-Flash. The editorial assessment supplied for this page rates reasoning and coding highly, speed as moderate, and cost competitiveness positively; these are comparative editorial judgments, not DeepSeek’s official scores.
When to choose DeepSeek-V4-Pro
Choose V4-Pro when the original model’s documented role matches the workload and you can verify the current serving status. It is a strong candidate for:
- Large codebase analysis and repository-scale software work.
- Complex coding tasks that benefit from deliberate reasoning and tool calls.
- Long documents, technical archives, and extended context-dependent questions.
- Agent workflows involving planning, function calls, test results, and iterative correction.
- Structured extraction or automation using JSON responses.
- Applications that can benefit from prompt caching and competitive per-token pricing.
Another option may be more appropriate when the application needs native image, audio, or video understanding; very low latency; a guaranteed independently served V4-Pro checkpoint; or a model selected specifically for high-volume simple requests. A Flash-class model may be preferable for speed and cost, while a native multimodal model is a better fit for image or other non-text inputs.
Overall, DeepSeek-V4-Pro remains technically notable for its very large context window, configurable reasoning, coding orientation, and tool support. Its most consequential practical issue is not a missing feature but model identity: the deepseek-v4-pro name may now route to another model. Verify the effective model, price, and behavior before using the identifier as a fixed production dependency.

