DeepSeek V3

DeepSeek-V3.2

by DeepSeek · Legacy open-weight model; former DeepSeek API aliases deepseek-chat and deepseek-reasoner were scheduled for discontinuation on 2026-07-24

DeepSeek-V3.2 is a 671B mixture-of-experts language model with approximately 37B active parameters per token, 128K context, DeepSeek Sparse Attention, thinking and non-thinking modes, tool use, function calling, JSON output, and MIT-licensed weights. It is primarily a legacy open-weight model after the scheduled retirement of its former first-party API aliases.

Text Reasoning Coding
DeepSeek-V3.2 is an open-weight reasoning and agent model built for general questions, coding, long-context analysis, and workflows that require external tools. Its main distinction is the combination of a large mixture-of-experts architecture, a documented 128K-token context window, and tool use that can continue during the model's thinking process. The model supports text input and text output only, so it is not a native vision or media-generation system. It was historically available through DeepSeek's API, but its former first-party API aliases are now treated as legacy following the provider's transition to a newer model generation.
Outputs

What DeepSeek-V3.2 can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Streaming JSON mode Structured output Prompt caching Batch API
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
7/10 Speed
10/10 Cost efficiency
Specifications

Technical details

Model family DeepSeek V3
Model type Reasoning
Context window 131K tokens
Maximum output 8K tokens
Release date 2025-12-01
Status Legacy open-weight model; former DeepSeek API aliases deepseek-chat and deepseek-reasoner were scheduled for discontinuation on 2026-07-24
Deprecation date 2026-07-24
Knowledge cutoff notes

No authoritative first-party knowledge-cutoff date was found for the exact DeepSeek-V3.2 model.

Model notes

DeepSeek-V3.2 succeeded DeepSeek-V3.2-Exp and introduced DeepSeek Sparse Attention through continued training. It supports thinking and non-thinking modes and integrates reasoning into tool use. The official model card documents text input and text output only, a 128K context window, 671B total parameters, and approximately 37B activated parameters per token. The model was released under the MIT License. The former API identifiers deepseek-chat and deepseek-reasoner were mapped to V3.2's non-thinking and thinking modes, respectively, before DeepSeek scheduled those legacy names for discontinuation on July 24, 2026. Pricing in this record reflects the model's official API period and is historical after the API transition. The maximum output value reflects the documented 8,192-token API limit; self-hosted inference limits can vary by serving stack.

Cost

Model pricing

Input $0.028 per 1M tokens cached; $0.28 per 1M tokens cache miss during official API availability
Output $0.42 per 1M tokens during official API availability
Model guide

DeepSeek-V3.2: Open-Weight Reasoning with Tool-Using Agents

DeepSeek-V3.2 is a 671-billion-parameter mixture-of-experts language model from DeepSeek, released on December 1, 2025. It combines thinking and non-thinking modes with DeepSeek Sparse Attention, long-context processing, function calling, structured JSON output, and reasoning during tool use. Its MIT-licensed weights remain useful for self-hosted reasoning and coding systems, although its former first-party API aliases were scheduled for discontinuation on July 24, 2026.

What is DeepSeek-V3.2?

DeepSeek-V3.2 is a large language model provided by DeepSeek and released on December 1, 2025. It is designed for general-purpose text generation, reasoning, programming, long-context analysis, and agentic workflows in which a model decides when to call an external tool.

The model is a mixture-of-experts system, often abbreviated MoE. It contains 671 billion total parameters, but approximately 37 billion parameters are activated for each token. In practical terms, the model has a very large overall capacity without using every parameter for every part of every response. That design can improve the balance between model capability and inference cost, although running the complete model locally still requires substantial hardware.

DeepSeek released the weights under the MIT License. This permits research, modification, and self-hosted deployment subject to the license terms. The open-weight release makes V3.2 different from a model that can only be accessed through a hosted chat product or a provider-controlled API.

Where V3.2 fits in DeepSeek's lineup

V3.2 succeeded DeepSeek-V3.2-Exp and retained the broad architecture of that experimental predecessor while adding DeepSeek Sparse Attention, or DSA. The model was also exposed through two API aliases: deepseek-chat for non-thinking use and deepseek-reasoner for thinking use.

Those aliases were scheduled for discontinuation on July 24, 2026, after DeepSeek introduced its V4 model generation. Based on the supplied lifecycle information, V3.2 should now be understood primarily as a legacy open-weight model rather than as DeepSeek's current first-choice hosted API model. Its weights and technical design can still be relevant for self-hosting, compatibility work, research, and deployments that specifically need the V3.2 model.

This lifecycle distinction matters when evaluating the model. Historical API prices and endpoint names describe the period when V3.2 was officially served, not a promise of current first-party availability. Developers starting a new hosted integration should verify DeepSeek's current model catalog before selecting V3.2.

Architecture, sparse attention, and context length

DeepSeek-V3.2 has a documented context length of 131,072 tokens, commonly described as 128K tokens. A context window is the amount of text the model can consider across the prompt and conversation, together with the generated response, subject to the serving system's limits.

V3.2 introduces DeepSeek Sparse Attention. Sparse attention is intended to reduce the computational cost of processing very long sequences by avoiding equally dense interaction between every token. The provider presents this as a way to make long-context processing more efficient while preserving model quality. The supplied research does not establish a universal performance improvement for every workload, so the practical result will depend on the serving stack, prompt structure, and task.

The model's maximum documented API output was 8,192 tokens. This is separate from the 128K-token context window: a large context does not mean that one response can contain 128K newly generated tokens. Self-hosted deployments may expose different output settings, but the available research only verifies the 8,192-token API limit for the official service.

Reasoning and tool use

V3.2 supports both thinking and non-thinking modes. Non-thinking mode is intended for ordinary responses where lower latency and shorter outputs are more important. Thinking mode gives the model room to work through more difficult problems before producing an answer, which is more suitable for complex reasoning, mathematics, coding, and multi-step analysis.

A notable feature is reasoning during tool use. The model can consider whether a tool is needed, select a tool, provide arguments for a function call, and continue its reasoning after the tool returns information. This makes it suitable for search agents, coding assistants, research workflows, and business automations that need several connected actions rather than a single text response.

Supported functions include tool calling, streaming responses, context caching, and structured JSON output. Function calling allows an application to expose operations such as a search function, database lookup, calculator, or code-management action. The model does not perform those operations by itself; the surrounding application must execute the requested function and return the result.

JSON mode is useful when an application needs predictable machine-readable output instead of prose. It does not remove the need for application-side validation, because a syntactically valid JSON response can still contain incorrect values or fail to follow a desired schema.

Coding and general-purpose work

DeepSeek-V3.2 is positioned for coding as well as general reasoning. Its combination of thinking modes, long context, structured output, and tool calls can support code explanation, debugging, repository analysis, implementation planning, and agent workflows that inspect or modify project files through external tools.

For example, a coding system could provide a repository summary, ask V3.2 to identify dependencies between modules, call a test-running tool, and then ask the model to interpret the test output and propose a patch. The model's role is to reason over the supplied information and request actions; the host application remains responsible for permissions, execution, validation, and safe deployment.

The supplied research provides editorial scores of 8 out of 10 for reasoning and coding. Those are evaluation judgments in this record, not provider-published benchmark results. They indicate that V3.2 is considered a strong fit for reasoning-heavy and programming-oriented workloads, but they should not be treated as a standardized comparison against every competing model.

Supported modalities and important limitations

The official model documentation identifies V3.2 as text-in and text-out. It does not natively generate or accept images, audio, video, speech, music, or embeddings. It also should not be confused with DeepSeek services that may provide visual understanding through another model or product family. A V3.2 deployment cannot be assumed to understand an uploaded image unless a separate vision-capable system is added around it.

Its lack of native media support rules it out for image analysis, video understanding, speech interfaces, and image or video generation. It is also a poor fit when a project requires a currently supported first-party DeepSeek API endpoint, a small model that can run on commodity hardware, or a simple low-latency response for every request.

The 671-billion-parameter total size is a major deployment consideration. Mixture-of-experts execution activates only part of the model for each token, but hosting the full weights and achieving useful throughput still requires substantial infrastructure. Quantization and serving optimizations may change the hardware requirements, but no specific hardware configuration is verified in the supplied research.

Historical pricing and API status

During its official API availability, DeepSeek-V3.2 was priced at approximately $0.28 per million cache-miss input tokens, $0.028 per million cache-hit input tokens, and $0.42 per million output tokens. Cache-hit pricing applied when reusable prompt content could be served from the provider's context cache. These figures are historical for the V3.2 API period and should not be interpreted as current first-party pricing after the legacy API transition.

The pricing made V3.2 particularly attractive for cost-sensitive reasoning and coding workloads, especially when repeated prompts or system instructions could benefit from caching. However, price alone does not determine total cost. Long reasoning traces, tool calls, repeated context, hosting requirements, and the engineering effort required for a self-managed deployment can all affect the real cost of a system.

Main strengths and trade-offs

  • Open deployment: MIT-licensed weights support research, modification, and self-hosted use.
  • Reasoning flexibility: Thinking and non-thinking modes allow a trade-off between deeper analysis and faster responses.
  • Agent workflows: Tool calls work in both thinking and non-thinking modes, supporting multi-step applications.
  • Long context: The 128K-token context window is useful for large documents, codebases, and extended task state.
  • Efficient architecture: Sparse attention is intended to reduce the cost of long-sequence processing, while the MoE design activates approximately 37B parameters per token.
  • Structured integration: JSON output, streaming, function calling, and caching support application development.
  • Deployment burden: The total model size makes local operation demanding despite partial parameter activation.
  • Legacy hosted status: The former official API aliases were scheduled for discontinuation, so new projects should not assume continuing first-party access.
  • No native media support: V3.2 is not a vision, audio, video, speech, or image-generation model.

When to choose DeepSeek-V3.2

Choose DeepSeek-V3.2 when you specifically need an open-weight reasoning model for self-hosted or research-oriented work, and when its long context and tool-using behavior are more important than small-model simplicity. It is a reasonable candidate for codebase analysis, research agents, multi-step automation, structured extraction, and deployments that can provide the infrastructure needed for a very large model.

Its historical API economics are also relevant for cost-sensitive applications that already have access to a compatible serving arrangement, particularly workloads that reuse large prompt prefixes and benefit from caching. The model's non-thinking mode can reduce unnecessary latency for straightforward requests, while thinking mode can be reserved for difficult tasks.

Another option is more appropriate when the project needs native image or audio input, media generation, lightweight local inference, or a currently supported hosted endpoint with a clear lifecycle. A smaller model may be preferable for high-throughput, low-latency applications. A vision-capable model is necessary for image understanding. A current first-party model should be evaluated instead when long-term API support matters more than preserving compatibility with V3.2.

Bottom line

DeepSeek-V3.2 is best understood as a large, open-weight text reasoning model with unusually strong support for tool-using agents and long-context workloads. Its 671B MoE architecture, approximately 37B active parameters per token, 128K context, thinking modes, and MIT license give it a useful position for technically capable teams that can manage substantial infrastructure.

Its main caveats are equally important: it is text-only, its documented API output limit was 8,192 tokens, local deployment is demanding, and its former first-party API aliases were scheduled for retirement. For a new project, the choice depends less on the historical price than on whether the project values open weights and V3.2-specific behavior enough to accept its deployment and lifecycle trade-offs.


Answers to Frequently Asked Questions

Is DeepSeek-V3.2 still suitable for new projects?
DeepSeek-V3.2 can still suit self-hosted, research, compatibility, codebase-analysis, and tool-using agent projects that specifically need its open weights and long-context reasoning. However, its former hosted API aliases were scheduled for discontinuation after DeepSeek introduced its V4 generation, so new integrations should verify current model availability and consider a currently supported model when long-term API access is important.
Is DeepSeek-V3.2 multimodal?
No. DeepSeek-V3.2 is documented as a text-in, text-out model and does not natively support images, audio, video, speech, music, or embeddings. Image or media understanding requires a separate vision- or media-capable system.
Can DeepSeek-V3.2 use tools and function calling?
Yes. DeepSeek-V3.2 can reason about whether a tool is needed, select a function, provide arguments, and continue reasoning after the tool returns a result. Applications can connect it to search, databases, calculators, code tools, and other external operations, but the host application must execute and validate those functions.
What is DeepSeek-V3.2?
DeepSeek-V3.2 is a large open-weight mixture-of-experts language model designed for text generation, reasoning, programming, long-context analysis, and tool-using agent workflows. It has 671 billion total parameters, activates approximately 37 billion parameters per token, and was released under the MIT License.
How large is DeepSeek-V3.2's context window?
DeepSeek-V3.2 has a documented context window of 131,072 tokens, commonly described as 128K tokens. Its historical official API output limit was 8,192 tokens, which is separate from the amount of input and conversation context the model can process.


Sources 9
Provider

About DeepSeek