DeepSeek-V3

DeepSeek-V3

by DeepSeek · Superseded and no longer current as a first-party hosted API model; open-weight checkpoint remains available for self-hosted and third-party deployment.

DeepSeek-V3 is a 671-billion-parameter open-weight Mixture-of-Experts model released in December 2024. It supports text generation, coding, long-context processing, streaming, JSON responses, and tool-oriented integrations. Its original hosted API pricing was very low, but the model has been superseded by newer generations in DeepSeek's rolling first-party API lineup.

Text Reasoning Coding
DeepSeek-V3 is a general-purpose, text-only language model designed for conversation, coding, mathematics, writing, and other long-context tasks. Its sparse Mixture-of-Experts architecture gives it 671 billion total parameters while activating about 37 billion for each token. The model remains relevant because DeepSeek released open-weight checkpoints for self-hosted and third-party deployment, although newer DeepSeek generations have replaced the original V3 in the provider's active hosted API lineup.
Outputs

What DeepSeek-V3 can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Streaming Fine-tuning JSON mode Prompt caching
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
8/10 Speed
10/10 Cost efficiency
Specifications

Technical details

Model family DeepSeek-V3
Model type General Purpose
Context window 128K tokens
Maximum output 8K tokens
Release date 2024-12-26
Status Superseded and no longer current as a first-party hosted API model; open-weight checkpoint remains available for self-hosted and third-party deployment.
Knowledge cutoff notes

No direct authoritative knowledge-cutoff date for the exact original DeepSeek-V3 checkpoint was found in DeepSeek's official model announcement, model repository, or API documentation. July 2024 is sometimes attributed to the shared V3 base lineage in secondary references, but it is not used here as a verified exact-model value.

Model notes

DeepSeek-V3 was released on December 26, 2024 with 671B total parameters and approximately 37B activated parameters per token. The original model was text-only and explicitly did not support multimodal input or output. DeepSeek's hosted API initially exposed the model through the deepseek-chat identifier, but that rolling identifier was subsequently upgraded to DeepSeek-V3-0324, DeepSeek-V3.1, DeepSeek-V3.1-Terminus, and newer generations. The listed prices are historical launch prices, not current prices for an active V3 endpoint. The 128K context value reflects published V3 family/model materials; hosted output limits varied by service version, with 8K commonly documented for the original non-reasoning API mode. DeepSeek does not publish a direct authoritative knowledge-cutoff date for the exact original V3 checkpoint. Editorial scores are comparative estimates, not vendor specifications.

Cost

Model pricing

Input $0.27 per 1M tokens for cache misses; $0.07 per 1M tokens for cache hits, historical launch pricing
Output $1.10 per 1M tokens, historical launch pricing
Model guide

DeepSeek-V3: An Open-Weight MoE Model for Low-Cost Text and Coding

DeepSeek-V3 is a 671-billion-parameter open-weight Mixture-of-Experts language model released by DeepSeek in December 2024. It activates approximately 37 billion parameters per token, supports text generation, coding, long-context work, streaming, JSON output, and tool-oriented API workflows. Its historical hosted pricing was unusually low, but the original V3 model is no longer the current model behind DeepSeek's rolling first-party API identifier.

What is DeepSeek-V3?

DeepSeek-V3 is a large language model from DeepSeek, released on December 26, 2024. It generates and understands text, making it suitable for tasks such as answering questions, drafting documents, summarizing material, translating languages, solving mathematics problems, and writing or reviewing code.

The model is open-weight rather than a conventional closed hosted model. DeepSeek published model checkpoints and technical materials, allowing organizations and researchers to run the model through suitable infrastructure or use third-party inference providers. Open-weight does not mean that deployment is simple: the full model is exceptionally large and generally requires distributed hardware, quantization, or a specialized hosting service.

DeepSeek-V3 is now best understood as an open-weight model and an important 2024 release, rather than as DeepSeek's current first-party hosted model. DeepSeek's rolling deepseek-chat API identifier was subsequently upgraded through newer V3-family generations and later models. Users who specifically need the original V3 behavior should therefore verify the checkpoint or endpoint version instead of assuming that the current rolling identifier serves the original weights.

Architecture and scale

DeepSeek-V3 uses a sparse Mixture-of-Experts, or MoE, architecture. In a conventional dense model, most or all parameters participate in processing every token. In an MoE model, a routing system selects a subset of specialized expert networks for each token. DeepSeek-V3 has 671 billion total parameters but activates approximately 37 billion parameters per token.

This distinction matters operationally. The total parameter count reflects the model's overall capacity and storage requirements, while the activated count helps explain why inference can be more computationally efficient than running a dense 671-billion-parameter model. It does not make the model lightweight: the complete downloadable model files are approximately 685 billion parameters when the multi-token-prediction module is included.

The technical design builds on DeepSeekMoE and Multi-head Latent Attention approaches used in earlier DeepSeek systems. DeepSeek also described an auxiliary-loss-free load-balancing method intended to reduce some quality trade-offs associated with conventional expert-balancing objectives. Multi-token prediction was included to improve training performance and support techniques such as speculative decoding during inference.

DeepSeek reported approximately 2.788 million H800 GPU hours for the complete training process. That figure is a provider-reported training claim, not a guarantee of the hardware or cost required for an individual deployment. The model was trained with an FP8 mixed-precision framework and later went through supervised fine-tuning and reinforcement-learning stages.

Capabilities and supported modalities

The original DeepSeek-V3 is a text-in, text-out model. It does not natively accept images, audio, or video, and it does not generate images, speech, music, audio, or video. It also does not provide native embeddings. These limitations distinguish it from multimodal assistants and specialized media-generation models.

Its useful capabilities are concentrated in language and code. DeepSeek-V3 can generate prose, follow instructions, summarize and transform text, assist with mathematics, translate content, and produce or explain code. It can also be used in conversational applications that need streaming responses, structured JSON output, or external tool and function calls.

Tool calling is an integration capability rather than a separate non-text output modality. An application can expose functions such as database lookup, retrieval, or business actions to the model, but the application remains responsible for executing those functions and validating the returned arguments. Similarly, JSON output helps an application receive machine-readable text; it should not be confused with a separately documented JSON-schema structured-output feature. The supplied model research records JSON mode as available through the API while recording structured output as unsupported.

Context window and output limits

Published DeepSeek-V3 model materials describe a context window of up to 128,000 tokens. A context window is the amount of input and output text that can be considered within one request, although the exact usable balance depends on the serving system and its output reservation.

The original hosted API commonly documented a maximum output setting of 8,192 tokens for the non-reasoning API mode. This limit should be treated as a historical service-specific value rather than a universal property of every self-hosted implementation. A third-party provider may impose different context, input, output, rate, or concurrency limits.

The 128K context is useful for long documents, large code repositories, extended conversations, and multi-document analysis. It does not guarantee that every long prompt will receive equally reliable attention throughout the entire context. For important work, users should still retrieve relevant passages, divide very large tasks into stages, and verify generated conclusions.

Reasoning, coding, and tool use

DeepSeek-V3 was positioned as a general-purpose model with strong reasoning and coding performance for its release period. It can work through mathematical problems, explain intermediate ideas, generate code, identify likely bugs, and help transform requirements into implementation steps. These are model capabilities rather than guarantees of correctness, so generated calculations and software should be checked.

DeepSeek-V3 is not the same as a dedicated reasoning model with a separately exposed reasoning process or a guaranteed deliberation budget. Its reasoning usefulness comes from general language-model behavior, instruction tuning, and the surrounding application design. For routine coding and text tasks, that can be sufficient; for especially difficult problems, a newer reasoning-focused model may be more appropriate if the provider or host makes one available.

Streaming allows an application to display generated text progressively instead of waiting for the entire response. Tool or function calling can connect the model to external services, but the model itself has no built-in authoritative live knowledge. Web search, retrieval, code execution, authentication, permissions, and action safety must be supplied by the integrating application. The original model's knowledge is static, and no authoritative exact knowledge-cutoff date was published for the original V3 checkpoint.

Historical pricing and current API status

At launch, DeepSeek announced historical hosted API pricing of $0.27 per million input tokens for cache misses, $0.07 per million input tokens for cache hits, and $1.10 per million output tokens. Cache-hit pricing applied when the service could reuse an eligible prompt prefix; it was not simply a second universal input rate. These prices were notably low compared with many contemporary frontier-model APIs.

Those figures should not be treated as current pricing for an active original V3 endpoint. DeepSeek's API change log records that the legacy deepseek-chat identifier was upgraded first to DeepSeek-V3, then to DeepSeek-V3-0324, DeepSeek-V3.1, DeepSeek-V3.1-Terminus, and later generations. The exact current price, availability, and behavior of a rolling identifier may therefore differ from the original V3 launch specification.

For a reproducible deployment, use the published open-weight checkpoint or a host that explicitly identifies the original DeepSeek-V3 version. For a managed API, confirm the provider's current model ID, pricing page, context limits, retention policy, and deprecation status before building against it.

Main strengths and limitations

Strengths

  • Open-weight access: Researchers and organizations can evaluate or deploy the checkpoint outside a single closed API, subject to the applicable model license and infrastructure requirements.
  • Efficient sparse design: Approximately 37 billion parameters are activated per token even though the model contains 671 billion total parameters.
  • Broad text capability: The model covers general conversation, writing, mathematics, translation, summarization, and coding rather than targeting only one narrow task.
  • Long-context support: The published 128K context window can support large documents, extended code, and multi-step text workflows.
  • Low historical API cost: Its original launch prices were especially attractive for cost-sensitive text and code workloads.
  • Integration features: Streaming, JSON responses, and tool or function calling support application-oriented use cases.

Limitations

  • Very demanding deployment: The total model size makes ordinary single-machine self-hosting impractical without substantial memory, distributed inference, or aggressive quantization.
  • Text only: The original V3 checkpoint does not accept or produce images, audio, video, speech, or music.
  • No verified exact knowledge cutoff: DeepSeek did not publish an authoritative cutoff date for the original checkpoint, so users should not infer that it knows current events.
  • Superseded hosted identity: The original model is no longer reliably represented by the provider's rolling first-party API identifier.
  • Integration-dependent features: Tool execution, web access, retrieval, JSON validation, and other application behaviors depend on the serving platform.
  • Operational uncertainty: Third-party hosting can differ in latency, availability, batching, quantization quality, privacy terms, and maximum request size.

When to choose DeepSeek-V3

Choose DeepSeek-V3 when you need an open-weight general language model for research, local or private inference, coding assistance, long-context text processing, or comparison of sparse Mixture-of-Experts architectures. It is particularly relevant when control over deployment matters more than having a turnkey multimodal assistant, and when historical or third-party inference pricing is a major consideration.

It can also be a useful choice for teams that want to test model behavior independently of a single provider's current API lineup. Running a fixed checkpoint makes experiments more reproducible than relying on a rolling model name that may silently change over time.

Another option may be more appropriate when the application requires image understanding, image generation, speech, video, native multimodal interaction, or a maintained first-party endpoint with clearly current specifications. A newer DeepSeek model may also be preferable for production API work if it provides better current support, reliability, or reasoning performance. Likewise, a smaller dense model may be a better fit when low latency, modest hardware, or simple deployment is more important than the capacity and openness of this very large checkpoint.

Practical evaluation guidance

Before adopting DeepSeek-V3, decide whether the requirement is for the original checkpoint or simply for a current DeepSeek text model. For the original checkpoint, verify the model files, license, quantization format, inference engine, and hardware plan. For hosted use, verify the exact model ID rather than relying on the historical deepseek-chat name.

Evaluate representative prompts rather than relying only on general scores. Test the model on the languages, codebases, document lengths, JSON schemas, tool arguments, and failure cases that matter to the application. Measure end-to-end latency and cost, including prompt caching, output length, infrastructure, and monitoring. For sensitive data, review the selected host's retention and training policies separately from DeepSeek's consumer-service policies.

DeepSeek-V3 remains a technically significant open-weight text model, but its best present-day role is as a large, reproducible checkpoint for capable text and code workloads—not as a current multimodal assistant or automatically current DeepSeek API target.


Answers to Frequently Asked Questions

Is DeepSeek-V3 still available through the current deepseek-chat API?
The rolling deepseek-chat identifier has been upgraded through newer DeepSeek-V3-family generations and later models, so it should not be assumed to serve the original V3 checkpoint. For reproducible use, verify the exact checkpoint or model ID with the provider and confirm current pricing, context limits, and availability.
What is DeepSeek-V3's context window and output limit?
Published model materials describe a context window of up to 128,000 tokens. The original hosted API commonly documented a maximum output of 8,192 tokens in non-reasoning mode, but self-hosted deployments and third-party providers may impose different limits.
Does DeepSeek-V3 support images, audio, or video?
No. The original DeepSeek-V3 checkpoint is a text-only model: it accepts text and produces text. It does not natively understand or generate images, audio, speech, music, or video, and it does not provide native embeddings.
What is DeepSeek-V3?
DeepSeek-V3 is an open-weight large language model released on December 26, 2024, for text generation, question answering, summarization, translation, mathematics, and coding. Its original checkpoint is designed for text-in, text-out use and can be self-hosted or accessed through third-party inference providers.
How many parameters does DeepSeek-V3 have?
DeepSeek-V3 has 671 billion total parameters and activates approximately 37 billion parameters per token through its sparse Mixture-of-Experts architecture. The model is still extremely large and generally requires distributed hardware, quantization, or specialized hosting.


Sources 8
Provider

About DeepSeek