DeepSeek V4

DeepSeek-V4-Flash

by DeepSeek · Retired; legacy API identifier temporarily routed to DeepSeek-V4.1-Flash from September 10, 2026

DeepSeek-V4-Flash was a lightweight DeepSeek V4 model for long-context reasoning, coding, document processing, and tool-using agents. It offered a 1-million-token context window, thinking and non-thinking modes, streaming, prompt caching, and historical low-cost API pricing. The original implementation was retired on September 10, 2026, and its legacy identifier was temporarily routed to DeepSeek-V4.1-Flash.

Text Reasoning Coding
DeepSeek-V4-Flash was the economical, speed-oriented member of DeepSeek’s V4 model family. It was designed for applications that needed to process very large prompts, generate code, reason through complex tasks, or support tool-using agents without the cost and latency associated with a larger flagship model. The model offered a 1-million-token context window, text generation, thinking and non-thinking modes, streaming, tool use, and automatic prompt caching. However, the original V4-Flash implementation was retired on September 10, 2026, and its former API identifier was temporarily redirected to DeepSeek-V4.1-Flash.
Outputs

What DeepSeek-V4-Flash can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Streaming Prompt caching
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family DeepSeek V4
Model type Lightweight
Context window 1M tokens
Release date 2026-04-24
Status Retired; legacy API identifier temporarily routed to DeepSeek-V4.1-Flash from September 10, 2026
Deprecation date 2026-09-10
Shutdown date 2026-09-10
Knowledge cutoff notes

No direct authoritative knowledge-cutoff date for the exact DeepSeek-V4-Flash model was found in the reviewed first-party documentation.

Model notes

DeepSeek-V4-Flash was announced as part of the DeepSeek-V4 preview on April 24, 2026 and entered public beta on July 31, 2026. DeepSeek described it as a 284-billion-parameter mixture-of-experts model with approximately 13 billion active parameters and a 1-million-token context window. Historical API pricing was listed as US$0.14 per million cache-miss input tokens, US$0.0028 per million cache-hit input tokens, and US$0.28 per million output tokens, subject to peak/off-peak pricing. The exact model was retired on September 10, 2026. The legacy identifier deepseek-v4-flash was temporarily routed to DeepSeek-V4.1-Flash and should not be treated as an independently active model. Editorial scores are comparative estimates, not provider-published ratings.

Cost

Model pricing

Input $0.14 per 1 million input tokens
Output $0.28 per 1 million output tokens
Model guide

DeepSeek-V4-Flash: A Retired Million-Token Model Built for Fast, Low-Cost Reasoning

DeepSeek-V4-Flash was DeepSeek’s lightweight V4 model for long-context reasoning, coding, document processing, and agent workflows. It combined a 1-million-token context window with lower historical API pricing and faster positioning than V4-Pro, but the original model was retired on September 10, 2026. Its legacy API identifier was temporarily routed to DeepSeek-V4.1-Flash, so it should not be treated as an independently active model for new deployments.

What was DeepSeek-V4-Flash?

DeepSeek-V4-Flash was a lightweight large language model from DeepSeek, a Chinese artificial intelligence company. It was introduced on April 24, 2026, as part of the DeepSeek-V4 preview and later entered public beta on July 31, 2026. The model was intended to provide a faster and less expensive alternative to DeepSeek-V4-Pro for workloads involving long documents, software development, reasoning, and AI agents.

The model accepted text and produced text, including ordinary answers, explanations, reasoning traces where enabled, and generated code. It was available through DeepSeek’s API under the identifier deepseek-v4-flash. During its preview period, DeepSeek documented both OpenAI-compatible Chat Completions access and Anthropic-compatible API access.

Despite its name, the original V4-Flash model was not a vision model. DeepSeek separately described the experimental deepseek-v4-flash-vision-exp model as adding image understanding. V4-Flash itself should therefore be evaluated as a text model rather than as a general multimodal assistant.

Model design and 1-million-token context

DeepSeek described V4-Flash as a mixture-of-experts, or MoE, model. An MoE model contains multiple specialized sections, called experts, but activates only a portion of them for each token processed. According to DeepSeek’s launch information, V4-Flash had approximately 284 billion total parameters and about 13 billion active parameters per token.

The distinction between total and active parameters helps explain the model’s positioning. Its total architecture was very large, but sparse activation was intended to reduce the computation required for each response. DeepSeek also described sparse-attention and token-compression techniques designed to make very long inputs more practical.

Its documented context window was up to 1 million tokens. A context window is the amount of text the model can consider in a single request, including the prompt and relevant conversation or document content. One million tokens is substantially larger than the context windows commonly used by smaller general-purpose models, making V4-Flash suitable for large codebases, lengthy technical documentation, extensive transcripts, or multi-document analysis.

The supplied specifications do not provide a verified maximum output-token limit for the original V4-Flash implementation. Applications should therefore avoid assuming that the full million-token capacity can be returned as output; the published figure describes context capacity, not a guaranteed response length.

Reasoning, coding, and agent capabilities

V4-Flash supported both non-thinking and thinking modes. Non-thinking mode was intended for more direct responses, while thinking mode allocated more processing to tasks that benefited from step-by-step reasoning. This made the model relevant to mathematics, technical analysis, planning, and other problems where a quick surface-level answer might be insufficient.

Software development was another central use case. The model was designed for code generation, code explanation, debugging, and code-agent workflows. Its long context could be useful when an application needed to provide multiple files, repository documentation, issue history, or tool results in one request. DeepSeek also positioned the model for agentic applications, where the language model decides when to call external tools and uses returned information to continue a task.

Tool use was supported, and the public-beta version added native support for DeepSeek’s Responses API format. DeepSeek stated that the model was adapted for Codex-style coding agents. Streaming was also supported, allowing an application to display a response progressively instead of waiting for the complete generation.

These features made V4-Flash more than a simple question-answering model, but they did not turn it into a fully autonomous software system. Tool execution, permissions, external data access, and application safeguards remained the responsibility of the surrounding software.

Historical API pricing

DeepSeek positioned V4-Flash as the cost-efficient option in the V4 family. During its documented availability, historical pricing was listed at US$0.14 per million input tokens for cache misses, US$0.0028 per million input tokens for cache hits, and US$0.28 per million output tokens. The provider also documented peak and off-peak pricing periods, and stated that prices could change.

Automatic prompt caching was available. When a repeated portion of a prompt was served from cache, the cache-hit input rate was much lower than the cache-miss rate. This could matter for applications that repeatedly send the same system instructions, tool definitions, repository context, or document prefixes.

These figures are historical rather than current prices for an active V4-Flash implementation. The original model was retired, so developers should not use the listed rates to estimate the cost of a new deployment without checking DeepSeek’s current API documentation and the model identifier actually selected by their application.

Modalities and supported features

The original DeepSeek-V4-Flash specification identifies text as both its input and output modality. It did not provide native image, audio, or video input, and it did not generate images, audio, or video. The model therefore suited text-centric applications such as coding assistants, document analysis, long-form reasoning, and tool-driven workflows.

CapabilityDeepSeek-V4-Flash
InputText
OutputText, including reasoning and code
Context windowUp to 1 million tokens
Reasoning modesThinking and non-thinking modes
Tool useSupported
StreamingSupported
Prompt cachingAutomatic caching documented
Image, audio, or video inputNot supported by the original model specification
Image, audio, or video outputNot supported

The supplied research does not verify fine-tuning support, batch API availability, structured-output support, or a separate JSON mode for the original model. Those capabilities should be treated as unknown rather than assumed from the model’s tool-use or API compatibility features.

Main strengths and limitations

Where V4-Flash was strong

  • Large context capacity: The 1-million-token window was its clearest technical distinction and supported very large documents, repositories, and multi-step prompts.
  • Cost-conscious positioning: Its historical token rates were designed to make long-context and high-volume workloads more affordable than using a larger flagship model.
  • Speed-oriented design: DeepSeek positioned it as faster and lighter than V4-Pro, making it more suitable for interactive applications and agent loops where latency matters.
  • Reasoning and coding: Thinking and non-thinking modes supported a choice between more direct responses and additional reasoning effort.
  • Agent integration: Tool use, streaming, prompt caching, and Responses API support made it suitable for developer-built agents.

Where it was limited

  • Retired implementation: The original model stopped being an independently active target on September 10, 2026.
  • Text-only design: It was not the appropriate V4 model for image understanding, audio, video, or media generation.
  • No verified output ceiling: DeepSeek documented the context window but did not provide a verified maximum output-token figure in the supplied research.
  • Historical pricing only: The published prices should not be treated as current availability or a current contract.
  • Compatibility ambiguity: Requests using the old API name were temporarily routed to DeepSeek-V4.1-Flash, meaning the identifier no longer reliably represented the original architecture.

When would V4-Flash have been the right choice?

Before retirement, V4-Flash was a sensible choice for a text-based application that needed a large working context, strong coding and reasoning capabilities, and lower cost or latency than a larger model. Examples included analyzing a long technical specification, reviewing a substantial codebase, maintaining context across an agent’s tool calls, or processing repeated prompts where automatic caching could reduce input costs.

Its speed-and-cost trade-off also made it more appropriate than a heavyweight model when the application handled many routine requests and did not need the strongest available reasoning quality for every interaction. A developer could use non-thinking mode for straightforward tasks and reserve thinking mode for more complex problems.

Another model would have been more appropriate when the application required image understanding, media generation, audio or video processing, or a currently supported canonical model identifier. For new DeepSeek deployments, the supplied research identifies DeepSeek-V4.1-Flash as the successor and states that the canonical current identifier became deepseek-flash. Developers should use the provider’s current documentation to confirm the supported successor, capabilities, and pricing rather than building new systems around the retired deepseek-v4-flash name.

Retirement and current status

DeepSeek retired the original V4-Flash model on September 10, 2026, following the release of DeepSeek-V4.1-Flash. The legacy identifiers deepseek-v4-flash and deepseek-v4-flash-vision-exp were temporarily retained for compatibility, but requests were routed to V4.1-Flash rather than the original implementations.

This status is important for both evaluation and migration. Historical descriptions of V4-Flash still explain its intended design, context capacity, pricing, and use cases, but they do not guarantee that an API request using the old identifier will reproduce the same behavior. New applications should select the current successor identifier documented by DeepSeek and retest prompts, tool calls, output handling, latency, and cost before production use.

In summary, DeepSeek-V4-Flash was notable for combining a million-token context window with a relatively economical, speed-focused design for reasoning, coding, and agents. Its most important practical limitation is no longer a missing feature: it is that the original model has been retired.


Answers to Frequently Asked Questions

Is DeepSeek-V4-Flash still available?
The original DeepSeek-V4-Flash implementation was retired on September 10, 2026, after the release of DeepSeek-V4.1-Flash. Legacy identifiers were temporarily retained for compatibility but routed requests to the newer model, so developers should use DeepSeek’s current documentation and canonical model identifier for new applications.
What did DeepSeek-V4-Flash cost historically?
Historical pricing was listed at US$0.14 per million input tokens for cache misses, US$0.0028 per million input tokens for cache hits, and US$0.28 per million output tokens. These rates are no longer reliable for new deployments because the original model was retired and prices may change.
Did DeepSeek-V4-Flash support images, audio, or video?
No. The original DeepSeek-V4-Flash specification supported text input and text output only, including generated code and reasoning. It did not natively accept images, audio, or video, and it did not generate media. DeepSeek described a separate experimental vision model for image understanding.
What was DeepSeek-V4-Flash?
DeepSeek-V4-Flash was a text-only mixture-of-experts language model designed for fast, low-cost reasoning, coding, long-document analysis, and AI-agent workflows. It supported thinking and non-thinking modes, tool use, streaming, prompt caching, and a context window of up to 1 million tokens.
How large was DeepSeek-V4-Flash’s context window?
DeepSeek-V4-Flash supported a documented context window of up to 1 million tokens. This made it suitable for processing large codebases, lengthy technical documents, extensive transcripts, and multi-document prompts. The million-token figure described context capacity, not a guaranteed maximum output length.


Sources 5
Provider

About DeepSeek