What is the DeepSeek API and when should you use it?
The DeepSeek API is a hosted developer platform for adding DeepSeek model capabilities to applications, coding tools, agents, and automated workflows. Its primary endpoint is https://api.deepseek.com. The platform follows OpenAI-compatible conventions, so developers familiar with the OpenAI client libraries or standard JSON HTTP requests can usually adapt quickly.
The platform currently exposes Chat Completions, the Responses API, an Anthropic-compatible Messages interface, and a Files API. The API supports text generation, reasoning, coding, streaming responses, function or tool calls, JSON Output, and image understanding through deepseek-flash.
DeepSeek is a reasonable choice when you want competitive token pricing, OpenAI-compatible interfaces, strong reasoning and coding performance, Chinese-language capability, or vision input through the Flash model. It may be a poor fit when your application requires a verified API-specific no-training default, a universal prompt-retention guarantee, mature image or video generation, or a provider-independent data-processing location.
How to get access and create an API key
API access is separate from the free consumer DeepSeek Web and App services. Create an account in the DeepSeek Platform, create an API key, and add balance as required for API usage. Every API request must include the key as a Bearer token.
Keep the key on your server or in a secret manager. Do not place it in browser JavaScript, mobile application binaries, source repositories, or client-side logs. A typical request uses these headers:
Authorization: Bearer $DEEPSEEK_API_KEY
Content-Type: application/jsonThe platform also provides a developer playground at platform.deepseek.com. Use it to inspect account access and experiment before embedding requests in an application.
Which API and model should you choose?
For most new integrations, choose between the OpenAI-compatible Chat Completions API and Responses API. Chat Completions is the simpler option for conversational requests and existing OpenAI-style integrations. Responses is better suited to semantic streaming events, response items, and agent-style workflows.
DeepSeek also offers an Anthropic-compatible Messages interface at https://api.deepseek.com/anthropic. This can reduce migration work for applications already built around Anthropic-format requests.
| Model | Best starting point | Vision | Documented concurrency |
|---|---|---|---|
deepseek-flash | Cost-sensitive general workloads, coding, reasoning, and image input | Yes | 2,500 |
deepseek-v4-pro | Higher-end reasoning or agent workloads where cost is less important | No | 500 |
deepseek-flash is the current identifier for DeepSeek-V4.1-Flash. Older V4 Flash identifiers may remain accepted as retired aliases, but new applications should use the current identifier. deepseek-v4-pro is text-focused and does not currently support vision input.
Make your first request
The following example uses Chat Completions. Set the API key in the environment before running it:
export DEEPSEEK_API_KEY="sk-your-key"
curl --fail-with-body https://api.deepseek.com/chat/completions
-H "Authorization: Bearer ${DEEPSEEK_API_KEY}"
-H "Content-Type: application/json"
-d '{
"model": "deepseek-flash",
"messages": [
{"role": "system", "content": "You are a concise technical assistant."},
{"role": "user", "content": "Explain what an API timeout is in one sentence."}
],
"thinking": {"type": "disabled"},
"stream": false
}'A normal request sends a model name and an array of messages. The system message sets broad behavior, while the user message contains the task. The thinking setting controls the reasoning mode in this request; the example disables it for a short answer.
The same request with the OpenAI Python SDK
DeepSeek's documented quick starts use the OpenAI SDK configured with DeepSeek's base URL:
# Install: pip install -U openai
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["DEEPSEEK_API_KEY"],
base_url="https://api.deepseek.com",
)
response = client.chat.completions.create(
model="deepseek-flash",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Give me three names for a weather app."},
],
thinking={"type": "disabled"},
)
print(response.choices[0].message.content)How to understand the response
Chat Completions returns a response object containing one or more choices. For a basic non-streaming request, the generated text is normally available at response.choices[0].message.content. The response also includes usage information that can be used for cost monitoring.
When tools are requested, the assistant message may contain tool calls rather than ordinary text. Your application must inspect the call, validate its arguments, execute the corresponding local function or service, and send the tool result back to the model. Never execute generated arguments without validation.
The Responses API uses response items rather than only chat messages. Its streaming events include names such as response.output_text.delta, and function calls are represented as output items. This format is useful when an application needs semantic event handling or agent-oriented workflows.
How DeepSeek API pricing works
DeepSeek charges by token usage rather than by a fixed monthly developer plan. Prices are listed per one million tokens and distinguish cached input, uncached input, and output. Peak and off-peak rates also apply. The documented ranges are:
| Model | Cached input | Uncached input | Output |
|---|---|---|---|
deepseek-flash | $0.003–$0.006 per 1M tokens | $0.15–$0.30 per 1M tokens | $0.60–$1.20 per 1M tokens |
deepseek-v4-pro | $0.022–$0.044 per 1M tokens | $0.66–$1.32 per 1M tokens | $1.98–$3.96 per 1M tokens |
Peak hours are listed as 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday. The lower end of each range represents off-peak pricing. Treat the official pricing page as authoritative because rates can change. DeepSeek charges against the account balance, using granted balance before topped-up balance when both are available.
What can developers build with the API?
Streaming output
Streaming sends generated output incrementally instead of waiting for the complete response. Chat Completions uses server-sent events and ends with data: [DONE]. The Responses API uses semantic events such as response.output_text.delta and a final completed, incomplete, or failed response event.
Streaming is useful for chat interfaces and long responses, but your application still needs connection timeouts, error handling, and a way to handle incomplete output.
Function and tool calling
Both listed models support tool calls. Define a function with its name, description, and parameter schema. The model can request the function, but your application—not DeepSeek—executes it. After execution, return the result to the model so it can produce a final answer.
Validate names, types, ranges, permissions, and authorization before executing a generated call. DeepSeek's documentation notes that generated arguments may not always conform perfectly to the declared schema.
JSON Output
DeepSeek supports JSON Output with response_format: {"type":"json_object"}. Tell the model explicitly to return JSON and describe the expected structure in the prompt:
json_response = client.chat.completions.create(
model="deepseek-flash",
messages=[
{"role": "system", "content": "Return valid JSON with keys title and tags."},
{"role": "user", "content": "Describe a secure API gateway."},
],
response_format={"type": "json_object"},
)
print(json_response.choices[0].message.content)This provides valid-JSON output, but it is not the same as a separately verified schema-constrained Structured Outputs feature. Parse and validate the returned JSON in your application.
Vision and uploaded files
deepseek-flash accepts images supplied through public URLs, base64 data URLs, or uploaded files. The Files API supports JPEG, PNG, GIF, and WebP images up to 64 MiB. Uploaded files are associated with the API key that created them.
Files may be retained permanently when no expiration is specified, or assigned an expiration between one hour and thirty days. Use expiration settings deliberately, especially when handling sensitive material. The public documentation does not establish a universal API prompt-retention period.
Advanced integration examples
Streaming with the Responses API
stream = client.responses.create(
model="deepseek-flash",
instructions="Answer clearly and briefly.",
input="What is exponential backoff?",
stream=True,
)
for event in stream:
if event.type == "response.output_text.delta":
print(event.delta, end="", flush=True)
print()Use this approach when your application wants semantic response events rather than only incremental Chat Completions deltas.
Referencing an uploaded image
with open("image.jpg", "rb") as image_file:
uploaded = client.files.create(file=image_file, purpose="user_data")
vision_response = client.chat.completions.create(
model="deepseek-flash",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "Describe this image."},
{"type": "file", "file_id": uploaded.id},
],
}],
)
print(vision_response.choices[0].message.content)SDKs and language support
Python and JavaScript or TypeScript applications can use the current OpenAI client libraries with DeepSeek's base URL. Anthropic-compatible integrations can use the Messages interface. PHP applications can call the API with ordinary HTTPS and cURL; no official DeepSeek PHP SDK was verified in the supplied documentation.
Limits and production considerations
Documented account-level concurrency is currently 2,500 for deepseek-flash and 500 for deepseek-v4-pro. Exceeding the applicable concurrency limit can return HTTP 429.
Other documented status codes include 400 for invalid formats, 401 for authentication failures, 402 for insufficient balance, 422 for invalid parameters, 500 for server errors, and 503 for overload. DeepSeek documents a ten-minute limit for requests that have not begun inference and keep-alive behavior for long-running requests.
Production clients should:
- Set explicit connection and read timeouts.
- Retry transient 429, 500, and 503 failures with exponential backoff.
- Do not retry authentication or parameter errors without changing the request.
- Monitor account balance, token usage, response errors, and concurrency.
- Exclude API keys and sensitive prompts from logs.
- Validate tool arguments and JSON responses before using them.
- Review file expiration and deletion behavior for uploaded content.
Privacy and data handling
DeepSeek's public privacy materials state that data may be stored and processed in the People's Republic of China. They describe retention as dependent on factors such as data type, sensitivity, purpose, and legal requirements. The materials consulted do not provide a clearly documented API-specific no-training default or a universal API retention period.
Organizations should review the applicable platform terms and privacy documentation before sending personal, confidential, regulated, or proprietary information. If your requirements demand a specific data residency policy, contractual retention period, or confirmed no-training treatment for API prompts, verify those controls directly before deployment.
Advantages and limitations
Where DeepSeek API is a good fit
- Applications that benefit from OpenAI-compatible Chat Completions or Responses interfaces.
- Cost-sensitive workloads that can use the documented Flash pricing.
- Coding, reasoning, mathematics, Chinese-language, and general text applications.
- Agent workflows that need tool calls, function calls, or Responses API events.
- Applications that need image understanding through
deepseek-flash. - Teams comfortable using standard HTTP or compatible Python and JavaScript SDKs.
Where it may be a poor fit
- Applications requiring image generation, video generation, or a broad consumer plugin ecosystem.
- Workloads that require vision support from the Pro model.
- Organizations that cannot accept processing or storage in China.
- Systems requiring a clearly documented API-specific no-training default or universal retention guarantee.
- Applications needing a separately verified schema-constrained Structured Outputs feature.
- High-volume systems that cannot accommodate the documented concurrency limits or variable service load.
A practical starting path
- Create a DeepSeek Platform account and API key.
- Start with
deepseek-flashfor ordinary text, coding, reasoning, or vision tests. - Use Chat Completions for a simple request and Responses when semantic streaming or agent-style workflows are needed.
- Record token usage and test both peak and off-peak cost assumptions.
- Add timeouts, exponential-backoff retries, balance monitoring, and response validation before production.
- Review privacy, retention, and data-location requirements before sending real user or business data.
