Developer platform

DeepSeek API

Developer overview for DeepSeek, including API access, pricing, SDK support, endpoints, capabilities and platform policies.

API key Required
Primary API OpenAI-compatible Responses API and Chat Completions API
SDK support Official examples use the OpenAI Python and JavaScript SDKs configured with DeepSeek's base URL. Anthropic-compatible integrations are also documented. PHP can use raw HTTP; no official DeepSeek PHP SDK was verified.
Rate limits Documented account-level concurrency is 2,500 for deepseek-flash and 500 for deepseek-v4-pro. Exceeding concurrency returns HTTP 429. DeepSeek also documents request keep-alive behavior and a ten-minute limit for requests that have not begun inference.
Platform

API overview

Endpoints

API access

Base URL https://api.deepseek.com
Primary API OpenAI-compatible Responses API and Chat Completions API
Pricing

API pricing

Pricing model Usage-based token pricing with separate cached-input, uncached-input, and output rates; peak and off-peak pricing applies.

deepseek-flash: $0.003–$0.006 cached input, $0.15–$0.30 uncached input, and $0.60–$1.20 output per 1M tokens. deepseek-v4-pro: $0.022–$0.044 cached input, $0.66–$1.32 uncached input, and $1.98–$3.96 output per 1M tokens.

Developer experience

SDKs & usability

SDKs Official examples use the OpenAI Python and JavaScript SDKs configured with DeepSeek's base URL. Anthropic-compatible integrations are also documented. PHP can use raw HTTP; no official DeepSeek PHP SDK was verified.
Ease of use High for developers familiar with OpenAI-compatible APIs. The platform uses standard HTTP, Bearer authentication, JSON requests, and compatible Python and JavaScript SDK configuration.
Documentation Good and actively updated, with API reference pages, quick starts, feature guides, pricing, error codes, files, vision, Responses API, and agent integration documentation.
Latency No fixed latency SLA or guaranteed response-time figure was verified. Latency varies by model, load, request size, reasoning effort, and peak/off-peak scheduling. Long-running requests may remain connected with keep-alive lines or SSE comments.
Features

API capabilities

✓ Streaming
✓ Function calling
✓ Assistants API
✓ File uploads
✓ Image input
✓ Playground
Feature notes

The current API platform supports Chat Completions, Responses, Anthropic-compatible Messages, streaming, function/tool calling, JSON Output, reasoning controls, user_id isolation, and image input through deepseek-flash. The Responses API supports semantic SSE events and function-call output items. The Files API supports JPEG, PNG, GIF, and WebP uploads up to 64 MiB, with optional expiration from one hour to thirty days or permanent retention when no expiration is set. deepseek-flash has a 1M-token context window and supports vision; deepseek-v4-pro is listed as not supporting vision. DeepSeek Harness is documented as a first-party developer-preview agent harness. No separately documented schema-constrained Structured Outputs feature or public fine-tuning API was verified. Legacy deepseek-v4-flash and deepseek-v4-flash-vision-exp names remain accepted but are retired aliases served by the current Flash model.

Examples

API examples

export DEEPSEEK_API_KEY="sk-your-key"

curl --fail-with-body https://api.deepseek.com/chat/completions \
  -H "Authorization: Bearer ${DEEPSEEK_API_KEY}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-flash",
    "messages": [
      {"role": "system", "content": "You are a concise technical assistant."},
      {"role": "user", "content": "Explain what an API timeout is in one sentence."}
    ],
    "thinking": {"type": "disabled"},
    "stream": false
  }' | jq -r '.choices[0].message.content'

# Streaming with JSON Output
curl --fail-with-body https://api.deepseek.com/chat/completions \
  -H "Authorization: Bearer ${DEEPSEEK_API_KEY}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-flash",
    "messages": [
      {"role": "system", "content": "Return valid JSON with keys answer and confidence."},
      {"role": "user", "content": "Summarize why retries need backoff."}
    ],
    "response_format": {"type": "json_object"},
    "stream": true,
    "stream_options": {"include_usage": true}
  }'
# Install: pip install -U openai
import json
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["DEEPSEEK_API_KEY"],
    base_url="https://api.deepseek.com",
)

response = client.chat.completions.create(
    model="deepseek-flash",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Give me three names for a weather app."},
    ],
    thinking={"type": "disabled"},
)
print(response.choices[0].message.content)

# Structured JSON output
json_response = client.chat.completions.create(
    model="deepseek-flash",
    messages=[
        {"role": "system", "content": "Return valid JSON with keys title and tags."},
        {"role": "user", "content": "Describe a secure API gateway."},
    ],
    response_format={"type": "json_object"},
)
print(json.loads(json_response.choices[0].message.content))

# Streaming
stream = client.responses.create(
    model="deepseek-flash",
    instructions="Answer clearly and briefly.",
    input="What is exponential backoff?",
    stream=True,
)
for event in stream:
    if event.type == "response.output_text.delta":
        print(event.delta, end="", flush=True)
print()

# Uploaded image file for vision input
with open("image.jpg", "rb") as image_file:
    uploaded = client.files.create(file=image_file, purpose="user_data")
vision_response = client.chat.completions.create(
    model="deepseek-flash",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "Describe this image."},
            {"type": "file", "file_id": uploaded.id},
        ],
    }],
)
print(vision_response.choices[0].message.content)
// Install: npm install openai
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.DEEPSEEK_API_KEY,
  baseURL: "https://api.deepseek.com",
});

const response = await client.chat.completions.create({
  model: "deepseek-flash",
  messages: [
    { role: "system", content: "You are a helpful assistant." },
    { role: "user", content: "Give me three names for a weather app." },
  ],
  thinking: { type: "disabled" },
});
console.log(response.choices[0].message.content);

// Streaming through the Responses API
const stream = await client.responses.create({
  model: "deepseek-flash",
  instructions: "Answer briefly and clearly.",
  input: "Explain why API clients use timeouts.",
  stream: true,
});
for await (const event of stream) {
  if (event.type === "response.output_text.delta") {
    process.stdout.write(event.delta);
  }
}
process.stdout.write("\n");

// Tool calling
const toolResponse = await client.chat.completions.create({
  model: "deepseek-flash",
  messages: [{ role: "user", content: "What is the weather in Boston?" }],
  tools: [{
    type: "function",
    function: {
      name: "get_weather",
      description: "Get current weather for a city.",
      parameters: {
        type: "object",
        properties: { city: { type: "string" } },
        required: ["city"],
      },
    },
  }],
});
const toolCall = toolResponse.choices[0].message.tool_calls?.[0];
if (toolCall) {
  console.log(JSON.parse(toolCall.function.arguments));
} else {
  console.log(toolResponse.choices[0].message.content);
}
<?php
$apiKey = getenv('DEEPSEEK_API_KEY');
if (!$apiKey) {
    throw new RuntimeException('DEEPSEEK_API_KEY is not set');
}

$url = 'https://api.deepseek.com/chat/completions';
$payload = [
    'model' => 'deepseek-flash',
    'messages' => [
        ['role' => 'system', 'content' => 'You are a helpful assistant.'],
        ['role' => 'user', 'content' => 'Explain what an API timeout is in one sentence.'],
    ],
    'thinking' => ['type' => 'disabled'],
    'stream' => false,
];

$ch = curl_init($url);
curl_setopt_array($ch, [
    CURLOPT_POST => true,
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_HTTPHEADER => [
        'Authorization: Bearer ' . $apiKey,
        'Content-Type: application/json',
    ],
    CURLOPT_POSTFIELDS => json_encode($payload, JSON_THROW_ON_ERROR),
    CURLOPT_TIMEOUT => 120,
]);

$body = curl_exec($ch);
if ($body === false) {
    $error = curl_error($ch);
    curl_close($ch);
    throw new RuntimeException('cURL error: ' . $error);
}
$status = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);

$data = json_decode($body, true, 512, JSON_THROW_ON_ERROR);
if ($status < 200 || $status >= 300) {
    $message = $data['error']['message'] ?? $body;
    throw new RuntimeException('DeepSeek API error (' . $status . '): ' . $message);
}

$content = $data['choices'][0]['message']['content'] ?? null;
if ($content === null) {
    throw new RuntimeException('No assistant content was returned');
}
echo $content . PHP_EOL;
?>
Policies

Data & usage

Data training

A clearly API-specific no-training default was not verified in the public documentation consulted. DeepSeek's public privacy policy states that DeepSeek group entities may process personal data for research and development, foundation-model training and optimization, analytics, and related purposes. Because the policy is framed around DeepSeek services and does not clearly distinguish hosted API data controls, organizations should review the applicable open-platform terms and avoid submitting sensitive data unless approved by their privacy and security teams.

Data retention

The public privacy policy states that retention varies according to the amount, type, sensitivity, purpose, and legal requirements associated with personal data, and that data may be stored and processed in China. For the Files API, uploaded images may be retained permanently when expiration fields are omitted or can be assigned an expiration between 3,600 and 2,592,000 seconds. A universal API prompt-retention period was not verified.

Rate limits

Documented account-level concurrency is 2,500 for deepseek-flash and 500 for deepseek-v4-pro. Exceeding concurrency returns HTTP 429. DeepSeek also documents request keep-alive behavior and a ten-minute limit for requests that have not begun inference.

Developer guide

DeepSeek API: Current Developer Platform, Models, Pricing, and Integration Guide

DeepSeek provides a usage-priced, OpenAI-compatible developer API at api.deepseek.com. Developers can use Chat Completions, the Responses API, an Anthropic-compatible Messages interface, the Files API, streaming, tool calling, JSON Output, and vision input through deepseek-flash. The current primary model identifiers are deepseek-flash and deepseek-v4-pro.
DeepSeek's developer platform lets applications send prompts, images, and tool instructions to hosted DeepSeek models over standard HTTP APIs. It is designed for developers who want OpenAI-compatible integration patterns, usage-based pricing, reasoning and coding capabilities, and access to vision input without adopting a provider-specific programming language. This guide covers access, model selection, requests, responses, pricing, capabilities, SDKs, limits, and production considerations.
DeepSeek's developer platform provides OpenAI-compatible Chat Completions and Responses APIs, an Anthropic-compatible interface, image uploads, streaming, tool calling, JSON Output, and usage-based pricing for DeepSeek Flash and V4 Pro.

What is the DeepSeek API and when should you use it?

The DeepSeek API is a hosted developer platform for adding DeepSeek model capabilities to applications, coding tools, agents, and automated workflows. Its primary endpoint is https://api.deepseek.com. The platform follows OpenAI-compatible conventions, so developers familiar with the OpenAI client libraries or standard JSON HTTP requests can usually adapt quickly.

The platform currently exposes Chat Completions, the Responses API, an Anthropic-compatible Messages interface, and a Files API. The API supports text generation, reasoning, coding, streaming responses, function or tool calls, JSON Output, and image understanding through deepseek-flash.

DeepSeek is a reasonable choice when you want competitive token pricing, OpenAI-compatible interfaces, strong reasoning and coding performance, Chinese-language capability, or vision input through the Flash model. It may be a poor fit when your application requires a verified API-specific no-training default, a universal prompt-retention guarantee, mature image or video generation, or a provider-independent data-processing location.

How to get access and create an API key

API access is separate from the free consumer DeepSeek Web and App services. Create an account in the DeepSeek Platform, create an API key, and add balance as required for API usage. Every API request must include the key as a Bearer token.

Keep the key on your server or in a secret manager. Do not place it in browser JavaScript, mobile application binaries, source repositories, or client-side logs. A typical request uses these headers:

Authorization: Bearer $DEEPSEEK_API_KEY
Content-Type: application/json

The platform also provides a developer playground at platform.deepseek.com. Use it to inspect account access and experiment before embedding requests in an application.

Which API and model should you choose?

For most new integrations, choose between the OpenAI-compatible Chat Completions API and Responses API. Chat Completions is the simpler option for conversational requests and existing OpenAI-style integrations. Responses is better suited to semantic streaming events, response items, and agent-style workflows.

DeepSeek also offers an Anthropic-compatible Messages interface at https://api.deepseek.com/anthropic. This can reduce migration work for applications already built around Anthropic-format requests.

ModelBest starting pointVisionDocumented concurrency
deepseek-flashCost-sensitive general workloads, coding, reasoning, and image inputYes2,500
deepseek-v4-proHigher-end reasoning or agent workloads where cost is less importantNo500

deepseek-flash is the current identifier for DeepSeek-V4.1-Flash. Older V4 Flash identifiers may remain accepted as retired aliases, but new applications should use the current identifier. deepseek-v4-pro is text-focused and does not currently support vision input.

Make your first request

The following example uses Chat Completions. Set the API key in the environment before running it:

export DEEPSEEK_API_KEY="sk-your-key"

curl --fail-with-body https://api.deepseek.com/chat/completions 
  -H "Authorization: Bearer ${DEEPSEEK_API_KEY}" 
  -H "Content-Type: application/json" 
  -d '{
    "model": "deepseek-flash",
    "messages": [
      {"role": "system", "content": "You are a concise technical assistant."},
      {"role": "user", "content": "Explain what an API timeout is in one sentence."}
    ],
    "thinking": {"type": "disabled"},
    "stream": false
  }'

A normal request sends a model name and an array of messages. The system message sets broad behavior, while the user message contains the task. The thinking setting controls the reasoning mode in this request; the example disables it for a short answer.

The same request with the OpenAI Python SDK

DeepSeek's documented quick starts use the OpenAI SDK configured with DeepSeek's base URL:

# Install: pip install -U openai
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["DEEPSEEK_API_KEY"],
    base_url="https://api.deepseek.com",
)

response = client.chat.completions.create(
    model="deepseek-flash",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Give me three names for a weather app."},
    ],
    thinking={"type": "disabled"},
)

print(response.choices[0].message.content)

How to understand the response

Chat Completions returns a response object containing one or more choices. For a basic non-streaming request, the generated text is normally available at response.choices[0].message.content. The response also includes usage information that can be used for cost monitoring.

When tools are requested, the assistant message may contain tool calls rather than ordinary text. Your application must inspect the call, validate its arguments, execute the corresponding local function or service, and send the tool result back to the model. Never execute generated arguments without validation.

The Responses API uses response items rather than only chat messages. Its streaming events include names such as response.output_text.delta, and function calls are represented as output items. This format is useful when an application needs semantic event handling or agent-oriented workflows.

How DeepSeek API pricing works

DeepSeek charges by token usage rather than by a fixed monthly developer plan. Prices are listed per one million tokens and distinguish cached input, uncached input, and output. Peak and off-peak rates also apply. The documented ranges are:

ModelCached inputUncached inputOutput
deepseek-flash$0.003–$0.006 per 1M tokens$0.15–$0.30 per 1M tokens$0.60–$1.20 per 1M tokens
deepseek-v4-pro$0.022–$0.044 per 1M tokens$0.66–$1.32 per 1M tokens$1.98–$3.96 per 1M tokens

Peak hours are listed as 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday. The lower end of each range represents off-peak pricing. Treat the official pricing page as authoritative because rates can change. DeepSeek charges against the account balance, using granted balance before topped-up balance when both are available.

What can developers build with the API?

Streaming output

Streaming sends generated output incrementally instead of waiting for the complete response. Chat Completions uses server-sent events and ends with data: [DONE]. The Responses API uses semantic events such as response.output_text.delta and a final completed, incomplete, or failed response event.

Streaming is useful for chat interfaces and long responses, but your application still needs connection timeouts, error handling, and a way to handle incomplete output.

Function and tool calling

Both listed models support tool calls. Define a function with its name, description, and parameter schema. The model can request the function, but your application—not DeepSeek—executes it. After execution, return the result to the model so it can produce a final answer.

Validate names, types, ranges, permissions, and authorization before executing a generated call. DeepSeek's documentation notes that generated arguments may not always conform perfectly to the declared schema.

JSON Output

DeepSeek supports JSON Output with response_format: {"type":"json_object"}. Tell the model explicitly to return JSON and describe the expected structure in the prompt:

json_response = client.chat.completions.create(
    model="deepseek-flash",
    messages=[
        {"role": "system", "content": "Return valid JSON with keys title and tags."},
        {"role": "user", "content": "Describe a secure API gateway."},
    ],
    response_format={"type": "json_object"},
)

print(json_response.choices[0].message.content)

This provides valid-JSON output, but it is not the same as a separately verified schema-constrained Structured Outputs feature. Parse and validate the returned JSON in your application.

Vision and uploaded files

deepseek-flash accepts images supplied through public URLs, base64 data URLs, or uploaded files. The Files API supports JPEG, PNG, GIF, and WebP images up to 64 MiB. Uploaded files are associated with the API key that created them.

Files may be retained permanently when no expiration is specified, or assigned an expiration between one hour and thirty days. Use expiration settings deliberately, especially when handling sensitive material. The public documentation does not establish a universal API prompt-retention period.

Advanced integration examples

Streaming with the Responses API

stream = client.responses.create(
    model="deepseek-flash",
    instructions="Answer clearly and briefly.",
    input="What is exponential backoff?",
    stream=True,
)

for event in stream:
    if event.type == "response.output_text.delta":
        print(event.delta, end="", flush=True)
print()

Use this approach when your application wants semantic response events rather than only incremental Chat Completions deltas.

Referencing an uploaded image

with open("image.jpg", "rb") as image_file:
    uploaded = client.files.create(file=image_file, purpose="user_data")

vision_response = client.chat.completions.create(
    model="deepseek-flash",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "Describe this image."},
            {"type": "file", "file_id": uploaded.id},
        ],
    }],
)

print(vision_response.choices[0].message.content)

SDKs and language support

Python and JavaScript or TypeScript applications can use the current OpenAI client libraries with DeepSeek's base URL. Anthropic-compatible integrations can use the Messages interface. PHP applications can call the API with ordinary HTTPS and cURL; no official DeepSeek PHP SDK was verified in the supplied documentation.

Limits and production considerations

Documented account-level concurrency is currently 2,500 for deepseek-flash and 500 for deepseek-v4-pro. Exceeding the applicable concurrency limit can return HTTP 429.

Other documented status codes include 400 for invalid formats, 401 for authentication failures, 402 for insufficient balance, 422 for invalid parameters, 500 for server errors, and 503 for overload. DeepSeek documents a ten-minute limit for requests that have not begun inference and keep-alive behavior for long-running requests.

Production clients should:

  • Set explicit connection and read timeouts.
  • Retry transient 429, 500, and 503 failures with exponential backoff.
  • Do not retry authentication or parameter errors without changing the request.
  • Monitor account balance, token usage, response errors, and concurrency.
  • Exclude API keys and sensitive prompts from logs.
  • Validate tool arguments and JSON responses before using them.
  • Review file expiration and deletion behavior for uploaded content.

Privacy and data handling

DeepSeek's public privacy materials state that data may be stored and processed in the People's Republic of China. They describe retention as dependent on factors such as data type, sensitivity, purpose, and legal requirements. The materials consulted do not provide a clearly documented API-specific no-training default or a universal API retention period.

Organizations should review the applicable platform terms and privacy documentation before sending personal, confidential, regulated, or proprietary information. If your requirements demand a specific data residency policy, contractual retention period, or confirmed no-training treatment for API prompts, verify those controls directly before deployment.

Advantages and limitations

Where DeepSeek API is a good fit

  • Applications that benefit from OpenAI-compatible Chat Completions or Responses interfaces.
  • Cost-sensitive workloads that can use the documented Flash pricing.
  • Coding, reasoning, mathematics, Chinese-language, and general text applications.
  • Agent workflows that need tool calls, function calls, or Responses API events.
  • Applications that need image understanding through deepseek-flash.
  • Teams comfortable using standard HTTP or compatible Python and JavaScript SDKs.

Where it may be a poor fit

  • Applications requiring image generation, video generation, or a broad consumer plugin ecosystem.
  • Workloads that require vision support from the Pro model.
  • Organizations that cannot accept processing or storage in China.
  • Systems requiring a clearly documented API-specific no-training default or universal retention guarantee.
  • Applications needing a separately verified schema-constrained Structured Outputs feature.
  • High-volume systems that cannot accommodate the documented concurrency limits or variable service load.
  1. Create a DeepSeek Platform account and API key.
  2. Start with deepseek-flash for ordinary text, coding, reasoning, or vision tests.
  3. Use Chat Completions for a simple request and Responses when semantic streaming or agent-style workflows are needed.
  4. Record token usage and test both peak and off-peak cost assumptions.
  5. Add timeouts, exponential-backoff retries, balance monitoring, and response validation before production.
  6. Review privacy, retention, and data-location requirements before sending real user or business data.
Sources 14