Developer platform

OpenAI API

Developer overview for OpenAI, including API access, pricing, SDK support, endpoints, capabilities and platform policies.

API key Required
Primary API Responses API
SDK support Official SDKs and client libraries are available for JavaScript/TypeScript, Python, .NET, Java, Go, and Ruby. The Agents SDK is officially available for TypeScript and Python. PHP clients are community-maintained, so raw HTTPS is the portable PHP approach
Rate limits Organization- and project-level limits vary by model and usage tier. Limits may include RPM, RPD, TPM, TPD, image-per-minute, audio, batch, vector-store, and long-context limits. Limits are visible in the developer console and response headers. Applicatio
Platform

API overview

Endpoints

API access

Base URL https://api.openai.com/v1
Primary API Responses API
Pricing

API pricing

Pricing model Usage-based pay-as-you-go pricing, primarily per million input, cached-input, cache-write, and output tokens, with separate charges for some tools and specialized services.

Current pricing varies by model and processing mode. As listed on the current pricing page, flagship examples include GPT-6 Astra at $10 per 1M short-context input tokens and $50 per 1M output tokens under Standard processing; GPT-6 Sol at $2 input and $1

Developer experience

SDKs & usability

SDKs Official SDKs and client libraries are available for JavaScript/TypeScript, Python, .NET, Java, Go, and Ruby. The Agents SDK is officially available for TypeScript and Python. PHP clients are community-maintained, so raw HTTPS is the portable PHP approach
Ease of use High for common use cases. The Responses API has a unified request format, official SDKs, automatic environment-variable authentication, direct HTTPS access, Playground testing, and helper properties such as output_text. Advanced multi-step tool workflows
Documentation High. Current official documentation includes a quickstart, API reference, guides for Responses, tools, structured outputs, files, rate limits, production, privacy, latency, SDKs, models, and migrations. Deprecated APIs are identified, although the rapidl
Latency No single universal latency guarantee is published for all models. Latency is mainly affected by model selection, generated-token count, input and output size, network time, tool calls, and sequential request count. Smaller models, shorter outputs, prompt
Features

API capabilities

✓ Streaming
✓ Function calling
✓ Assistants API
✓ File uploads
✓ Fine-tuning
✓ Web search
✓ Image input
✓ Structured outputs
✓ Playground
Feature notes

The Responses API is the recommended architecture for new direct API integrations. It accepts text, image, and file inputs and can return text or structured JSON. It supports streaming, function calling, built-in web search, file search, image generation, computer-use workflows, and other tools depending on model and account access. The Agents SDK for TypeScript and Python adds code-first orchestration, handoffs, guardrails, tracing, and tool execution. The legacy Assistants API was deprecated and scheduled to shut down on August 26, 2026; it should not be used for new integrations. Agent Builder is also being deprecated with a scheduled shutdown on November 30, 2026. Current model availability, supported tools, and limits vary by model, account, region, and project.

Examples

API examples

set -euo pipefail

: "${OPENAI_API_KEY:?Set OPENAI_API_KEY first}"

response=$(curl --fail-with-body --silent --show-error https://api.openai.com/v1/responses \
  -H "Authorization: Bearer ${OPENAI_API_KEY}" \
  -H "Content-Type: application/json" \
  -d @- <<'JSON'
{
  "model": "gpt-6-astra",
  "instructions": "You are a concise technical assistant.",
  "input": "Explain idempotency in one sentence."
}
JSON
)

printf '%s\n' "$response" | jq -r '.output_text // .error.message // "No text output returned"'

# Structured JSON output
curl --fail-with-body --silent --show-error https://api.openai.com/v1/responses \
  -H "Authorization: Bearer ${OPENAI_API_KEY}" \
  -H "Content-Type: application/json" \
  -d @- <<'JSON' | jq -r '.output_text // .error.message // "No text output returned"'
{
  "model": "gpt-6-astra",
  "input": "Classify this request: I need a refund for an accidental purchase.",
  "text": {
    "format": {
      "type": "json_schema",
      "name": "classification",
      "strict": true,
      "schema": {
        "type": "object",
        "properties": {
          "category": {"type": "string", "enum": ["billing", "technical", "general"]},
          "urgent": {"type": "boolean"}
        },
        "required": ["category", "urgent"],
        "additionalProperties": false
      }
    }
  }
}
JSON
# Install: pip install openai
import json
import os
from openai import OpenAI

client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])

response = client.responses.create(
    model="gpt-6-astra",
    instructions="You are a concise technical assistant.",
    input=[
        {"role": "user", "content": "Explain API retries in one sentence."},
        {"role": "assistant", "content": "Retries repeat a failed request according to a controlled policy."},
        {"role": "user", "content": "Now give one practical rule."}
    ],
)
print(response.output_text)

# Streaming
stream = client.responses.create(
    model="gpt-6-astra",
    input="List three benefits of server-sent event streaming.",
    stream=True,
)
for event in stream:
    if event.type == "response.output_text.delta":
        print(event.delta, end="", flush=True)
print()

# Structured output
structured = client.responses.create(
    model="gpt-6-astra",
    input="Classify: The payment was charged twice.",
    text={
        "format": {
            "type": "json_schema",
            "name": "classification",
            "strict": True,
            "schema": {
                "type": "object",
                "properties": {
                    "category": {"type": "string", "enum": ["billing", "technical", "general"]},
                    "urgent": {"type": "boolean"},
                },
                "required": ["category", "urgent"],
                "additionalProperties": False,
            },
        }
    },
)
print(json.loads(structured.output_text))

# Tool calling
weather_tool = {
    "type": "function",
    "name": "get_weather",
    "description": "Get the current weather for a city.",
    "parameters": {
        "type": "object",
        "properties": {"city": {"type": "string"}},
        "required": ["city"],
        "additionalProperties": False,
    },
}
first = client.responses.create(
    model="gpt-6-astra",
    input="What is the weather in Boston?",
    tools=[weather_tool],
)

follow_up = list(first.output)
for item in first.output:
    if item.type == "function_call" and item.name == "get_weather":
        args = json.loads(item.arguments)
        tool_result = {"city": args["city"], "temperature_f": 58, "condition": "cloudy"}
        follow_up.append({
            "type": "function_call_output",
            "call_id": item.call_id,
            "output": json.dumps(tool_result),
        })

if len(follow_up) > len(first.output):
    final = client.responses.create(
        model="gpt-6-astra",
        input=follow_up,
        tools=[weather_tool],
    )
    print(final.output_text)

try:
    client.responses.create(model="gpt-6-astra", input="Test request")
except Exception as exc:
    print(f"OpenAI API request failed: {exc}")
// Install: npm install openai
import OpenAI from "openai";

const client = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });

try {
  const response = await client.responses.create({
    model: "gpt-6-astra",
    instructions: "You are a concise technical assistant.",
    input: [
      { role: "user", content: "What is an API?" },
      { role: "assistant", content: "An API is an interface that lets software systems communicate." },
      { role: "user", content: "Give one example." }
    ]
  });
  console.log(response.output_text);

  const stream = await client.responses.create({
    model: "gpt-6-astra",
    input: "Explain streaming in two sentences.",
    stream: true
  });
  for await (const event of stream) {
    if (event.type === "response.output_text.delta") process.stdout.write(event.delta);
  }
  process.stdout.write("\n");

  const structured = await client.responses.create({
    model: "gpt-6-astra",
    input: "Classify: The application crashes after login.",
    text: {
      format: {
        type: "json_schema",
        name: "classification",
        strict: true,
        schema: {
          type: "object",
          properties: {
            category: { type: "string", enum: ["billing", "technical", "general"] },
            urgent: { type: "boolean" }
          },
          required: ["category", "urgent"],
          additionalProperties: false
        }
      }
    }
  });
  console.log(JSON.parse(structured.output_text));
} catch (error) {
  if (error instanceof OpenAI.APIError) {
    console.error(`OpenAI API error ${error.status}: ${error.message}`);
  } else {
    console.error(error);
  }
}
<?php
$apiKey = getenv('OPENAI_API_KEY');
if (!$apiKey) {
    fwrite(STDERR, "OPENAI_API_KEY is not set\n");
    exit(1);
}

$payload = [
    'model' => 'gpt-6-astra',
    'instructions' => 'You are a concise technical assistant.',
    'input' => 'Explain API retries in one sentence.'
];

$ch = curl_init('https://api.openai.com/v1/responses');
curl_setopt_array($ch, [
    CURLOPT_POST => true,
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_HTTPHEADER => [
        'Authorization: Bearer ' . $apiKey,
        'Content-Type: application/json'
    ],
    CURLOPT_POSTFIELDS => json_encode($payload, JSON_THROW_ON_ERROR),
    CURLOPT_TIMEOUT => 120
]);

$body = curl_exec($ch);
if ($body === false) {
    $message = curl_error($ch);
    curl_close($ch);
    throw new RuntimeException('Network error: ' . $message);
}

$status = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);
$data = json_decode($body, true, 512, JSON_THROW_ON_ERROR);

if ($status < 200 || $status >= 300) {
    $message = $data['error']['message'] ?? 'Unknown OpenAI API error';
    throw new RuntimeException("HTTP $status: $message");
}

echo $data['output_text'] ?? "No text output returned\n";

// Structured JSON output can be requested by adding a text.format object.
$structuredPayload = [
    'model' => 'gpt-6-astra',
    'input' => 'Classify: The payment was charged twice.',
    'text' => [
        'format' => [
            'type' => 'json_schema',
            'name' => 'classification',
            'strict' => true,
            'schema' => [
                'type' => 'object',
                'properties' => [
                    'category' => ['type' => 'string', 'enum' => ['billing', 'technical', 'general']],
                    'urgent' => ['type' => 'boolean']
                ],
                'required' => ['category', 'urgent'],
                'additionalProperties' => false
            ]
        ]
    ]
];

// Send $structuredPayload with the same cURL configuration when structured output is needed.
Policies

Data & usage

Data training

OpenAI states that data sent to the API is not used to train or improve OpenAI models by default unless the customer explicitly opts in. API customers remain responsible for complying with applicable laws, usage policies, and their own data-protection obligations.

Data retention

Abuse-monitoring logs are generally retained for up to 30 days unless longer retention is legally required or reasonably necessary for safety. Some features retain application state, including conversations, files, vector stores, and stored Responses data. Eligible organizations may request Modified Abuse Monitoring or Zero Data Retention, but approval is required and some endpoints or capabilities remain ineligible. The Responses API can retain application state when stored responses, conversations, background mode, or other stateful features are used.

Rate limits

Organization- and project-level limits vary by model and usage tier. Limits may include RPM, RPD, TPM, TPD, image-per-minute, audio, batch, vector-store, and long-context limits. Limits are visible in the developer console and response headers. Applicatio

Developer guide

OpenAI API Guide: Responses, Tools, Pricing, and Examples

OpenAI's current developer platform is centered on the Responses API, with official SDKs, hosted tools, structured outputs, multimodal inputs, file handling, conversation state, realtime capabilities, and the Agents SDK. This guide explains how beginners can obtain credentials, make a first request, understand responses, estimate usage-based costs, and evaluate production considerations.
The OpenAI API lets developers add model-based text, vision, audio, image, video, search, retrieval, moderation, and agent features to their own applications. For new direct model integrations, OpenAI recommends the Responses API. It provides one request pattern for text generation, streaming, function calling, structured JSON output, image and file inputs, hosted tools, and conversation continuation. This guide covers the practical path from creating an API key to deploying a reliable integration.
The OpenAI API uses the Responses API for new direct model integrations and supports text, multimodal inputs, streaming, tools, structured outputs, files, conversation state, and agent runtimes. Access requires a server-side API key, pricing is usage-based, and production systems must account for rate limits, privacy, retention, model availability, and tool security.

What is the OpenAI API, and when should you use it?

The OpenAI API is a hosted developer platform. Your application sends an HTTPS request containing a model, instructions, and input; OpenAI processes the request and returns generated output or other response items. Your application can then display the result, store it, or use it to trigger additional code.

The recommended starting point for new direct model integrations is the Responses API at POST /v1/responses. It supports ordinary text generation as well as streaming, image and file inputs, hosted tools, custom function calling, structured outputs, and conversation continuation. Chat Completions remains available for compatible existing integrations, but new applications should generally begin with Responses.

The API is intended for server-side applications, backends, automation systems, search and retrieval experiences, customer-support tools, document workflows, and agentic applications. It is less suitable when you need a completely offline model, fixed infrastructure with no hosted service dependency, or a cost that cannot vary with usage.

Who is it for?

It can serve a beginner building a small server endpoint as well as an intermediate developer implementing tools, document processing, stateful conversations, or agent orchestration. Basic requests require only an API key, an available model, and an SDK or HTTPS client. More advanced applications must also account for tool execution, retries, rate limits, privacy settings, response storage, and model-specific pricing.

How do you get API access?

  1. Create an account and project in the OpenAI developer platform.
  2. Create an API key using the platform's key-management tools.
  3. Store the key in a server-side environment variable such as OPENAI_API_KEY.
  4. Install an official SDK or send HTTPS requests to the API base URL.
  5. Monitor usage, billing, errors, request IDs, and rate-limit headers from the beginning.

The standard base URL is https://api.openai.com/v1. Requests authenticate with a bearer token in the Authorization header. Never place an API key in browser JavaScript, a mobile application, a public repository, or another client-side bundle. A safer design sends requests from your server, which keeps the credential outside the user's device.

Dashboard and Playground

The developer dashboard provides API-key management, projects, usage monitoring, billing controls, the Chat Playground, Agent Builder, and related development tools. The Playground uses the same API infrastructure, so its requests count toward API usage and pricing. It is useful for testing prompts and configurations before moving the request into application code.

Which API and model should you choose?

For a new direct model request, start with the Responses API. Choose a currently available model based on the task, required modality, latency, output quality, context needs, and budget. Model availability and exact pricing can change, so confirm the selected model in the current model and pricing documentation before deploying it.

RequirementRelevant API capability
Ordinary text or reasoning responseResponses API
Incremental output while generation is in progressResponses streaming
Application code must perform an actionFunction calling or a hosted tool
Answers must follow a defined JSON schemaStructured Outputs
Documents or images must be suppliedImage and file inputs, Files API, or file search
Durable multi-session state or orchestrationConversations API, Agents SDK, or current Agents APIs

The Responses API is the primary interface for a single model interaction. For lightweight continuation, pass previous_response_id. For durable state shared across sessions, devices, workers, or jobs, use the Conversations API. For managed agent workflows involving tools, handoffs, guardrails, sessions, tracing, or sandbox behavior, use the Agents SDK or current Agents APIs rather than the retired Assistants API.

Making your first request

Install the official Python SDK with pip install openai, set OPENAI_API_KEY in the server environment, and make a Responses request. The model name below is an example from the supplied current platform information; verify that it is available to your account before running the example.

import os
from openai import OpenAI

client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])

response = client.responses.create(
    model="gpt-6-astra",
    instructions="You are a concise technical assistant.",
    input="Explain idempotent API retries in two sentences."
)

print(response.output_text)

The same operation can be sent directly over HTTPS:

curl -sS https://api.openai.com/v1/responses 
  -H "Authorization: Bearer $OPENAI_API_KEY" 
  -H "Content-Type: application/json" 
  -d '{
    "model": "gpt-6-astra",
    "instructions": "You are a concise technical assistant.",
    "input": "Explain idempotent API retries in two sentences."
  }'

In production, check the HTTP status, handle structured API errors, log the request ID, and avoid exposing the complete response or sensitive input in logs unless your data-handling policy permits it.

How to understand a response

A Responses API result is not limited to one plain text string. It can contain output items such as messages, reasoning-related items, tool calls, and tool results. The official SDK provides response.output_text as a convenient way to extract generated text when text is present.

When the response requests a function call, your application—not the model—executes the function. Your code reads the function name and arguments, validates them, performs the operation, and sends the result back with the associated call identifier. The model can then produce a final user-facing response. Never treat model-generated arguments as automatically trusted input; validate permissions, types, ranges, and the requested operation in application code.

Continuing a conversation

For lightweight server-managed continuation, send the earlier response's identifier as previous_response_id:

second = client.responses.create(
    model="gpt-6-astra",
    previous_response_id=response.id,
    input="Now rewrite those tips as a checklist."
)

print(second.output_text)

Responses are stored for 30 days by default unless storage is disabled. Conversation objects and their items have separate persistence behavior, so select the state mechanism deliberately when data must survive longer or be shared across application components.

How OpenAI API pricing works

OpenAI API pricing is primarily usage-based rather than a single fixed monthly API subscription. Model inference is generally charged according to input tokens, cached input tokens where applicable, and output tokens. Audio, images, video, realtime sessions, and other modalities can use different billing units.

Hosted capabilities can add separate charges. Examples include web-search calls, file-search storage and tool calls, hosted containers or code-interpreter sessions, image generation, video generation, and realtime audio. Batch, Flex, Fast, regional-processing, Scale Tier, and Reserved Tier options can alter price or performance characteristics.

There is no universal request price to apply across every model and feature. Pricing changes frequently, and the current official pricing page should be checked before cost modeling or procurement. Track both input and output usage in your application, and remember that retries, long conversation history, tool calls, file processing, and multimodal inputs can increase total consumption.

Capabilities available through the API

Streaming output

Streaming uses server-sent events to deliver output incrementally. It improves perceived latency because the application can display partial output before the complete response is ready. The SDK emits events such as response.output_text.delta for text fragments.

stream = client.responses.create(
    model="gpt-6-astra",
    input="Explain server-sent events briefly.",
    stream=True
)

for event in stream:
    if getattr(event, "type", "") == "response.output_text.delta":
        print(event.delta, end="", flush=True)
print()

Streaming does not remove the need to handle errors or incomplete output. Your user interface should account for disconnects, cancellation, moderation or policy handling, and the possibility that a response contains non-text events.

Function calling and hosted tools

Function calling allows the model to request an operation described by your application, such as looking up an order or calculating a shipping estimate. Your server defines the function schema, receives the call, executes trusted application code, and submits the result. The model does not directly gain access to your database or business systems.

OpenAI also provides hosted tools, including web search, file search, code interpreter, image generation, computer-use-related capabilities, remote MCP, shell, and tool search. Availability depends on the model, endpoint, account, and tool-specific restrictions. Treat tools as separate cost, security, and reliability boundaries.

Structured Outputs

Structured Outputs constrains a response to a supplied JSON Schema for supported models and schemas. In the Responses API, configure the text format with a strict schema:

schema = {
    "type": "object",
    "properties": {
        "answer": {"type": "string"},
        "confidence": {"type": "number"}
    },
    "required": ["answer", "confidence"],
    "additionalProperties": False
}

structured = client.responses.create(
    model="gpt-6-astra",
    input="State whether HTTPS encrypts data in transit.",
    text={
        "format": {
            "type": "json_schema",
            "name": "answer",
            "strict": True,
            "schema": schema
        }
    }
)

print(structured.output_text)

Structured Outputs is different from older JSON mode. JSON mode aims to produce valid JSON, while Structured Outputs is intended to enforce the supplied schema within the supported subset of JSON Schema. Your application should still validate the returned data and handle refusal, truncation, and transport errors.

Files, images, and multimodal input

Compatible models can accept image URLs, uploaded images, file identifiers, and document inputs. The Files API supports uploading, retrieving, downloading, deleting, setting expiration policies, and handling files for specific purposes. Files can also be used in file search, batch processing, evaluations, and other workflows.

Image and file processing has additional privacy, safety, availability, and pricing considerations. Confirm that the chosen model and endpoint support the input type, and avoid uploading sensitive material unless your retention and contractual requirements are satisfied.

Agents, audio, and realtime applications

The wider platform includes APIs and tools for realtime communication, audio, image generation, video generation, embeddings, moderation, files, batches, evaluations, and agent workflows. The Agents SDK adds orchestration features such as tools, handoffs, agents-as-tools, guardrails, sessions, tracing, streamed runs, voice agents, realtime agents, and sandbox agents. It uses the Responses API by default for OpenAI model calls while providing a higher-level runtime.

SDKs and supported languages

Official SDKs and libraries are available for Python, JavaScript and TypeScript, Go, Java, .NET, and Ruby. The official Python and JavaScript SDKs expose the Responses API directly. PHP applications can use the REST API with cURL or another HTTP client; a provider-specific PHP SDK is not required.

Use one current SDK generation consistently. Keep SDK versions controlled, test model and feature availability in a non-production project, and consult the current API reference when a request uses tools, structured outputs, streaming, files, or state.

Advanced example: declaring a function

The following JavaScript pattern declares a function. It demonstrates the declaration step; production code must inspect returned function-call items, validate the arguments, execute the function, and submit the result in a follow-up request.

import OpenAI from "openai";

const client = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });

const response = await client.responses.create({
  model: "gpt-6-astra",
  input: "What is the weather in Paris?",
  tools: [{
    type: "function",
    name: "get_weather",
    description: "Get the current weather for a city.",
    parameters: {
      type: "object",
      properties: { city: { type: "string" } },
      required: ["city"],
      additionalProperties: false
    },
    strict: true
  }]
});

console.log(response.output);

In a real application, the function should call an approved weather service, enforce authorization and input validation, and return a bounded result. Do not allow a model to choose arbitrary URLs, shell commands, database queries, or privileged operations without an explicit security layer.

Limits and production considerations

Rate limits, retries, and latency

Limits vary by organization, project, model, endpoint, usage tier, and processing mode. Responses expose request and token headers such as x-ratelimit-limit-requests, x-ratelimit-remaining-requests, and x-ratelimit-reset-requests, along with token equivalents.

Production clients should inspect these headers, respect retry-after when supplied, use bounded exponential backoff for retryable failures, and log request IDs. Retries must be designed carefully: repeating a read-only generation may be acceptable, but repeating a tool call that charges a customer or changes state can create duplicate side effects. Use idempotency controls in your own application where appropriate.

Latency depends on the model, prompt and output size, reasoning effort, tool use, queueing, network conditions, and processing tier. Streaming improves time to first visible output, but it does not necessarily reduce the total generation time. Fast and enterprise capacity options may improve performance for eligible accounts.

Privacy and data retention

OpenAI states that API data is not used to train or improve its models by default unless a customer explicitly opts in. Abuse-monitoring logs are generally retained for up to 30 days by default. Responses are stored for 30 days by default and can be created with storage disabled; files and conversations have their own persistence behavior.

Eligible organizations may request controls such as Zero Data Retention or Modified Abuse Monitoring, but endpoint, feature, geographic, and contractual restrictions apply. Image and file inputs may receive specialized safety scanning. Review the applicable data controls before sending personal, confidential, regulated, or customer-owned information.

Legacy features and availability caveats

The legacy Assistants API should not be used for new integrations. The recommended migration paths are the Responses API, Conversations API, Agents SDK, or current Agents APIs. Fine-tuning remains documented for eligible existing users, but current pricing information describes the fine-tuning platform as being wound down and no longer accessible to new users. Do not make a new project dependent on fine-tuning without confirming account eligibility.

Individual models, tools, regions, modalities, processing tiers, and schemas can have separate restrictions. Confirm availability and exact behavior against the current documentation rather than assuming that every capability works with every model.

When is the OpenAI API a good or poor choice?

It is a good choice when:

  • You want a hosted API with official SDKs and a unified current interface.
  • Your application needs text plus image, file, audio, search, or tool capabilities.
  • You want streaming, function calling, structured JSON output, or managed conversation state.
  • You prefer usage-based billing and do not want to operate model infrastructure yourself.
  • You need an upgrade path from a basic model request to tools or agent orchestration.

It may be a poor choice when:

  • Your application must run entirely offline or inside infrastructure that cannot call an external service.
  • Variable usage-based costs are unacceptable or difficult to monitor.
  • Your requirements depend on a model, tool, region, retention policy, or fine-tuning workflow that your account does not support.
  • Your system cannot tolerate network latency, provider outages, changing model availability, or endpoint evolution.
  • You cannot establish an acceptable privacy and data-retention arrangement for the information being processed.

For a first implementation, begin with one server-side Responses request, add usage and error logging, and then introduce streaming, structured outputs, files, tools, or Agents SDK orchestration only when the application needs them. This keeps the initial integration understandable while leaving room for more advanced workflows.

Sources 18