What is the OpenAI API, and when should you use it?
The OpenAI API is a hosted developer platform. Your application sends an HTTPS request containing a model, instructions, and input; OpenAI processes the request and returns generated output or other response items. Your application can then display the result, store it, or use it to trigger additional code.
The recommended starting point for new direct model integrations is the Responses API at POST /v1/responses. It supports ordinary text generation as well as streaming, image and file inputs, hosted tools, custom function calling, structured outputs, and conversation continuation. Chat Completions remains available for compatible existing integrations, but new applications should generally begin with Responses.
The API is intended for server-side applications, backends, automation systems, search and retrieval experiences, customer-support tools, document workflows, and agentic applications. It is less suitable when you need a completely offline model, fixed infrastructure with no hosted service dependency, or a cost that cannot vary with usage.
Who is it for?
It can serve a beginner building a small server endpoint as well as an intermediate developer implementing tools, document processing, stateful conversations, or agent orchestration. Basic requests require only an API key, an available model, and an SDK or HTTPS client. More advanced applications must also account for tool execution, retries, rate limits, privacy settings, response storage, and model-specific pricing.
How do you get API access?
- Create an account and project in the OpenAI developer platform.
- Create an API key using the platform's key-management tools.
- Store the key in a server-side environment variable such as
OPENAI_API_KEY. - Install an official SDK or send HTTPS requests to the API base URL.
- Monitor usage, billing, errors, request IDs, and rate-limit headers from the beginning.
The standard base URL is https://api.openai.com/v1. Requests authenticate with a bearer token in the Authorization header. Never place an API key in browser JavaScript, a mobile application, a public repository, or another client-side bundle. A safer design sends requests from your server, which keeps the credential outside the user's device.
Dashboard and Playground
The developer dashboard provides API-key management, projects, usage monitoring, billing controls, the Chat Playground, Agent Builder, and related development tools. The Playground uses the same API infrastructure, so its requests count toward API usage and pricing. It is useful for testing prompts and configurations before moving the request into application code.
Which API and model should you choose?
For a new direct model request, start with the Responses API. Choose a currently available model based on the task, required modality, latency, output quality, context needs, and budget. Model availability and exact pricing can change, so confirm the selected model in the current model and pricing documentation before deploying it.
| Requirement | Relevant API capability |
|---|---|
| Ordinary text or reasoning response | Responses API |
| Incremental output while generation is in progress | Responses streaming |
| Application code must perform an action | Function calling or a hosted tool |
| Answers must follow a defined JSON schema | Structured Outputs |
| Documents or images must be supplied | Image and file inputs, Files API, or file search |
| Durable multi-session state or orchestration | Conversations API, Agents SDK, or current Agents APIs |
The Responses API is the primary interface for a single model interaction. For lightweight continuation, pass previous_response_id. For durable state shared across sessions, devices, workers, or jobs, use the Conversations API. For managed agent workflows involving tools, handoffs, guardrails, sessions, tracing, or sandbox behavior, use the Agents SDK or current Agents APIs rather than the retired Assistants API.
Making your first request
Install the official Python SDK with pip install openai, set OPENAI_API_KEY in the server environment, and make a Responses request. The model name below is an example from the supplied current platform information; verify that it is available to your account before running the example.
import os
from openai import OpenAI
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
response = client.responses.create(
model="gpt-6-astra",
instructions="You are a concise technical assistant.",
input="Explain idempotent API retries in two sentences."
)
print(response.output_text)The same operation can be sent directly over HTTPS:
curl -sS https://api.openai.com/v1/responses
-H "Authorization: Bearer $OPENAI_API_KEY"
-H "Content-Type: application/json"
-d '{
"model": "gpt-6-astra",
"instructions": "You are a concise technical assistant.",
"input": "Explain idempotent API retries in two sentences."
}'In production, check the HTTP status, handle structured API errors, log the request ID, and avoid exposing the complete response or sensitive input in logs unless your data-handling policy permits it.
How to understand a response
A Responses API result is not limited to one plain text string. It can contain output items such as messages, reasoning-related items, tool calls, and tool results. The official SDK provides response.output_text as a convenient way to extract generated text when text is present.
When the response requests a function call, your application—not the model—executes the function. Your code reads the function name and arguments, validates them, performs the operation, and sends the result back with the associated call identifier. The model can then produce a final user-facing response. Never treat model-generated arguments as automatically trusted input; validate permissions, types, ranges, and the requested operation in application code.
Continuing a conversation
For lightweight server-managed continuation, send the earlier response's identifier as previous_response_id:
second = client.responses.create(
model="gpt-6-astra",
previous_response_id=response.id,
input="Now rewrite those tips as a checklist."
)
print(second.output_text)Responses are stored for 30 days by default unless storage is disabled. Conversation objects and their items have separate persistence behavior, so select the state mechanism deliberately when data must survive longer or be shared across application components.
How OpenAI API pricing works
OpenAI API pricing is primarily usage-based rather than a single fixed monthly API subscription. Model inference is generally charged according to input tokens, cached input tokens where applicable, and output tokens. Audio, images, video, realtime sessions, and other modalities can use different billing units.
Hosted capabilities can add separate charges. Examples include web-search calls, file-search storage and tool calls, hosted containers or code-interpreter sessions, image generation, video generation, and realtime audio. Batch, Flex, Fast, regional-processing, Scale Tier, and Reserved Tier options can alter price or performance characteristics.
There is no universal request price to apply across every model and feature. Pricing changes frequently, and the current official pricing page should be checked before cost modeling or procurement. Track both input and output usage in your application, and remember that retries, long conversation history, tool calls, file processing, and multimodal inputs can increase total consumption.
Capabilities available through the API
Streaming output
Streaming uses server-sent events to deliver output incrementally. It improves perceived latency because the application can display partial output before the complete response is ready. The SDK emits events such as response.output_text.delta for text fragments.
stream = client.responses.create(
model="gpt-6-astra",
input="Explain server-sent events briefly.",
stream=True
)
for event in stream:
if getattr(event, "type", "") == "response.output_text.delta":
print(event.delta, end="", flush=True)
print()Streaming does not remove the need to handle errors or incomplete output. Your user interface should account for disconnects, cancellation, moderation or policy handling, and the possibility that a response contains non-text events.
Function calling and hosted tools
Function calling allows the model to request an operation described by your application, such as looking up an order or calculating a shipping estimate. Your server defines the function schema, receives the call, executes trusted application code, and submits the result. The model does not directly gain access to your database or business systems.
OpenAI also provides hosted tools, including web search, file search, code interpreter, image generation, computer-use-related capabilities, remote MCP, shell, and tool search. Availability depends on the model, endpoint, account, and tool-specific restrictions. Treat tools as separate cost, security, and reliability boundaries.
Structured Outputs
Structured Outputs constrains a response to a supplied JSON Schema for supported models and schemas. In the Responses API, configure the text format with a strict schema:
schema = {
"type": "object",
"properties": {
"answer": {"type": "string"},
"confidence": {"type": "number"}
},
"required": ["answer", "confidence"],
"additionalProperties": False
}
structured = client.responses.create(
model="gpt-6-astra",
input="State whether HTTPS encrypts data in transit.",
text={
"format": {
"type": "json_schema",
"name": "answer",
"strict": True,
"schema": schema
}
}
)
print(structured.output_text)Structured Outputs is different from older JSON mode. JSON mode aims to produce valid JSON, while Structured Outputs is intended to enforce the supplied schema within the supported subset of JSON Schema. Your application should still validate the returned data and handle refusal, truncation, and transport errors.
Files, images, and multimodal input
Compatible models can accept image URLs, uploaded images, file identifiers, and document inputs. The Files API supports uploading, retrieving, downloading, deleting, setting expiration policies, and handling files for specific purposes. Files can also be used in file search, batch processing, evaluations, and other workflows.
Image and file processing has additional privacy, safety, availability, and pricing considerations. Confirm that the chosen model and endpoint support the input type, and avoid uploading sensitive material unless your retention and contractual requirements are satisfied.
Agents, audio, and realtime applications
The wider platform includes APIs and tools for realtime communication, audio, image generation, video generation, embeddings, moderation, files, batches, evaluations, and agent workflows. The Agents SDK adds orchestration features such as tools, handoffs, agents-as-tools, guardrails, sessions, tracing, streamed runs, voice agents, realtime agents, and sandbox agents. It uses the Responses API by default for OpenAI model calls while providing a higher-level runtime.
SDKs and supported languages
Official SDKs and libraries are available for Python, JavaScript and TypeScript, Go, Java, .NET, and Ruby. The official Python and JavaScript SDKs expose the Responses API directly. PHP applications can use the REST API with cURL or another HTTP client; a provider-specific PHP SDK is not required.
Use one current SDK generation consistently. Keep SDK versions controlled, test model and feature availability in a non-production project, and consult the current API reference when a request uses tools, structured outputs, streaming, files, or state.
Advanced example: declaring a function
The following JavaScript pattern declares a function. It demonstrates the declaration step; production code must inspect returned function-call items, validate the arguments, execute the function, and submit the result in a follow-up request.
import OpenAI from "openai";
const client = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });
const response = await client.responses.create({
model: "gpt-6-astra",
input: "What is the weather in Paris?",
tools: [{
type: "function",
name: "get_weather",
description: "Get the current weather for a city.",
parameters: {
type: "object",
properties: { city: { type: "string" } },
required: ["city"],
additionalProperties: false
},
strict: true
}]
});
console.log(response.output);In a real application, the function should call an approved weather service, enforce authorization and input validation, and return a bounded result. Do not allow a model to choose arbitrary URLs, shell commands, database queries, or privileged operations without an explicit security layer.
Limits and production considerations
Rate limits, retries, and latency
Limits vary by organization, project, model, endpoint, usage tier, and processing mode. Responses expose request and token headers such as x-ratelimit-limit-requests, x-ratelimit-remaining-requests, and x-ratelimit-reset-requests, along with token equivalents.
Production clients should inspect these headers, respect retry-after when supplied, use bounded exponential backoff for retryable failures, and log request IDs. Retries must be designed carefully: repeating a read-only generation may be acceptable, but repeating a tool call that charges a customer or changes state can create duplicate side effects. Use idempotency controls in your own application where appropriate.
Latency depends on the model, prompt and output size, reasoning effort, tool use, queueing, network conditions, and processing tier. Streaming improves time to first visible output, but it does not necessarily reduce the total generation time. Fast and enterprise capacity options may improve performance for eligible accounts.
Privacy and data retention
OpenAI states that API data is not used to train or improve its models by default unless a customer explicitly opts in. Abuse-monitoring logs are generally retained for up to 30 days by default. Responses are stored for 30 days by default and can be created with storage disabled; files and conversations have their own persistence behavior.
Eligible organizations may request controls such as Zero Data Retention or Modified Abuse Monitoring, but endpoint, feature, geographic, and contractual restrictions apply. Image and file inputs may receive specialized safety scanning. Review the applicable data controls before sending personal, confidential, regulated, or customer-owned information.
Legacy features and availability caveats
The legacy Assistants API should not be used for new integrations. The recommended migration paths are the Responses API, Conversations API, Agents SDK, or current Agents APIs. Fine-tuning remains documented for eligible existing users, but current pricing information describes the fine-tuning platform as being wound down and no longer accessible to new users. Do not make a new project dependent on fine-tuning without confirming account eligibility.
Individual models, tools, regions, modalities, processing tiers, and schemas can have separate restrictions. Confirm availability and exact behavior against the current documentation rather than assuming that every capability works with every model.
When is the OpenAI API a good or poor choice?
It is a good choice when:
- You want a hosted API with official SDKs and a unified current interface.
- Your application needs text plus image, file, audio, search, or tool capabilities.
- You want streaming, function calling, structured JSON output, or managed conversation state.
- You prefer usage-based billing and do not want to operate model infrastructure yourself.
- You need an upgrade path from a basic model request to tools or agent orchestration.
It may be a poor choice when:
- Your application must run entirely offline or inside infrastructure that cannot call an external service.
- Variable usage-based costs are unacceptable or difficult to monitor.
- Your requirements depend on a model, tool, region, retention policy, or fine-tuning workflow that your account does not support.
- Your system cannot tolerate network latency, provider outages, changing model availability, or endpoint evolution.
- You cannot establish an acceptable privacy and data-retention arrangement for the information being processed.
For a first implementation, begin with one server-side Responses request, add usage and error logging, and then introduce streaming, structured outputs, files, tools, or Agents SDK orchestration only when the application needs them. This keeps the initial integration understandable while leaving room for more advanced workflows.
