What is the StepFun API and when should you use it?
StepFun Open Platform is the developer-facing API for integrating StepFun models into applications. It is intended for developers building chat interfaces, document and image analysis, coding tools, search-assisted applications, agent workflows, voice experiences, and other software that needs model-generated responses.
The platform follows OpenAI-compatible conventions. That means developers familiar with the OpenAI API can usually reuse the same general client pattern while changing the API key, base URL, and model name. StepFun documents both the newer Responses API and the Chat Completions API. Responses is the primary choice for new applications using supported models such as Step 5 Preview and Step 3.7 Flash; Chat Completions remains available for broader model compatibility.
This is a usage-based developer service rather than a fixed-price consumer subscription. You need a platform account and API key, and your costs depend on the model and the resources used.
Who is it for?
- Developers who want an OpenAI-compatible interface and SDK workflow.
- Teams building Chinese-language or multimodal applications.
- Applications that need text, vision, audio, image, or realtime voice models.
- Products requiring streaming, function calling, file input, web search, or structured JSON responses.
- Developers evaluating model-specific pricing rather than purchasing a single unlimited API plan.
StepFun may be a poor fit when you require globally uniform availability, independently verified enterprise privacy guarantees, or API features that are documented consistently across all regions and models.
How to get access and create API credentials
API access requires a StepFun Open Platform account and an API key. The supplied platform research confirms that API keys are required, but does not verify a single universal account-creation or key-generation workflow. Use the current platform console and documentation for the account and key process.
Individual accounts require identity verification for the published rate-limit tiers. Keep the key on your server or in a secret manager; do not place it in browser JavaScript, a mobile application, or source control. A typical environment-variable setup is:
export STEPFUN_API_KEY="your_api_key"The international API base URL supplied by the platform is https://api.stepfun.ai/v1. The documentation and console are hosted separately at platform.stepfun.ai.
Which API and model should you choose?
| Need | Starting point | Important distinction |
|---|---|---|
| New application using a supported newer model | Responses API | Recommended for Step 5 Preview and Step 3.7 Flash according to the supplied research. |
| Broad compatibility with existing integrations | Chat Completions API | Still supported, but model-specific features may differ. |
| Text and reasoning | Step 5 Preview or Step 3.7 Flash | Check the current model documentation before relying on a particular feature. |
| Lower-cost general model usage | Step 3.5 Flash | Its published international token prices are lower than the listed Step 5 Preview and Step 3.7 Flash prices. |
| Realtime voice | Supported realtime models | Uses a WebSocket-based realtime interface rather than an ordinary synchronous request. |
Model capabilities are not identical. Confirm the current model page before combining reasoning, vision, audio, tools, structured output, or realtime features in production. StepFun's current platform materials highlight Step 5 Preview, Step 3.7 Flash, Step 3.5 Flash, and StepAudio 3.
Making your first request
The following request uses the Responses API through raw HTTP. It keeps the example simple: send text, select a model, and read the returned response object.
curl https://api.stepfun.ai/v1/responses
-H "Authorization: Bearer $STEPFUN_API_KEY"
-H "Content-Type: application/json"
-d '{
"model": "step-3.7-flash",
"input": "Explain what an API key is in two sentences."
}'Use the exact model identifier shown in the current StepFun model documentation. The research identifies model families and commercial names, but does not provide a guaranteed identifier spelling for every model.
The same API can be accessed with the OpenAI Python SDK by changing the base URL and supplying the StepFun key:
from openai import OpenAI
client = OpenAI(
api_key="your_api_key",
base_url="https://api.stepfun.ai/v1",
)
response = client.responses.create(
model="step-3.7-flash",
input="Explain what an API key is in two sentences.",
)
print(response.output_text)StepFun officially documents OpenAI-compatible Python and JavaScript SDK usage. Raw HTTP is available for other languages. The research does not verify a separate first-party StepFun Python, JavaScript, or PHP SDK requirement.
Understanding the response
A Responses API result is a structured response object rather than just a bare text string. It can contain generated output items, usage information, tool calls, and other metadata. SDKs may provide a convenience property such as output_text for ordinary text output, while applications that use tools or multimodal results should inspect the structured output items.
For production code, handle at least these cases:
- Normal text output displayed to the user or passed to another application component.
- Streaming events when the request is streamed.
- Function or tool calls that require your application to execute code and return the result.
- Errors caused by invalid credentials, unsupported model features, malformed input, or rate limits.
- Usage data used for billing and monitoring.
How StepFun API pricing works
StepFun uses usage-based pricing. Depending on the service, billing can be based on input tokens, output tokens, cached input tokens, characters, duration, images, search calls, or file storage. Step Plan subscription access is separate from API usage billing.
The supplied research lists these current international token prices:
| Model | Cache-miss input | Cache-hit input | Output |
|---|---|---|---|
| Step 5 Preview | $1.00 per million tokens | $0.05 per million tokens | $2.70 per million tokens |
| Step 3.7 Flash | $0.20 per million tokens | $0.04 per million tokens | $1.15 per million tokens |
| Step 3.5 Flash | $0.10 per million tokens | $0.02 per million tokens | $0.30 per million tokens |
Other listed prices include $0.006 per internet search call, $0.02 per image search call, and $0.08 per GB per day for file storage. These prices and model availability can change, so use the current pricing documentation before estimating a production budget.
To estimate a real request, account for both input and output tokens, repeated prompt prefixes that may qualify as cached input, tool or search calls, and the storage lifetime of uploaded files. Do not confuse API charges with consumer or Step Plan subscription prices.
Important API capabilities
Multimodal input
StepFun provides models and services covering language, vision, audio, image generation, and realtime voice. The platform supports image input, and the wider ecosystem includes file and media processing. Capability support is model-specific, so an application should select a model documented for the input type it sends rather than assuming every model accepts every modality.
Streaming output
Streaming lets an application receive partial output while the model is still generating instead of waiting for the complete response. StepFun supports streaming through server-sent events for ordinary API interactions. It also provides a WebSocket interface for realtime use cases.
Streaming is useful for chat interfaces and long responses, but the client must assemble incremental events, handle interrupted connections, and avoid treating an incomplete stream as a completed answer.
Function and tool calling
Function calling lets the model request an operation described by your application, such as looking up an order or querying an internal database. The model does not execute your function automatically. Your server must validate the arguments, perform the operation, and send the result back to the model.
A simplified Responses request can describe a function like this:
{
"model": "step-3.7-flash",
"input": "Find the status of order 1842.",
"tools": [
{
"type": "function",
"name": "get_order_status",
"description": "Look up an order by its identifier",
"parameters": {
"type": "object",
"properties": {
"order_id": { "type": "string" }
},
"required": ["order_id"],
"additionalProperties": false
}
}
]
}Tool schemas should be narrow and validated. Treat tool arguments as untrusted input, enforce authorization in application code, and never allow a model-generated request to bypass business rules.
Files and document workflows
The platform exposes file APIs and supports file upload. Uploaded content can be used in supported file-analysis or multimodal workflows, while storage is billed at the listed rate of $0.08 per GB per day. The exact accepted file types, request parameters, retention behavior, and model compatibility should be checked in the current file API documentation.
Structured JSON outputs
Structured outputs are supported, including JSON Schema-based output. This is useful when the result must be consumed by software rather than displayed directly to a person, such as extracting fields from a document or returning a classification object.
Even with a schema, validate the returned data in your application. A schema can constrain the shape of a response, but it does not guarantee that the values are factually correct or safe to use without further checks.
Web search and realtime services
The platform supports web search and image search as billable tool calls, and provides realtime voice interfaces over WebSocket. These are separate concerns from ordinary text generation: search introduces per-call charges and external-source variability, while realtime voice requires connection management and model-specific support.
Advanced examples and development tools
For a streamed request, add the streaming option supported by the selected API and consume the server-sent events as they arrive. The exact event names should come from the current StepFun Responses or Chat Completions reference rather than being hard-coded from another provider's implementation.
For JavaScript, the documented OpenAI-compatible pattern is to configure the OpenAI client with the StepFun key and base URL, then call the corresponding Responses or Chat Completions method. Raw HTTP remains the most portable option when using another language or an HTTP framework.
StepFun also provides AI Studio and developer-oriented programs, including an Agent Builder Program and Startup Program. These can help with experimentation or access programs, but they do not replace the production API documentation, billing rules, or security review.
Limits and production considerations
Individual rate limits are tied to cumulative cash top-up tiers. The published examples are:
| Tier | Concurrency | Requests per minute | Tokens per minute |
|---|---|---|---|
| V0 | 5 | 100 | 500,000 |
| V4 | 130 | 2,600 | 13,000,000 |
Enterprise limits are available through sales. These are account-level published tiers, not a guarantee that every model or endpoint has identical behavior. Implement retries with backoff for temporary rate-limit responses, bound concurrency, and monitor token usage.
No public numeric latency target or SLA was verified in the supplied research. The API provides standard synchronous HTTP, SSE streaming, and WebSocket realtime interfaces, but production teams should measure latency for their own prompts, regions, model choices, and traffic patterns.
Privacy also requires review. The platform's user-privacy materials state that de-identified and securely processed conversation data may be used to optimize models when the relevant experience program is enabled; users can disable that program. Temporary conversations receive separate treatment and are described as not being used for model training or response optimization. Review the current privacy policy before sending confidential, regulated, or personal data.
Model retirement and migration notices matter because model availability changes. Pin the model name you have tested, monitor StepFun migration notices, and maintain a fallback or migration plan instead of assuming a model will remain available indefinitely.
Advantages and limitations
Advantages
- OpenAI-compatible HTTP and SDK patterns reduce the effort needed to prototype.
- Responses, Chat Completions, streaming, tools, files, structured outputs, and realtime interfaces cover a broad range of application designs.
- Published token prices include relatively low-cost options such as Step 3.5 Flash.
- Language, vision, audio, image, and voice services are available across the wider platform, subject to model-specific support.
Limitations
- Model features, endpoint support, pricing, and retirement status can change quickly.
- Identity verification and cumulative top-up tiers affect individual rate limits.
- No public numeric latency or SLA target was verified.
- Availability, documentation, billing, and registration may vary by region, with the ecosystem oriented toward China and some access requiring Chinese mobile-number verification.
- Privacy and data-use settings must be reviewed before sending sensitive content.
When is StepFun API a good or poor choice?
StepFun is a reasonable choice when you want an OpenAI-compatible API, need Chinese-language or multimodal capabilities, value usage-based pricing, or want one platform covering text, vision, tools, files, search, and realtime voice. It is especially worth evaluating when the published model prices and the required feature set match your workload.
It is a weaker choice when your application depends on globally consistent account access, mature English-language documentation for every feature, a publicly documented SLA, fixed long-term model availability, or independently verified enterprise privacy guarantees. Before production deployment, test the exact model and endpoint you plan to use, confirm regional access and billing, review privacy settings, and load-test against the applicable rate tier.
