What the ByteDance Seed API is and when to use it
The current developer API for ByteDance Seed models is provided through BytePlus ModelArk. ModelArk is a managed model-serving platform: it provides the console, model activation, API keys, endpoints, billing, playground, documentation, and operational controls needed to call Seed models from an application.
ModelArk supports OpenAI-compatible integration patterns, including Chat Completions and Responses APIs. This makes it a practical option for developers who already understand standard chat-message requests or want to connect an existing OpenAI-compatible client to BytePlus infrastructure.
Use ModelArk when you need access to versioned Seed models for text generation, reasoning, coding, multimodal understanding, tool use, file-based workflows, agents, or fine-tuning. It is less suitable if you need one globally uniform consumer product, identical behavior across every model, or a single pricing and availability policy for the entire ByteDance AI ecosystem.
Getting access and obtaining credentials
- Create or use a BytePlus account and open the ModelArk console.
- Activate the Seed model required by your project. Model availability can depend on the account, region, project, and model.
- Create a long-term ModelArk API key in the console.
- Store the key in an environment variable such as
ARK_API_KEYrather than placing it directly in source code. - Use the regional base URL supplied by the console and documentation. An example documented endpoint is
https://ark.ap-southeast.bytepluses.com/api/v3.
Requests authenticate with a bearer token in the Authorization header. For larger or more controlled deployments, ModelArk also provides inference endpoints and additional access-control options.
Choosing an API and model
For a conventional conversation or a simple text-generation request, start with the OpenAI-compatible Chat API. It uses a model identifier and an array of role-based messages. Chat Completions also supports streaming, tool calls, and structured response formats for compatible models.
The Responses API is intended for newer agent-oriented workflows, including persistent response context, file-based multimodal inputs, and tool-call continuation. Choose it when the application needs more than a straightforward request-and-answer exchange.
ModelArk uses versioned model identifiers. Examples documented for the current platform include seed-2-0-lite-260428, seed-2-0-mini-260428, seed-2-0-pro-260328, and seed-1-8-251228. These identifiers and their capabilities can change, so production applications should pin a specific version and check the current model reference before changing it.
Making your first request
The following curl example uses the Chat Completions-compatible endpoint. It sends a system instruction and a user message, then prints the returned text with jq.
curl --fail-with-body -sS https://ark.ap-southeast.bytepluses.com/api/v3/chat/completions
-H "Authorization: Bearer ${ARK_API_KEY:?Set ARK_API_KEY}"
-H "Content-Type: application/json"
-d '{
"model": "seed-2-0-lite-260428",
"messages": [
{"role": "system", "content": "You are a concise technical assistant."},
{"role": "user", "content": "Explain API rate limits in two sentences."}
],
"stream": false
}' | jq -r '.choices[0].message.content'The important request fields are the model identifier, the messages, and the stream setting. The exact model must be activated and supported for the selected account and region.
Understanding the response
A successful Chat Completions response places the generated assistant message in the first choice, commonly accessed as choices[0].message.content. Applications should still handle cases where the response contains tool calls, an empty content value, or an error instead of ordinary text.
For production code, check the HTTP status, preserve useful error details, and avoid assuming that every model returns the same optional fields. Model-specific documentation should be treated as authoritative for response parameters and compatibility.
How ModelArk pricing works
ModelArk uses mixed usage-based pricing rather than one subscription price for all Seed models. Language models are generally billed according to input and output token usage. Cached input and cache storage can have separate treatment, and batch inference uses separate discounted token prices where available.
Other model categories use different units. Image generation may be priced per image, while video and dedicated model-unit services can use model-specific token, hourly, monthly, or other calculations.
Published examples include seed-2-0-lite-260428 at USD 0.25 per million input tokens and USD 2.00 per million output tokens for prompts up to 128K tokens. The same published tier lists seed-2-0-mini-260428 at USD 0.10 per million input tokens and USD 0.40 per million output tokens. These are model- and version-specific examples, not a universal Seed price. Check the current pricing page before estimating costs, particularly when prompt length, caching, batch processing, or a different model is involved.
Capabilities available to developers
Streaming output
Set stream to true to receive output incrementally rather than waiting for the complete response. Streaming can improve perceived responsiveness for long answers and complex reasoning tasks, but the client must process partial events and handle interrupted connections.
Function and tool calling
Compatible models can request application-defined functions using JSON-schema-described tools. Your application remains responsible for executing the function, validating its arguments, and sending the result back to the model. Tool calling is useful for actions such as looking up service status, querying an internal system, or invoking a controlled business operation.
Structured outputs
ModelArk documents structured output support through response_format, including JSON object and JSON Schema modes. This is useful when application code needs predictable fields instead of free-form prose. The documentation identifies structured outputs as beta for some model versions, so verify compatibility before depending on it in production.
Multimodal input and files
Applicable Seed models can support image, video, audio, and text understanding. The Files API accepts files and URLs for later requests, including image, video, document, and agent workflows, subject to model and file-type restrictions.
Uploaded video files are documented as having a seven-day default storage period in applicable workflows, with configurable validity periods ranging from one to thirty days. Supported workflows can also use BytePlus TOS-backed storage when longer or customer-controlled retention is required.
Agents and fine-tuning
ModelArk provides agent APIs with configurable tools, skills, MCP servers, files, sessions, and persistent agent configuration. It also provides fine-tuning workflows and post-fine-tuning inference for eligible models and deployment configurations. These features require more setup than a basic chat request and should be evaluated against the specific model and account documentation.
Python SDK example with structured output
The OpenAI Python package can be configured with ModelArk’s compatible base URL. This example requests a JSON Schema response from a versioned Seed model.
# Install: pip install --upgrade "openai>=1.0"
import json
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["ARK_API_KEY"],
base_url="https://ark.ap-southeast.bytepluses.com/api/v3",
timeout=1800.0,
)
response = client.chat.completions.create(
model="seed-2-0-lite-260428",
messages=[
{"role": "system", "content": "You are a precise technical assistant."},
{"role": "user", "content": "Give two recommendations for a reliable API client."},
],
response_format={
"type": "json_schema",
"json_schema": {
"name": "recommendations",
"strict": True,
"schema": {
"type": "object",
"properties": {
"recommendations": {
"type": "array",
"items": {"type": "string"}
}
},
"required": ["recommendations"],
"additionalProperties": False
}
}
}
)
content = response.choices[0].message.content
print(json.dumps(json.loads(content), indent=2))Because structured outputs are model-dependent and may be beta, production code should handle unsupported parameters and validate the returned JSON before using it.
JavaScript example with a tool definition
The following example shows the same OpenAI-compatible client style with a function tool. ModelArk returns a tool call when the model decides that the application function is needed; your application must execute it and handle the result.
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.ARK_API_KEY,
baseURL: "https://ark.ap-southeast.bytepluses.com/api/v3",
timeout: 1800000
});
const completion = await client.chat.completions.create({
model: "seed-2-0-lite-260428",
messages: [
{ role: "user", content: "Check the status of the payments service." }
],
tools: [
{
type: "function",
function: {
name: "lookup_service_status",
description: "Return the service status for a named service.",
parameters: {
type: "object",
properties: { service: { type: "string" } },
required: ["service"],
additionalProperties: false
},
strict: true
}
}
],
tool_choice: "auto"
});
const message = completion.choices?.[0]?.message;
if (!message) throw new Error("No assistant message returned");
console.log(message.content ?? "");
for (const call of message.tool_calls ?? []) {
console.log(call.function.name, JSON.parse(call.function.arguments));
}SDKs, playground, and developer tools
ModelArk documents OpenAI-compatible Python and JavaScript integrations. BytePlus or ModelArk SDK examples are also available for Python, Go, Java, and other supported environments. Use one consistent SDK generation and verify the current package documentation before copying an example into a project.
The ModelArk console and playground allow developers to test models before integrating them into an application. Documentation includes API references, model capability information, authentication guidance, pricing, streaming, files, agents, multimodal workflows, and reasoning-related guidance.
Limits and production considerations
Language-model limits are primarily expressed as requests per minute and tokens per minute. Quotas depend on the account, project, model, and endpoint and can be viewed in the ModelArk console. Limit increases can be requested through the platform. Image and video models may use image-per-minute or model-specific quotas instead.
There is no single public latency guarantee for the entire ModelArk catalog. Streaming can reduce perceived waiting time, while flexible service tiers, provisioned endpoints, or model units can support workloads with more controlled throughput or latency requirements.
Production applications should use bounded exponential backoff for transient failures, respect RPM and TPM limits, keep API keys out of client-side code, log request identifiers and errors safely, and avoid assuming that one model’s parameters work on every other Seed model. Pin model versions, monitor deprecations, and test changes against representative prompts.
BytePlus states that ModelArk customer data is not used to train or optimize models unless the customer separately approves that use. Data may still be processed for service delivery, safety and abuse detection, troubleshooting, security investigations, and aggregated service statistics. Review the applicable service terms before sending confidential or regulated information.
When ModelArk is a good or poor choice
It is a good choice when:
- You want developer access to ByteDance Seed models through managed APIs.
- Your team already uses OpenAI-compatible Chat or Responses patterns.
- You need a combination of text, coding, multimodal, tool, file, agent, or fine-tuning workflows.
- You can manage model activation, regional endpoints, versioned identifiers, and usage-based billing.
- You want a playground and documented platform features before building a production integration.
It may be a poor choice when:
- You require one consistent global endpoint and identical capabilities across all models.
- You need transparent, fixed pricing independent of model version, prompt length, caching, or media type.
- Your users need a standalone consumer assistant rather than an application-development platform.
- Your deployment cannot accommodate region-specific availability, account activation, or changing model versions.
- You need a universal native web-search feature; documented web search is provided through the Info Quest MCP integration rather than guaranteed for every Seed model.
In practical terms, BytePlus ModelArk is best evaluated as a family of managed model services. The API surface is familiar, but the details that matter in production—model support, pricing, quotas, retention, and feature compatibility—must be checked for the exact model, region, and account configuration you plan to use.
