What is Meta Model API, and who is it for?
Meta Model API is Meta’s current self-serve developer platform for calling hosted Meta models over HTTP. The main hosted model family is Muse Spark, with additional services for image generation, speech transcription, and media segmentation. The platform is intended for developers building applications, assistants, automations, and agents rather than for ordinary end-user chat.
The primary API base URL is https://api.meta.ai/v1. New applications should generally start with the Responses API because it is designed for agentic and multi-step workloads. It supports reasoning continuity, server-managed response state, search grounding, background responses, file references, and the broader current feature set.
Meta also supports two compatibility-oriented surfaces:
- Responses API: The recommended choice for new applications and agent workflows.
- Chat Completions: A drop-in option for applications that already use the OpenAI messages-array format.
- Messages API: An Anthropic-compatible interface for clients and tools built around that protocol.
All three use bearer-token authentication and access the platform’s supported models, although individual capabilities depend on the selected model and endpoint.
Getting access and creating an API key
Meta Model API is documented as a public-preview, self-serve platform with expanded global access. Developers use the Meta developer dashboard and Playground to create or manage access, inspect available capabilities, and test requests.
Authentication uses a bearer token. Store the token in an environment variable such as MODEL_API_KEY rather than placing it directly in source code, a browser application, or a public repository. Requests include the token in the HTTP Authorization header:
Authorization: Bearer ${MODEL_API_KEY}Before sending production data, check the current account, privacy, deletion, retention, and contractual terms. The reviewed documentation does not specify one universal retention period for every endpoint and request type.
Choosing the API surface and model
Use the Responses API when you are starting a new project, especially if the application may need tools, web search, files, background work, or multi-turn reasoning. The previous_response_id field can connect a new request to an earlier response and preserve server-managed context across turns.
Choose Chat Completions when you already have an OpenAI-compatible application and want to change the base URL and credentials with minimal restructuring. It remains supported, but it does not provide the same reasoning continuity for external API keys.
Choose the Messages API when your application or framework expects Anthropic’s request and response conventions.
Hosted capabilities described in the current documentation include Muse Spark, Muse Image, Muse Voice Transcribe, and Segment Anything Model 3.1. Muse Glimmer is different: it is an open-weight model intended for self-hosted inference and should not be treated as a hosted Meta Model API endpoint.
Making a first request
The following request uses the recommended Responses endpoint. Set the API key before running it:
export MODEL_API_KEY="your-api-key"curl -sS https://api.meta.ai/v1/responses
-H "Authorization: Bearer ${MODEL_API_KEY}"
-H "Content-Type: application/json"
-d '{
"model": "muse-spark-1.3",
"instructions": "You are a concise technical assistant.",
"input": "Explain why API clients should use exponential backoff.",
"stream": false
}'The request specifies a model, a high-level instruction, user input, and whether the response should be streamed. For a simple application, the returned response can be displayed to the user or passed to the next part of your program.
The same request with the OpenAI Python SDK
Meta documents compatibility with the OpenAI SDK. This example uses one consistent current SDK pattern and points the client at Meta’s base URL:
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.meta.ai/v1",
api_key=os.environ["MODEL_API_KEY"],
)
response = client.responses.create(
model="muse-spark-1.3",
instructions="You are a concise technical assistant.",
input="Give one practical tip for handling HTTP 429 responses.",
)
print(response.output_text)Meta also provides official Python and TypeScript libraries. The OpenAI-compatible route can be useful when an existing application already uses the OpenAI client pattern, while Meta’s own SDKs may be preferable when you want platform-specific documentation and types.
Understanding the response
The Responses API returns a response object that can contain one or more output items. For ordinary text generation, SDKs expose the convenient output_text value. Lower-level integrations can inspect the output items and their content parts directly.
Do not assume every response contains only text. Tool calls, citations, structured data, reasoning-related items, and other supported output types may require inspecting the response structure rather than reading a single string.
For multi-turn workflows, save the response identifier when you need to continue the interaction. A later request can use previous_response_id to reference the earlier response and preserve server-managed state where supported.
How Meta Model API pricing works
Meta Model API uses pay-as-you-go pricing. Muse Spark token prices distinguish input tokens, cached input tokens, and output tokens. A token is a small piece of text processed by the model; output tokens are generated by the model.
| Capability or tier | Documented price |
|---|---|
| Muse Spark Standard input | $1.25 per million tokens |
| Muse Spark Standard cached input | $0.15 per million tokens |
| Muse Spark Standard output | $4.25 per million tokens |
| Contributor input | $0.10 per million tokens |
| Contributor cached input | $0.002 per million tokens |
| Contributor output | $0.20 per million tokens |
| Muse Image | $0.01 per successfully generated image |
| Muse Voice Transcribe | $0.18 per hour of audio processed |
| Segment Anything Model 3.1 | $2.50 per 1,000 images or $0.20 per 1,000 video frames |
Contributor pricing is lower, but prompts and completions may be used to train future Meta models. Standard-tier Muse Spark prompts and completions are documented as not being used to train Meta models. If that distinction matters for your data, evaluate the tier and current terms before implementation.
Capabilities available to developers
Streaming responses
Streaming sends output incrementally through server-sent events instead of waiting for the complete response. It can reduce the time before users see the first generated text and is useful for chat interfaces. Streaming is also available for streamed tool-call arguments and speech transcription.
const stream = await client.responses.create({
model: "muse-spark-1.3",
input: "Stream three short API reliability tips.",
stream: true,
});
for await (const event of stream) {
if (event.type === "response.output_text.delta") {
process.stdout.write(event.delta);
}
}Function and tool calling
Tool calling lets the model request an action from your application, such as looking up an order or querying an internal service. Your server remains responsible for executing the function, validating its arguments, enforcing permissions, and sending the result back. Meta documents function tools, parallel tool calls, tool search, and agent-oriented loops.
Because tool calls can cause real-world side effects, treat model-generated arguments as untrusted input. Validate schemas and authorization in application code instead of allowing the model to call sensitive systems directly.
Structured outputs
Structured outputs constrain a response to a developer-provided JSON schema. This is useful when the result must be consumed by software, such as an invoice extractor, classifier, or workflow router. A schema reduces formatting ambiguity, but your application should still validate the returned data and handle refusals, missing fields, and API errors.
Files and multimodal input
Supported models and endpoints can accept image, document, video, and other file inputs. You can upload files and reference them by file ID across requests, which is useful for document analysis and multi-step workflows. Image and video understanding, file references, and other multimodal features depend on the selected model and endpoint.
Web search and media services
The Responses API supports web search grounding with inline citations. This allows an application to request current information while exposing supporting references in the response workflow. Meta also documents Muse Image for image generation and editing, Muse Voice Transcribe for file and realtime WebSocket transcription, and Segment Anything Model 3.1 for image and video segmentation.
Advanced request patterns
For a multi-step agent, a typical pattern is to send a user request to the Responses API, allow the model to request an approved tool, execute that tool on your server, and continue the response with the tool result. You can also use previous response references when the workflow needs continuity across turns.
Background responses are intended for work that should continue without holding an interactive request open. The platform also supports encrypted reasoning replay and server-managed state in the Responses workflow. These features are useful for longer agent tasks, but they should be tested with realistic workloads because queueing, tool execution, and reasoning effort affect latency.
Playground, SDKs, and developer tools
The Meta Model API dashboard includes a browser-based Playground for testing chat, streaming, tools, structured output, and file inputs without writing an application first. It is useful for checking prompts, comparing request shapes, and confirming that a capability is available to a selected model.
Meta provides official Python and TypeScript packages. The platform also documents compatibility with the OpenAI and Anthropic SDKs. Whichever client you choose, keep the base URL, authentication method, model identifier, and API method consistent. Avoid mixing examples from different SDK generations.
Rate limits and production considerations
Limits apply at the team level rather than separately to each API key. The documented limits include:
| Area | Limit |
|---|---|
| Standard Muse Spark | 3,000 requests per minute and 4,000,000 tokens per minute |
| Contributor Muse Spark | 100 requests per minute and 3,000,000 tokens per minute |
| Muse Image | 150 requests per minute |
| Background response submissions | 600 submissions per minute per team by default |
| Muse Voice Transcribe | 128 concurrent streams and 16,000 streams per hour |
Successful token-based responses include rate-limit headers such as x-ratelimit-limit-tokens, x-ratelimit-remaining-tokens, x-ratelimit-limit-requests, and x-ratelimit-remaining-requests. Monitor these values and implement exponential backoff with jitter after HTTP 429 responses.
Meta does not publish a universal latency guarantee in the reviewed documentation. Streaming can improve time to first visible output, but total latency varies with the model, request size, reasoning effort, tools, queueing, and media processing.
For production systems, keep keys in a secret manager, set request timeouts, retry only appropriate transient failures, log request identifiers and status codes, monitor token usage, validate tool arguments, and avoid sending sensitive data until your organization has reviewed the applicable data-handling terms.
When Meta Model API is a good or poor choice
Meta Model API is a good fit when you want a current Meta-hosted model platform with a recommended agent-oriented Responses API, OpenAI- and Anthropic-compatible request surfaces, multimodal input, tools, structured outputs, web search grounding, file references, and separate media services. It is especially convenient for teams that already understand one of those compatible SDK patterns.
It may be a poor fit when you require a mature, fully stable production contract during the public-preview period, a universal latency SLA, a clearly documented retention period for every endpoint, or a single fixed pricing model across all capabilities. Fine-tuning was not verified in the reviewed hosted API documentation. Contributor-tier training terms may also make that tier unsuitable for applications that cannot permit prompts and completions to be used for model improvement.
Before committing to the platform, test the exact model and endpoint you plan to use, measure latency and token consumption with representative requests, confirm regional availability, and review the current pricing, privacy, and retention documentation.
