What is the Reka API and when should you use it?
Reka API is a usage-based developer platform for calling Reka models and selected models through Reka Gateway. The primary interface is an OpenAI-compatible Chat Completions API at https://api.reka.ai/v1. It accepts messages and returns assistant responses in a familiar format, which makes it a practical option for developers already using OpenAI-style client libraries.
The Chat API is intended for text conversations, multimodal prompts, streaming interfaces, and application-controlled tool use. Supported multimodal inputs include text, images, short videos, and audio, depending on the model and endpoint. Reka also offers separate services for use cases that need more than an individual chat request:
- Chat API: General text and multimodal inference, streaming, and tools.
- Reka Gateway: A shared interface for Reka models and selected additional models.
- Reka Research: Web-assisted research and analysis with tool activity, source metadata, streaming, and documented JSON-schema output.
- Reka Vision: Persistent video and image upload, indexing, semantic search, video question answering, tagging, and highlight-clip generation.
Use Chat when you need a normal model request or multimodal analysis. Use Vision when you need a searchable, persistent media collection. Use Research when web search, connected files, research steps, or strict structured responses are central to the workflow.
How do you get access and obtain an API key?
Create an account on the Reka Platform and generate a key from the API Keys area. The API is usage-based, so the account needs available credits before requests can succeed. If the balance reaches zero, requests can fail with an insufficient-balance error.
For the Chat API, the current documentation uses the X-Api-Key header. Some newer Gateway examples show Bearer-token authentication, so check the authentication instructions for the exact product surface and account configuration you are using. Do not expose a key in browser JavaScript, a mobile application, a public repository, or client-side HTML. Send requests through a server or another trusted environment.
Before building a production integration, confirm four values in the documentation for the selected service: the base URL, authentication header, model or endpoint identifier, and applicable pricing and retention rules. Chat, Research, Gateway, and Vision are related products but are not interchangeable endpoints.
Which API and model should you choose?
For a first integration, start with the Chat API and a generally useful available model such as reka-flash. Publicly documented model identifiers include reka-flash and reka-edge; the Gateway catalog also lists identifiers such as reka-flash-3 and reka-edge-2603. The available catalog can depend on the account and may change, so query or consult the current model documentation rather than permanently assuming that every identifier is available.
| Requirement | Recommended surface | Why |
|---|---|---|
| Ordinary chat or multimodal inference | Chat API | OpenAI-compatible messages and chat completions |
| Web-assisted research and analysis | Reka Research | Research tools, source metadata, streaming, and structured output |
| Search across a media library | Reka Vision | Upload, indexing, semantic search, video Q&A, and metadata workflows |
| Calling several available model providers through one interface | Reka Gateway | Shared gateway access to Reka and selected additional models |
There is no publicly documented persistent Assistants API or provider-managed thread-and-run runtime in the supplied documentation. If your application needs durable conversations, tool state, or file ownership, keep that state in your own database and application layer.
How to make your first Reka API request
The simplest current pattern uses the OpenAI Python package with Reka's base URL and an API key. Install the package and set the key in your environment:
pip install openai
export REKA_API_KEY="your_api_key"Then send a chat-completions request:
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.reka.ai/v1",
api_key=os.environ["REKA_API_KEY"],
)
response = client.chat.completions.create(
model="reka-flash",
messages=[
{"role": "system", "content": "You are a concise technical assistant."},
{"role": "user", "content": "Explain HTTP streaming in two sentences."},
],
)
print(response.choices[0].message.content)The equivalent REST request uses POST /v1/chat/completions:
curl --fail-with-body -sS https://api.reka.ai/v1/chat/completions
-H "X-Api-Key: ${REKA_API_KEY}"
-H "Content-Type: application/json"
-d '{
"model": "reka-flash",
"messages": [
{"role": "user", "content": "Explain HTTP streaming in two sentences."}
],
"stream": false
}'In JavaScript, the current OpenAI SDK pattern is similar:
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.reka.ai/v1",
apiKey: process.env.REKA_API_KEY,
});
const response = await client.chat.completions.create({
model: "reka-flash",
messages: [
{ role: "user", content: "Explain multi-turn chat in two sentences." },
],
});
console.log(response.choices[0].message.content);How to understand the response
A successful Chat API response follows the familiar chat-completions structure. The generated assistant message is normally available at response.choices[0].message.content in the Python SDK, or the corresponding choices[0].message.content property in JavaScript.
Do not assume that every response contains ordinary text. A request may return tool calls instead of a final answer, and streamed responses arrive as incremental chunks. Your application should check the response type and handle empty or tool-oriented content appropriately. For production code, also inspect HTTP status codes and error objects rather than printing only the response body.
How Reka API pricing works
Reka Chat uses pay-as-you-go pricing based primarily on input and output token usage. The published Chat rates in the supplied research are:
| Model | Input | Output |
|---|---|---|
| Reka Edge | $0.10 per million tokens | $0.10 per million tokens |
| Reka Flash | $0.80 per million tokens | $2.00 per million tokens |
| Reka Core | $2.00 per million tokens | $6.00 per million tokens |
Image, video, and audio inputs may create additional charges. Reka Gateway presents model-specific prices and states that it has no platform fee or markup. Prices and model availability can change, so production systems should use the current catalog and pricing pages rather than hard-coding old assumptions.
Reka Vision uses different units. Published developer pricing includes video indexing at $0.05 per input video minute, image upload at $10 per million images, image search at $0.003 per search, and video clip generation at $0.04 per input video minute, with an additional charge for 1080p generation. These Vision charges should not be mixed with Chat token pricing when estimating costs.
What capabilities can developers access?
Multimodal input
Chat requests can contain text and supported image, video, or audio inputs. For example, an image can be supplied as an image_url content item:
const response = await client.chat.completions.create({
model: "reka-flash",
messages: [
{
role: "user",
content: [
{
type: "image_url",
image_url: {
url: "https://v0.docs.reka.ai/_images/000000245576.jpg",
},
},
{ type: "text", text: "Describe the main subject of this image." },
],
},
],
});
console.log(response.choices[0].message.content);Input support is model- and endpoint-specific. For large media collections, use Vision's upload and indexing workflow instead of repeatedly attaching the same files to individual Chat requests.
Streaming responses
Set stream to true to receive generated text incrementally. Streaming is useful for interactive interfaces because the user can see output while the model is still producing it:
const stream = await client.chat.completions.create({
model: "reka-flash",
stream: true,
messages: [
{ role: "user", content: "Write a short explanation of API streaming." },
],
});
for await (const chunk of stream) {
const text = chunk.choices[0]?.delta?.content;
if (text) process.stdout.write(text);
}
process.stdout.write("n");Streaming changes how responses are consumed; it does not remove the need for timeout handling, error handling, and usage monitoring.
Tool and function calling
Chat supports tool definitions and tool-choice controls. A tool definition describes a function your application can execute. The model can request that function, but your application—not Reka—must perform the operation and return the result in a subsequent conversation step.
const response = await client.chat.completions.create({
model: "reka-flash",
messages: [
{ role: "user", content: "What is the weather in Boston?" },
],
tools: [
{
type: "function",
function: {
name: "get_weather",
description: "Get the current weather for a city",
parameters: {
type: "object",
properties: { city: { type: "string" } },
required: ["city"],
},
},
},
],
tool_choice: "auto",
});
console.log(response.choices[0].message.tool_calls);Validate tool arguments before execution and apply your own authorization, rate limits, and safety checks. Tool calling is an application integration pattern, not evidence that Reka automatically has access to your systems.
Files and media
File and media handling is available through multimodal Chat inputs and the separate Vision service. Vision supports local multipart uploads or source URLs for videos and images, followed by indexing and operations such as search, question answering, tagging, and clip generation.
Reka also documents fine-tuning as a service involving its engineering team. The supplied documentation does not establish a self-serve Chat fine-tuning endpoint, so do not design an integration around one without confirming a current commercial arrangement.
Structured output
Reka Research documents a response_format object that uses a JSON schema. This is the appropriate documented route when an application needs strict, machine-readable output from a research request. General Chat documentation also discusses prompting for assistant output, but that should not be treated as equivalent to strict schema enforcement unless the endpoint explicitly supports it.
When should you use Research or Vision?
Reka Research uses the specialized reka-flash-research capability for web search and analysis workflows. It can expose tool-related reasoning steps, return source metadata, support streaming, and work with connected private files in supported workflows. It is a better fit than ordinary Chat when the application needs a research process rather than a single model response.
Reka Vision is intended for persistent media workflows. A typical flow is:
- Upload an image or video to the Vision service.
- Index the media for later retrieval.
- Search the indexed collection or ask questions about a video.
- Request tags, summaries, or generated highlight clips.
Vision also provides streaming question answering and an MCP server for compatible agent clients. Its base URL is https://vision-agent.api.reka.ai, separate from the Chat base URL.
SDKs, documentation, and playgrounds
The easiest SDK path is to use the official OpenAI Python or JavaScript package with Reka's compatible base URL. Official examples also cover Go, Java, and REST requests. A dedicated Reka SDK was not verified in the supplied research.
Documentation is available at docs.reka.ai, and the developer playground or Gateway experience is available at developer.reka.ai. Use the documentation for the exact endpoint rather than assuming that Chat authentication, Vision authentication, pricing, or retention rules apply everywhere.
Important limits and production considerations
General Chat rate limits are not comprehensively specified in the reviewed public reference. Vision documents rolling 24-hour per-key limits of 100 image uploads, 100 image searches, 50 video uploads, 50 video searches, and 10 clip-generation requests. Enterprise customers can request higher limits.
Other production considerations include:
- Separate product surfaces: Chat, Research, Gateway, and Vision have different URLs, authentication conventions, pricing units, and supported features.
- Changing catalogs: Model identifiers and availability can change, so confirm them before deployment and monitor failures caused by unavailable models.
- Media retention: Uploaded files may be temporarily stored and may be deleted within a period such as 24 hours depending on the feature. Vision documentation states that indexed video storage is automatically deleted after 30 days, while enterprise arrangements can differ.
- Training choices: The current privacy policy states that interactions from free usage, including promotional API credits, may be used to train and improve models, except for content originating from Google Drive. Paying users can choose whether Reka may use prompts and outputs for improvement through account settings or by contacting Reka.
- Operational resilience: Implement retries for transient failures, explicit timeouts, balance monitoring, request logging that excludes secrets, and handling for rate-limit and insufficient-balance errors.
- Application state: Store conversation history, user permissions, tool results, and durable file metadata in your own application unless a separately documented service provides that state.
Advantages and limitations for developers
Reka API is especially attractive when OpenAI-compatible integration, multimodal input, video understanding, or efficient model access is important. The same platform covers ordinary chat, research tools, and media indexing, and selected models can be deployed through a gateway or locally where supported.
The main limitations are the complexity created by multiple API surfaces and the incomplete public picture around general Chat rate limits. Consumer and developer product boundaries can change, model availability is not guaranteed across accounts, and retention and training terms require careful review for sensitive workloads. Teams looking for a mature provider-managed assistant runtime, broad native client ecosystem, or fully uniform authentication and billing model may find the platform less convenient.
A practical starting path
- Create a Reka Platform account and generate an API key.
- Test
POST https://api.reka.ai/v1/chat/completionswith the OpenAI Python or JavaScript SDK. - Use an available general-purpose model such as
reka-flashfor the first prototype. - Add streaming if the application has an interactive user interface.
- Add tools only after defining validation, authorization, and failure handling in your application.
- Move to Research for web-assisted analysis or documented structured output.
- Move to Vision for persistent video or image indexing, search, video Q&A, tagging, or clip generation.
- Before production, verify current model IDs, prices, limits, retention rules, and data-use settings for the exact endpoint.
