Reka Flash 3 is a 21-billion-parameter causal language model from Reka AI. It was released as a research preview on March 10, 2025, and is aimed at practical language-model workloads where latency, operating cost, and deployment flexibility matter. Its main uses include coding assistance, general chat, instruction following, reasoning, and function calling.
The model is notable because it is available as an open-weight release under the Apache 2.0 license as well as through Reka Gateway. That combination gives developers two broad deployment choices: use a hosted API for convenience, or download and run the model in an environment that supports its deployment requirements. Reka describes the model as suitable for low-latency applications and local or on-device use.
Reka Flash 3 should not be confused with the separate multimodal models used in products such as Reka Vision and Reka Research. The model covered here is a text-oriented model: it accepts text and produces text. It does not natively accept images, audio, or video in the specifications supplied for this record, and it does not generate those media types.
What is Reka Flash 3?
Reka Flash 3 is a relatively compact reasoning model in Reka AI's model lineup. Its 21B parameter size is substantially smaller than many flagship language models, which is relevant to both speed and deployment cost. Parameter count alone does not determine overall quality, but a smaller model can be easier to serve and more practical for applications that handle many requests or operate under hardware constraints.
Reka trained the model from scratch and used reinforcement learning to improve its reasoning behavior. The provider's release materials describe a model that can expose reasoning inside special reasoning tags. Applications can also use budget forcing, a technique that limits or controls how much reasoning the model performs before producing its answer. This can help developers balance answer quality against response time and token usage.
The original public release was positioned as a research preview. It was not presented as a fully aligned, personality-focused consumer assistant. Reka's documentation notes that the model is primarily English-focused and has not undergone extensive alignment or persona training. Those details matter when deciding whether it is suitable for a polished end-user chatbot or for a more controlled developer workflow.
Where it fits in the Reka AI lineup
Reka Flash 3 occupies the open, text-reasoning side of Reka AI's catalog. Reka's broader ecosystem includes multimodal and video-oriented products, but those products should not be treated as evidence that this specific model accepts visual or audiovisual input.
For hosted use, the canonical API model identifier is reka-flash-3 on Reka Gateway. The current Gateway listing reports a 64K-token context window. The downloadable model release originally documented a 32K context window, so the two figures refer to different access paths or versions of the published serving configuration rather than a single universally applicable limit.
Reka Flash 3.1 is a later and distinct successor. It should not be treated as an alias, automatic replacement, or alternate name for Reka Flash 3. If an application depends on the original model's behavior, identifier, license, or deployment package, it should evaluate the successor separately.
Core capabilities and supported inputs
| Capability | Reka Flash 3 |
|---|---|
| Model type | Reasoning language model |
| Parameters | 21 billion |
| Text input | Yes |
| Text output | Yes |
| Image, audio, or video input | Not supported for this text-model record |
| Image, audio, or video output | No |
| Context window | 32K in the original open release; 64K through the current Reka Gateway listing |
| Function or tool use | Supported |
| Streaming | Supported through the API |
| Fine-tuning | No verified value supplied |
| Maximum output tokens | No verified provider-published value supplied |
The context window is the amount of text the model can consider in one request, including the conversation, instructions, retrieved documents, and generated response. A 64K context can accommodate substantial source material, but it does not guarantee that every long document will be handled equally well. For knowledge-intensive applications, retrieval or another source of current information may still be necessary because no authoritative knowledge-cutoff date was supplied for the exact model.
Reasoning and coding
Reasoning is the model's central differentiator. Reka Flash 3 can produce reasoning-tagged responses and supports budget forcing, allowing an application to impose a practical limit on the model's reasoning effort. In a fast interactive application, a smaller reasoning budget may reduce latency. For a difficult programming or multi-step problem, allowing more reasoning may be preferable, although the supplied research does not establish a universal quality threshold or benchmark result.
Coding is another primary use case. The model is intended for code generation, code explanation, instruction following, and developer workflows that require function calling. It can be useful for drafting implementation ideas, transforming code, explaining errors, and producing structured instructions for external tools. These are capability areas supported by the model's positioning and API record; they should not be interpreted as a guarantee that it will outperform larger coding-specialist or frontier models on every repository or programming language.
Function calling is particularly relevant when the model is placed inside an application rather than used only as a chat interface. A developer can have the model select or prepare a call to an application-defined function, such as querying a database or starting a workflow. The model's tool support does not mean that it independently has web access, current information, or permission to execute arbitrary actions. Those capabilities depend on the surrounding application and the tools exposed to it.
Speed, cost, and deployment trade-offs
Reka Gateway lists Reka Flash 3 at $0.10 per 1 million input tokens and $0.20 per 1 million output tokens. These are usage prices for the hosted API, not a recurring consumer subscription and not a complete estimate of the cost of self-hosting. Infrastructure, electricity, storage, hardware, and engineering time still affect the economics of a local deployment.
The low API rates and 21B size make the model attractive for high-volume or latency-sensitive workloads. Its smaller scale can be a practical advantage over larger reasoning models when a task is routine, when responses must arrive quickly, or when an application needs to keep inference costs predictable. The trade-off is that a compact model may be less appropriate for especially difficult reasoning, broad knowledge tasks, complex long-running agents, or applications that need extensive alignment and conversational polish.
The open-weight Apache 2.0 release provides more control than a hosted-only model. Developers can inspect the published model package, adapt their serving setup, and pursue local deployment where their hardware and software stack support it. However, local availability should not be read as a promise that the model will run efficiently on every laptop, phone, or edge device. The supplied research confirms local and on-device positioning, but does not provide a universal hardware requirement or performance guarantee.
Pricing and API access
For the current Reka Gateway listing, pricing is:
- Input: $0.10 per 1 million tokens.
- Output: $0.20 per 1 million tokens.
- Access: Hosted API using the
reka-flash-3model identifier. - Streaming: Supported.
The available research does not verify a separate cached-input discount, batch API price, fine-tuning price, or maximum output-token limit for this model. It is therefore safer to calculate application costs from the published input and output rates and to verify any additional Gateway terms before production deployment.
Main limitations
- Text-only operation: This model record supports text input and text output, not native image, audio, or video understanding.
- English emphasis: Reka describes the model as primarily English-focused, so multilingual applications should test their target languages rather than assume broad parity.
- Limited alignment and persona training: The model was not extensively tuned as a polished general-purpose consumer assistant. Developers may need to provide stronger system instructions, output validation, and safety controls.
- No verified knowledge cutoff: The supplied sources do not establish a precise cutoff date, and the model should not be assumed to know current events without retrieval.
- Research-preview status: The original release was a research preview. Model behavior, API availability, and surrounding product support may change.
- No verified maximum output limit: Applications that depend on a precise completion size should check the current Gateway documentation.
These limitations are important because the model's low cost and open license may otherwise encourage developers to use it for tasks outside its strongest operating range. A text reasoning model is not automatically a replacement for a multimodal model, a current-information search system, or a heavily aligned consumer assistant.
When to choose Reka Flash 3
Choose Reka Flash 3 when the application needs a relatively fast and inexpensive reasoning model for text-based work. It is a sensible candidate for:
- coding assistance and code transformation;
- instruction-following systems with controlled prompts;
- function-calling workflows that connect language understanding to application tools;
- high-volume text inference where per-token cost matters;
- local, private, or on-device experimentation supported by suitable hardware;
- developer research involving open-weight reasoning models; and
- applications that benefit from controlling reasoning effort with budget forcing.
Another option may be more appropriate when the application requires image, audio, or video input, extensive multilingual performance, a mature consumer-chat experience, strong persona alignment, or reliable access to current information. A larger reasoning model may be preferable for unusually complex tasks if additional quality is worth higher latency and cost. A multimodal Reka product may be more suitable for visual or video understanding, but that does not change the capabilities of the Reka Flash 3 text model itself.
Bottom line
Reka Flash 3 is best understood as an efficient, open-weight text reasoning model rather than as a general multimedia assistant. Its combination of 21B parameters, Apache 2.0 licensing, hosted API access, function calling, reasoning controls, and low published token prices gives developers meaningful flexibility. The strongest case for using it is a cost-sensitive or latency-sensitive coding and reasoning workflow that can be tested and controlled by the application owner.
Its limitations are equally clear: the model is primarily English-focused, lacks native multimodal input and output, has no verified knowledge-cutoff or maximum-output specification in the supplied research, and was not extensively aligned for consumer-style conversations. Those trade-offs make Reka Flash 3 a practical specialized choice, but not a universal substitute for larger, more current, more multilingual, or multimodal systems.

