Reka Flash

Reka Flash 3

by Reka AI · Current and accessible; open-weight research-preview model with API availability through Reka Gateway

A 21-billion-parameter open-weight reasoning model from Reka AI focused on coding, instruction following, function calling, low-latency inference and local deployment. It is available through Reka Gateway and can also be downloaded under the Apache 2.0 license.

Text Reasoning Coding
Reka Flash 3 is a compact, open-weight language model for developers who need fast reasoning, coding assistance, and local deployment rather than multimodal input or content generation. Released as a research preview on March 10, 2025, it combines a 21B-parameter architecture with reinforcement learning, reasoning traces, and budget forcing. The downloadable release originally documented a 32K context window, while the current Reka Gateway API listing provides a 64K context window.
Outputs

What Reka Flash 3 can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Streaming
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Reka Flash
Model type Reasoning
Context window 66K tokens
Release date 2025-03-10
Status Current and accessible; open-weight research-preview model with API availability through Reka Gateway
Knowledge cutoff notes

No authoritative provider-published knowledge-cutoff date was found for the exact Reka Flash 3 model.

Model notes

Reka Flash 3 is a 21-billion-parameter causal language model trained from scratch and released under the Apache 2.0 license. The original open release documented a 32K context window, while the current Reka Gateway listing provides a 64K context window for the reka-flash-3 API model. The model uses reasoning tags and supports budget forcing, allowing applications to limit its reasoning length. Reka describes it as primarily English-focused and notes that it has not undergone extensive alignment or persona training. A separate multimodal version of the model is used in Reka products such as Reka Vision and Reka Research; that should not be conflated with the text model record here. The 3.1 successor is a materially distinct model and is not treated as an alias of Reka Flash 3.

Cost

Model pricing

Input $0.10 per 1 million tokens through Reka Gateway
Output $0.20 per 1 million tokens through Reka Gateway
Model guide

Reka Flash 3: A Fast, Open 21B Reasoning Model for Local AI

Reka Flash 3 is a 21-billion-parameter open-weight reasoning model from Reka AI, designed for coding, general chat, instruction following, function calling, low-latency applications, and local or on-device deployment.

Reka Flash 3 is a 21-billion-parameter causal language model from Reka AI. It was released as a research preview on March 10, 2025, and is aimed at practical language-model workloads where latency, operating cost, and deployment flexibility matter. Its main uses include coding assistance, general chat, instruction following, reasoning, and function calling.

The model is notable because it is available as an open-weight release under the Apache 2.0 license as well as through Reka Gateway. That combination gives developers two broad deployment choices: use a hosted API for convenience, or download and run the model in an environment that supports its deployment requirements. Reka describes the model as suitable for low-latency applications and local or on-device use.

Reka Flash 3 should not be confused with the separate multimodal models used in products such as Reka Vision and Reka Research. The model covered here is a text-oriented model: it accepts text and produces text. It does not natively accept images, audio, or video in the specifications supplied for this record, and it does not generate those media types.

What is Reka Flash 3?

Reka Flash 3 is a relatively compact reasoning model in Reka AI's model lineup. Its 21B parameter size is substantially smaller than many flagship language models, which is relevant to both speed and deployment cost. Parameter count alone does not determine overall quality, but a smaller model can be easier to serve and more practical for applications that handle many requests or operate under hardware constraints.

Reka trained the model from scratch and used reinforcement learning to improve its reasoning behavior. The provider's release materials describe a model that can expose reasoning inside special reasoning tags. Applications can also use budget forcing, a technique that limits or controls how much reasoning the model performs before producing its answer. This can help developers balance answer quality against response time and token usage.

The original public release was positioned as a research preview. It was not presented as a fully aligned, personality-focused consumer assistant. Reka's documentation notes that the model is primarily English-focused and has not undergone extensive alignment or persona training. Those details matter when deciding whether it is suitable for a polished end-user chatbot or for a more controlled developer workflow.

Where it fits in the Reka AI lineup

Reka Flash 3 occupies the open, text-reasoning side of Reka AI's catalog. Reka's broader ecosystem includes multimodal and video-oriented products, but those products should not be treated as evidence that this specific model accepts visual or audiovisual input.

For hosted use, the canonical API model identifier is reka-flash-3 on Reka Gateway. The current Gateway listing reports a 64K-token context window. The downloadable model release originally documented a 32K context window, so the two figures refer to different access paths or versions of the published serving configuration rather than a single universally applicable limit.

Reka Flash 3.1 is a later and distinct successor. It should not be treated as an alias, automatic replacement, or alternate name for Reka Flash 3. If an application depends on the original model's behavior, identifier, license, or deployment package, it should evaluate the successor separately.

Core capabilities and supported inputs

CapabilityReka Flash 3
Model typeReasoning language model
Parameters21 billion
Text inputYes
Text outputYes
Image, audio, or video inputNot supported for this text-model record
Image, audio, or video outputNo
Context window32K in the original open release; 64K through the current Reka Gateway listing
Function or tool useSupported
StreamingSupported through the API
Fine-tuningNo verified value supplied
Maximum output tokensNo verified provider-published value supplied

The context window is the amount of text the model can consider in one request, including the conversation, instructions, retrieved documents, and generated response. A 64K context can accommodate substantial source material, but it does not guarantee that every long document will be handled equally well. For knowledge-intensive applications, retrieval or another source of current information may still be necessary because no authoritative knowledge-cutoff date was supplied for the exact model.

Reasoning and coding

Reasoning is the model's central differentiator. Reka Flash 3 can produce reasoning-tagged responses and supports budget forcing, allowing an application to impose a practical limit on the model's reasoning effort. In a fast interactive application, a smaller reasoning budget may reduce latency. For a difficult programming or multi-step problem, allowing more reasoning may be preferable, although the supplied research does not establish a universal quality threshold or benchmark result.

Coding is another primary use case. The model is intended for code generation, code explanation, instruction following, and developer workflows that require function calling. It can be useful for drafting implementation ideas, transforming code, explaining errors, and producing structured instructions for external tools. These are capability areas supported by the model's positioning and API record; they should not be interpreted as a guarantee that it will outperform larger coding-specialist or frontier models on every repository or programming language.

Function calling is particularly relevant when the model is placed inside an application rather than used only as a chat interface. A developer can have the model select or prepare a call to an application-defined function, such as querying a database or starting a workflow. The model's tool support does not mean that it independently has web access, current information, or permission to execute arbitrary actions. Those capabilities depend on the surrounding application and the tools exposed to it.

Speed, cost, and deployment trade-offs

Reka Gateway lists Reka Flash 3 at $0.10 per 1 million input tokens and $0.20 per 1 million output tokens. These are usage prices for the hosted API, not a recurring consumer subscription and not a complete estimate of the cost of self-hosting. Infrastructure, electricity, storage, hardware, and engineering time still affect the economics of a local deployment.

The low API rates and 21B size make the model attractive for high-volume or latency-sensitive workloads. Its smaller scale can be a practical advantage over larger reasoning models when a task is routine, when responses must arrive quickly, or when an application needs to keep inference costs predictable. The trade-off is that a compact model may be less appropriate for especially difficult reasoning, broad knowledge tasks, complex long-running agents, or applications that need extensive alignment and conversational polish.

The open-weight Apache 2.0 release provides more control than a hosted-only model. Developers can inspect the published model package, adapt their serving setup, and pursue local deployment where their hardware and software stack support it. However, local availability should not be read as a promise that the model will run efficiently on every laptop, phone, or edge device. The supplied research confirms local and on-device positioning, but does not provide a universal hardware requirement or performance guarantee.

Pricing and API access

For the current Reka Gateway listing, pricing is:

  • Input: $0.10 per 1 million tokens.
  • Output: $0.20 per 1 million tokens.
  • Access: Hosted API using the reka-flash-3 model identifier.
  • Streaming: Supported.

The available research does not verify a separate cached-input discount, batch API price, fine-tuning price, or maximum output-token limit for this model. It is therefore safer to calculate application costs from the published input and output rates and to verify any additional Gateway terms before production deployment.

Main limitations

  • Text-only operation: This model record supports text input and text output, not native image, audio, or video understanding.
  • English emphasis: Reka describes the model as primarily English-focused, so multilingual applications should test their target languages rather than assume broad parity.
  • Limited alignment and persona training: The model was not extensively tuned as a polished general-purpose consumer assistant. Developers may need to provide stronger system instructions, output validation, and safety controls.
  • No verified knowledge cutoff: The supplied sources do not establish a precise cutoff date, and the model should not be assumed to know current events without retrieval.
  • Research-preview status: The original release was a research preview. Model behavior, API availability, and surrounding product support may change.
  • No verified maximum output limit: Applications that depend on a precise completion size should check the current Gateway documentation.

These limitations are important because the model's low cost and open license may otherwise encourage developers to use it for tasks outside its strongest operating range. A text reasoning model is not automatically a replacement for a multimodal model, a current-information search system, or a heavily aligned consumer assistant.

When to choose Reka Flash 3

Choose Reka Flash 3 when the application needs a relatively fast and inexpensive reasoning model for text-based work. It is a sensible candidate for:

  • coding assistance and code transformation;
  • instruction-following systems with controlled prompts;
  • function-calling workflows that connect language understanding to application tools;
  • high-volume text inference where per-token cost matters;
  • local, private, or on-device experimentation supported by suitable hardware;
  • developer research involving open-weight reasoning models; and
  • applications that benefit from controlling reasoning effort with budget forcing.

Another option may be more appropriate when the application requires image, audio, or video input, extensive multilingual performance, a mature consumer-chat experience, strong persona alignment, or reliable access to current information. A larger reasoning model may be preferable for unusually complex tasks if additional quality is worth higher latency and cost. A multimodal Reka product may be more suitable for visual or video understanding, but that does not change the capabilities of the Reka Flash 3 text model itself.

Bottom line

Reka Flash 3 is best understood as an efficient, open-weight text reasoning model rather than as a general multimedia assistant. Its combination of 21B parameters, Apache 2.0 licensing, hosted API access, function calling, reasoning controls, and low published token prices gives developers meaningful flexibility. The strongest case for using it is a cost-sensitive or latency-sensitive coding and reasoning workflow that can be tested and controlled by the application owner.

Its limitations are equally clear: the model is primarily English-focused, lacks native multimodal input and output, has no verified knowledge-cutoff or maximum-output specification in the supplied research, and was not extensively aligned for consumer-style conversations. Those trade-offs make Reka Flash 3 a practical specialized choice, but not a universal substitute for larger, more current, more multilingual, or multimodal systems.


Answers to Frequently Asked Questions

What are the main limitations of Reka Flash 3?
Reka Flash 3 is primarily English-focused, lacks native multimodal capabilities, and was not extensively aligned for polished consumer-style conversations. The supplied information does not verify a precise knowledge cutoff, maximum output-token limit, or universal hardware requirement. Developers should also account for its research-preview status and test performance for their specific use case.
What are the context window and API prices for Reka Flash 3?
The original open release documented a 32K-token context window, while the current Reka Gateway listing reports 64K tokens. Reka Gateway lists pricing at $0.10 per 1 million input tokens and $0.20 per 1 million output tokens. The hosted API uses the model identifier reka-flash-3.
Does Reka Flash 3 support images, audio, or video?
No. Reka Flash 3 is a text-oriented model that accepts text and produces text. It does not natively support image, audio, or video input or output in the specifications described for this model.
What is Reka Flash 3?
Reka Flash 3 is a 21-billion-parameter causal reasoning language model from Reka AI, released as a research preview on March 10, 2025. It is designed for text-based tasks such as coding assistance, instruction following, general chat, reasoning, and function calling.
Can Reka Flash 3 be run locally?
Yes. Reka Flash 3 is available as an open-weight model under the Apache 2.0 license, allowing developers to download and deploy it in compatible environments. It is also available through the hosted Reka Gateway API. Local deployment still depends on suitable hardware, software, storage, and infrastructure.


Sources 5
Provider

About Reka AI