Reka Flash

Reka Flash 3.1

by Reka AI · Available as an open-weight model; official API availability was announced but the current public baseline API catalog does not list the exact model identifier

Reka Flash 3.1 is a 21B open-weight reasoning model released by Reka AI in July 2025. It focuses on coding, mathematical reasoning, local inference, quantization, and fine-tuning for agentic workflows. The public checkpoint is text-only, has a 32,768-token context window, and is released under Apache 2.0. Current API availability and pricing are not clearly documented.

Text Reasoning Coding
Reka Flash 3.1 is an open-weight reasoning model aimed at developers and researchers who want a relatively efficient model for coding, mathematics, local deployment, or agentic systems. Reka reports that it improves substantially over Reka Flash 3, including a 10-point gain on the full LiveCodeBench v5 benchmark. The public checkpoint contains 21 billion parameters, supports a 32,768-token context window, and is released under the Apache 2.0 license. It is not documented as a native image, audio, or video model, and current public API documentation does not clearly list the exact Flash 3.1 identifier.
Outputs

What Reka Flash 3.1 can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Fine-tuning
Model profile

Performance characteristics

7/10 Reasoning
8/10 Coding
7/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Reka Flash
Model type Reasoning
Context window 33K tokens
Release date 2025-07-10
Status Available as an open-weight model; official API availability was announced but the current public baseline API catalog does not list the exact model identifier
Knowledge cutoff notes

No direct authoritative knowledge-cutoff date was found in the official Reka model card or release announcement.

Model notes

Reka Flash 3.1 is a 21B open-weight model released in Llama-compatible format under the Apache 2.0 license. Reka reports a 10-point improvement over Reka Flash 3 on the full LiveCodeBench v5 benchmark. A 3.5-bit quantized release and the Reka Quant library support lower-memory deployment. Reka's July 2025 announcement stated that the model was available through its API as reka-flash-3.1, but the current public API model documentation lists reka-flash and reka-edge as baseline models and does not list Flash 3.1. The model card does not publish a verified knowledge cutoff or maximum output limit. Product-level references to a multimodal version powering Reka Research and Reka Vision should not be treated as proof that the public RekaAI/reka-flash-3.1 checkpoint accepts image, audio, or video input.

Model guide

Reka Flash 3.1: An Open-Weight 21B Model for Coding and Agentic Workflows

Reka Flash 3.1 is a 21-billion-parameter open-weight reasoning model from Reka AI, released on July 10, 2025. It is designed mainly for coding, mathematical reasoning, local inference, fine-tuning, and agentic applications. The model is distributed in a Llama-compatible format under the Apache 2.0 license, with a 3.5-bit quantized release and the Reka Quant library for lower-memory deployment. Its documented public checkpoint is text-only, despite multimodal versions of the Flash 3.1 technology being used in some Reka products.

What is Reka Flash 3.1?

Reka Flash 3.1 is a 21-billion-parameter reasoning language model developed by Reka AI and released on July 10, 2025. It is an updated version of Reka Flash 3, with its main emphasis on coding, mathematical reasoning, and use as a base model for agentic applications.

In practical terms, the model is intended for tasks such as writing and debugging code, solving structured problems, generating technical explanations, and serving as a component inside software that takes multiple steps to complete a task. It is not primarily positioned as a consumer chatbot or as an image-generation system. Its open-weight release also makes it relevant to organizations that want to run, evaluate, or fine-tune a model outside a fully managed hosted service.

Reka released the model in a Llama-compatible format under the Apache 2.0 license. The official release includes a lower-bit quantized version and the Reka Quant library, giving users more options when their available hardware cannot comfortably run the full model.

Where it fits in the Reka AI lineup

Reka Flash 3.1 belongs to Reka AI's Flash family of language models. It was announced as an API model with the identifier reka-flash-3.1, but its most clearly documented and verifiable current form is the public open-weight checkpoint hosted as RekaAI/reka-flash-3.1.

There is an important availability distinction for developers. Reka's July 2025 announcement said that Flash 3.1 was available through the Reka API. However, the current public API model documentation supplied for this review lists reka-flash and reka-edge as baseline models and does not list Flash 3.1 among those entries. Exact hosted availability and pricing should therefore be checked directly before starting a production integration.

Reka also states that a multimodal version of Flash 3.1 is used in products such as Reka Research and Reka Vision. That product-level information should not be confused with the public checkpoint. The official model card for RekaAI/reka-flash-3.1 documents a text language model rather than a checkpoint with native image, audio, or video input.

Capabilities and performance focus

Flash 3.1 is primarily a text-in, text-out reasoning model. Its strongest intended areas are software development and tasks that benefit from step-by-step problem solving. Examples include producing code from a specification, explaining an error, suggesting a fix, writing tests, and working through mathematical or algorithmic questions.

Reka reports that Flash 3.1 improves by 10 points over Flash 3 on the full LiveCodeBench v5 benchmark. The company also describes it as competitive with models including Qwen3-32B, o3-mini, and Gemini 2.5 Flash Thinking on coding-related tasks. These are provider-reported comparisons, not guarantees of performance on every programming language, codebase, or benchmark configuration.

The model's documented language focus is English. It can understand and generate some other languages, but the model card characterizes that support as limited rather than presenting Flash 3.1 as a broadly optimized multilingual model.

Reasoning, coding, and agentic use

Reka Flash 3.1 was post-trained with supervised fine-tuning followed by large-scale reinforcement learning using verifiable rewards. Reka's technical description identifies an RLOO-based reinforcement-learning approach. In simple terms, the training process used feedback that can be checked for tasks such as coding or mathematics, encouraging solutions that satisfy objective requirements.

This training focus helps explain the model's positioning. Flash 3.1 is not merely a general text generator; it is intended to produce answers that are more useful when correctness can be tested. A developer might use it to generate an implementation and then run a test suite, or to ask for a mathematical solution whose result can be verified independently.

The model is also presented as a foundation for agentic systems. An agentic application is software that divides a goal into steps, uses tools or code, observes results, and continues until it reaches an answer. Flash 3.1 can be fine-tuned for these workflows, but the supplied research does not verify a built-in tool-calling or function-calling interface for the public checkpoint. Tool support should therefore be implemented and tested at the application or serving layer rather than assumed from the model's agentic positioning.

Context window and output limits

The documented context length is 32,768 tokens. A token is a unit of text used by the model; it may represent a word, part of a word, punctuation, or another text fragment. A 32,768-token context allows the model to consider a substantial prompt, source file, technical discussion, or collection of related passages in one request, although the usable amount depends on how much space is reserved for the response.

No verified maximum output-token limit is published in the supplied sources. The model card also does not provide a verified knowledge-cutoff date. Applications that require predictable response sizes should set and test their own serving limits rather than relying on an undocumented maximum.

Deployment, quantization, and license

Flash 3.1 can be used with Transformers-compatible tooling or vLLM, and the release is supplied in a Llama-compatible format. These options make it suitable for users who want to run the model themselves rather than depend exclusively on a hosted endpoint.

Running a 21-billion-parameter model can require significant memory, especially at higher numerical precision. Reka therefore released a 3.5-bit quantized version and the Reka Quant library. Quantization stores model values with fewer bits to reduce memory use and can make local inference more practical. The trade-off is that aggressive quantization can reduce quality, so users should validate the chosen format against their own coding and reasoning tasks.

Reka reports that quantizing Flash 3.1 to the Q3_K_S format with Reka Quant causes substantially less benchmark degradation than a conventional quantization routine. This is a provider claim about the company's quantization approach, not a universal guarantee for every hardware setup or workload.

The model is released under the Apache 2.0 license. That permissive license is useful for research, experimentation, and many commercial deployments, subject to the license terms and any obligations associated with surrounding software, data, or hosted services.

Supported modalities and tool support

The public Flash 3.1 checkpoint is documented as text-only: it accepts text and produces text. The supplied model data marks image, audio, and video input as unsupported, and it does not identify image, audio, video, speech, music, or embedding output. It is therefore not the right checkpoint for directly analyzing an image or video, generating an image, or synthesizing audio.

Reka's broader product ecosystem includes multimodal systems, and the company says a multimodal Flash 3.1 version contributes to products such as Reka Research and Reka Vision. Those product capabilities do not establish that the downloadable public checkpoint accepts non-text inputs.

Native tool use, streaming, structured output, caching, and batch API support are not verified in the supplied model record. A deployment platform may add some of these features, but they should be treated as serving-layer capabilities until confirmed for the specific endpoint or runtime.

Pricing and API availability

No reliable current per-token input or output price is provided for Flash 3.1 in the supplied research. The current public Reka pricing documentation does not list a separate verified price for this exact model. Users considering hosted use should confirm both whether reka-flash-3.1 is still available and what billing terms apply.

The open-weight release provides an alternative to per-token hosted pricing. Self-hosting can be attractive when workloads are large, data must remain within a controlled environment, or the team needs to fine-tune the model. It also transfers responsibility for hardware, deployment, monitoring, scaling, and performance optimization to the user. Quantized versions may reduce infrastructure requirements, but they do not make operational costs disappear.

Main strengths and limitations

  • Strengths: coding and mathematical reasoning are central design goals; the model is open weight and Apache 2.0 licensed; Llama-compatible deployment supports common inference tooling; quantized variants can lower memory requirements; and fine-tuning is relevant for specialized agentic workflows.
  • Limitations: the public checkpoint is text-only; English is the primary language focus; no verified knowledge cutoff or maximum output limit is published; native tool support is not confirmed; and current API availability and pricing are unclear.
  • Operational trade-off: local deployment offers control and potentially better economics at scale, but a 21B model still requires meaningful compute and technical maintenance. A smaller hosted model may be easier to operate for short, high-volume requests.

When to choose Reka Flash 3.1

Choose Flash 3.1 when you need an open-weight model for coding, mathematical reasoning, research experimentation, or local inference. It is particularly worth evaluating when Apache 2.0 licensing, fine-tuning, quantization, or control over the serving environment matters more than having a turnkey consumer product.

It can also be a good candidate for an agentic prototype in which the application supplies tools, executes generated code, or validates results externally. The model's reasoning-oriented post-training and reported coding improvement make that use case plausible, but the application should include tests and safeguards rather than treating generated answers as automatically correct.

Another option may be more appropriate when the primary requirement is native image, audio, or video understanding; a polished consumer chat experience; verified built-in function calling; a clearly documented hosted price; or a smaller model optimized for minimal latency and hardware cost. For multimodal work, use a Reka product or model whose documentation explicitly confirms the needed input type instead of assuming that the public Flash 3.1 checkpoint has the same capabilities.

Bottom line

Reka Flash 3.1 is best understood as an open-weight coding and reasoning model rather than a general multimodal assistant. Its 21B size, Apache 2.0 license, Llama-compatible format, quantized release, and fine-tuning potential make it interesting for developers who want deployment flexibility. Its main uncertainties are equally practical: the public checkpoint's text-only scope, the absence of a verified output limit and knowledge cutoff, and unclear current API listing and pricing. Those trade-offs make hands-on evaluation important before selecting it for production.


Answers to Frequently Asked Questions

Is Reka Flash 3.1 available through an API, and how much does it cost?
Reka announced Flash 3.1 as an API model with the identifier reka-flash-3.1, but current public API documentation supplied for the review does not list it among the baseline models. No reliable current per-token price is provided, so availability and pricing should be confirmed directly with Reka.
How can Reka Flash 3.1 be deployed, and what license does it use?
Reka Flash 3.1 is provided in a Llama-compatible format and can be used with Transformers-compatible tooling or vLLM. Reka also released a 3.5-bit quantized version and the Reka Quant library to reduce memory requirements. The model is licensed under Apache 2.0.
Can Reka Flash 3.1 process images, audio, or video?
The public RekaAI/reka-flash-3.1 checkpoint is documented as text-only. It accepts text and produces text, while image, audio, and video inputs are not supported. Reka products may use multimodal versions, but those capabilities should not be assumed for the downloadable checkpoint.
What is Reka Flash 3.1?
Reka Flash 3.1 is a 21-billion-parameter open-weight reasoning language model developed by Reka AI and released on July 10, 2025. It is designed mainly for coding, mathematical reasoning, technical explanations, and agentic software workflows.
What are the main capabilities of Reka Flash 3.1?
The model focuses on text-based coding and reasoning tasks, including generating and debugging code, writing tests, solving mathematical or algorithmic problems, and producing technical explanations. Reka reports a 10-point improvement over Flash 3 on LiveCodeBench v5, although real-world performance can vary.


Sources 5
Provider

About Reka AI