Reka Flash

Reka Flash

by Reka AI · Current and publicly available baseline model

Reka Flash is a fast, cost-efficient 21-billion-parameter multimodal model for analyzing text, images, short videos, and audio. It offers a 128K-token context window, structured output, function calling, and public API pricing, but does not provide verified native media generation or a documented maximum output limit.

Text Reasoning Coding
Reka Flash is a 21-billion-parameter multimodal language model from Reka that prioritizes speed and operating cost while accepting text, images, video, and audio as input. Its 128K-token context window, video understanding, structured output, and function-calling support make it a practical choice for production applications that need to analyze mixed media and return usable text or structured results.
Outputs

What Reka Flash can produce

Text
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Tool use Streaming Structured output
Model profile

Performance characteristics

7/10 Reasoning
7/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Reka Flash
Model type Multimodal
Context window 128K tokens
Release date 2024-02-12
Status Current and publicly available baseline model
Knowledge cutoff notes

No direct authoritative knowledge-cutoff date was found for the current rolling reka-flash model identity.

Model notes

Reka Flash is a current rolling model identity exposed publicly as reka-flash. Reka's October 2024 update describes a 21-billion-parameter model with 128K context, interleaved text, image, video, and audio input, function calling, structured output, and improved instruction following. Reka has also described experimental speech-token generation, but public API documentation does not list speech output as a standard capability. The exact knowledge cutoff, maximum output-token limit, fine-tuning support, batch API support, prompt caching, and separate legacy JSON-mode support were not directly verified for the current reka-flash endpoint. The API pricing page lists separate image, video, and audio charges in addition to token pricing.

Cost

Model pricing

Input $0.80 per 1 million input tokens; $0.01 per image; $0.06 per minute of video; $0.015 per minute of audio
Output $2.00 per 1 million output tokens
Model guide

Reka Flash: A Fast Multimodal Model for Short-Video and Document Analysis

Reka Flash is Reka's fast, cost-efficient multimodal language model for applications that combine text with images, video, and audio. The current public reka-flash model supports a 128K-token context window, structured outputs, function calling, and text responses through the Reka API. It is best suited to latency-sensitive document analysis, short-video understanding, structured extraction, multilingual chat, coding, and lightweight tool-using agents rather than native media generation or frontier-level reasoning.

What is Reka Flash?

Reka Flash is a general-purpose multimodal language model provided by Reka. In practical terms, it can read text and inspect images, video, and audio within the same application workflow, then respond primarily with text. Reka introduced the model on February 12, 2024, and later updates positioned it as a fast model for general workloads, multimodal understanding, coding, and agent development.

The current public model identifier is reka-flash. Reka lists it as a baseline model available through the Reka API. The model is described as having 21 billion parameters, although that figure is a provider-published specification rather than an independent measurement.

Reka Flash is not an image, video, music, or speech-generation model. Its main value is understanding supplied content and producing text, structured data, or function calls. Reka has discussed experimental speech-token generation, but its public API documentation does not present speech generation as a standard output capability for the current reka-flash endpoint.

Where Reka Flash fits in Reka's lineup

Reka Flash occupies the fast, relatively cost-efficient position in Reka's public model catalog. That makes it a different type of choice from a model selected primarily for maximum reasoning depth or highly specialized video intelligence. Reka's wider ecosystem also includes products such as Reka Vision and Reka Clip, but those are separate offerings; Reka Flash itself is the model evaluated here.

The model is aimed at developers and product teams that need a broadly capable multimodal endpoint without automatically choosing a larger or more expensive system. Its public positioning emphasizes low latency, general instruction following, multilingual understanding, coding, and the building blocks needed for tool-using assistants.

Inputs, context, and output types

Reka Flash accepts interleaved text, images, video, and audio. “Interleaved” means an application can arrange different kinds of content together with text rather than treating the model as text-only with a separate image question. This is useful for prompts such as asking the model to compare a written specification with a product photograph, interpret a chart alongside an explanation, or summarize spoken content in a video.

Reka documents a 128K-token context window. A context window is the total amount of text and represented input content the model can consider in one request or conversation, subject to the provider's handling of each modality. This is large enough for long documents, extensive extracted transcripts, and substantial application context, although it should not be interpreted as an unlimited video or document length guarantee.

For images, Reka states that the updated model can process arbitrary image resolutions and aspect ratios. Its documented image-understanding capabilities include optical character recognition, or OCR, plus analysis of documents, tables, charts, and diagrams.

Reka also describes video understanding, including temporal content, conversations, and environmental sounds contained in video. Provider materials report support for videos of approximately three to five minutes in a single request. Longer material may require streaming or application-level processing, such as dividing the source into segments and combining the resulting analysis.

The standard public endpoint is best treated as a text-output model with multimodal input. Verified output capabilities include ordinary text responses, structured output, and function calls. The supplied research does not verify image, video, audio, music, or speech generation for the current public model.

Reasoning, coding, and tool use

Reka Flash is designed for general reasoning and instruction-following tasks, but the available research does not establish a specialized frontier reasoning mode or provide benchmark results for the current rolling model identity. An editorial assessment rates its reasoning capability at 7 out of 10; that score is an evaluation aid, not a score published by Reka.

Its practical reasoning strength is most apparent in multimodal interpretation. For example, an application can ask it to extract fields from an invoice, explain a chart, identify events in a short video, or combine information from a document and an image. These tasks involve organizing evidence from supplied content rather than relying only on a long chain of abstract reasoning.

Reka positions Flash for coding and multilingual use. The research supports coding as a target workload and gives it an editorial coding score of 7 out of 10, but no supplied benchmark should be treated as a definitive measure of programming performance. It is a reasonable candidate for code explanation, transformation, classification, and tool-assisted developer workflows, while highly complex software engineering may benefit from a model with stronger verified reasoning or coding evaluations.

Function calling lets the model return a structured request for an application-defined function. The application can then execute an operation such as looking up a record, validating extracted data, or sending a task to another service. Reka also documents structured output, which is useful when the application needs predictable fields rather than free-form prose. These capabilities make Flash suitable for document pipelines, multimodal question answering, classification, and lightweight agents.

Pricing and API access

Reka's published API pricing lists the following rates for Reka Flash:

Usage typePrice
Input tokens$0.80 per 1 million tokens
Output tokens$2.00 per 1 million tokens
Images$0.01 per image
Video$0.06 per minute
Audio$0.015 per minute

Multimodal charges are separate from token pricing according to the supplied Reka pricing information. A request that includes video, for example, can incur both the applicable video-minute charge and token charges for the request and response. Actual costs depend on the amount of content processed and the generated response.

The canonical public model name is reka-flash. Reka provides API documentation and an API model-list endpoint that can be used to confirm which models are available to a particular account. The research verifies public API availability, but it does not verify a separate fine-tuning service, batch API, prompt caching, or a documented maximum output-token limit for this exact model identity.

What Reka Flash does well

  • Mixed-media understanding: It can accept text, images, video, and audio, making it useful when important information is spread across different formats.
  • Long working context: The documented 128K-token context window supports large documents, transcripts, and application prompts.
  • Short-video analysis: It can examine temporal content and audio in short videos, rather than treating every request as a still-image problem.
  • Structured application integration: Function calling and structured output help developers connect model responses to software workflows.
  • Speed and cost positioning: Reka presents Flash as a fast, cost-efficient model, and its published token rates are lower than what many high-end frontier systems charge for comparable general-purpose usage.
  • Broad everyday coverage: It is intended for general chat, multilingual understanding, coding, instruction following, and multimodal analysis rather than only one narrow task.

Limitations and unresolved specifications

The main limitation is that Reka Flash is primarily an understanding and text-response model. It should not be selected when the application requires native image, video, music, or public speech generation.

Several operational details are not clearly documented for the current rolling reka-flash identity. The supplied research does not verify a maximum output-token limit, knowledge-cutoff date, fine-tuning availability, batch processing support, prompt caching, or a separate legacy JSON-mode feature. Structured output is documented, but that should not automatically be treated as a distinct JSON-mode guarantee.

Video length also requires planning. Reka's approximate three-to-five-minute single-request description is useful guidance, not a universal promise for every file, account, or request configuration. Longer recordings may need segmentation, streaming, or a separate workflow. Likewise, the 128K context figure does not remove practical limits created by media size, processing time, file handling, or application design.

Reka Flash may also be less appropriate for tasks where maximum reasoning reliability is more important than speed and cost. The supplied research does not provide current independent benchmark results showing how it compares with specific frontier competitors, so claims of superiority should be avoided.

When to choose Reka Flash

Choose Reka Flash when an application needs fast, economical analysis of mixed media and can work with text or structured responses. Good examples include:

  • Extracting fields from invoices, forms, tables, charts, and other documents.
  • Answering questions about short videos, including spoken dialogue and visible events.
  • Building multilingual support, classification, summarization, or content-review workflows.
  • Creating an assistant that inspects an uploaded image or document and then calls a business function.
  • Processing audio or video alongside text when the result can be returned as a transcript, summary, label, or structured record.
  • Implementing latency-sensitive coding, question-answering, or general chat features through the Reka API.

Another option may be more appropriate if the product needs generated images, synthetic video, native speech output, a confirmed large output limit, or a documented knowledge cutoff. A model selected specifically for advanced long-form reasoning may also be preferable for difficult mathematical, research, or software-engineering tasks where additional latency and cost are acceptable.

Bottom line

Reka Flash is a practical multimodal model for turning text, images, short videos, and audio into useful textual or structured results. Its combination of 128K context, video understanding, function calling, structured output, and published usage-based pricing gives it a clear role in production workflows that value speed and cost control. The trade-off is that several important operational details remain unverified, and the public model should not be treated as a media-generation system or as a guaranteed frontier-level reasoning model.


Answers to Frequently Asked Questions

How much does Reka Flash cost through the Reka API?
Published pricing lists $0.80 per 1 million input tokens, $2.00 per 1 million output tokens, $0.01 per image, $0.06 per minute of video, and $0.015 per minute of audio. Multimodal charges can apply in addition to token charges.
What is the context window and video capacity of Reka Flash?
Reka documents a 128K-token context window. Provider materials describe support for approximately three to five minutes of video in a single request, although longer videos may require segmentation, streaming, or application-level processing.
Can Reka Flash generate images, video, music, or speech?
No. The current public reka-flash endpoint is primarily an understanding and text-output model. Verified outputs include text, structured output, and function calls; image, video, music, and standard speech generation are not documented capabilities.
What is Reka Flash?
Reka Flash is a general-purpose multimodal language model from Reka that can analyze text, images, video, and audio, then return text, structured data, or function calls. Its public model identifier is reka-flash.
What types of content can Reka Flash analyze?
Reka Flash accepts interleaved text, images, video, and audio. It can perform tasks such as OCR, document and table analysis, chart interpretation, short-video understanding, transcript summarization, and combining information from multiple media types.


Sources 5
Provider

About Reka AI