What is Reka Core?
Reka Core is a general-purpose multimodal language model provided by Reka. Reka announced it on April 15, 2024, describing it as a model trained from scratch on thousands of GPUs over several months. In practical terms, it is designed to work with more than ordinary text prompts: a request can involve written instructions together with images, video, audio, or long documents.
The model belongs to Reka’s original Core, Flash, and Edge lineup. Core was the largest and most capable member of that family at launch, while Flash emphasized a lighter and faster profile and Edge targeted efficient or local deployment. Current Reka documentation identifies Flash and Edge as baseline models that are always available for public access. Core remains listed on Reka’s pricing page, but access may depend on the account, platform configuration, or deployment arrangement.
This distinction matters when evaluating Core. Its published specifications make it attractive for demanding analysis, but being listed for pricing is not the same as receiving the same unrestricted availability as a baseline public model.
What Core can process
Reka Core accepts four principal input modalities:
- Text: ordinary prompts, questions, instructions, and documents.
- Images: visual inspection and image-based question answering.
- Video: analysis of video content and questions about events or information contained in a recording.
- Audio: audio-aware workflows, although the supplied documentation does not establish a separate speech-transcription or audio-generation product capability.
Its output is text. Core does not natively generate images, video, or audio according to the supplied model record. A useful way to understand the model is therefore “multimodal input with text output”: it can examine several kinds of media and explain, summarize, compare, or reason about them in written form.
128K context and reasoning capabilities
Core has a documented context window of 128,000 tokens. A token is a small unit of text or other model input; the context window is the amount of information the model can consider during one interaction. A 128K window is large enough for extended documents, collections of related materials, or long multimodal prompts, although the usable amount depends on how the API represents uploaded media and on the surrounding instructions.
The model was designed for reasoning across language, mathematics, visual information, and video content. Potential uses include asking questions about a long technical document, comparing several images, locating an event in a video, or combining written evidence with visual evidence before producing a conclusion. Reka’s technical report and launch materials presented Core as competitive with leading models from its 2024 release period across language, vision, multimodal chat, and video question-answering evaluations. Those are historical provider-reported comparisons, not a guarantee of current performance against newer models.
The supplied research does not verify a knowledge-cutoff date for Core. It also does not document a maximum output-token limit. Developers should therefore avoid assuming that the full 128K context can be returned as output or that a particular response length is guaranteed.
Coding, function calling, and tool use
Reka positioned Core for coding and agentic workflows, and the model record rates its coding capability as useful for software-related tasks. It can generate and explain code, help inspect code-related material, and reason about programming problems. The available information does not establish a language-by-language compatibility matrix or a guaranteed execution environment, so generated code still needs normal testing and review.
Function calling is a more important limitation. Reka’s current function-calling documentation states that only Reka Flash supports function calling at the documented API level. As a result, tool use should not be assumed for Core. Core may be able to describe an action or produce structured instructions in text, but that is different from the API directly invoking an external function on the model’s behalf.
The research also does not verify a separate legacy JSON-mode capability, structured-output guarantee, fine-tuning support, prompt caching, or batch API support for the exact Core model. These omissions do not prove that such features are impossible in every deployment, but they should be treated as unverified rather than assumed.
Pricing and availability
Reka’s current API pricing page lists the following rates for Reka Core:
| Usage type | Listed price |
|---|---|
| Input text tokens | $2.00 per 1 million tokens |
| Output text tokens | $6.00 per 1 million tokens |
| Image input | $0.02 per image |
| Video input | $0.08 per minute |
| Audio input | $0.02 per minute |
These are usage prices rather than a consumer subscription fee. The token charges apply to text input and output, while media inputs have separate listed rates. The actual cost of a multimodal request depends on the amount of text and the duration or number of media items submitted.
The dated model identifier associated with the original public release was reka-core-20240501; reka-core was documented as a shorthand alias. Developers should check the current model list and their account permissions before building a production integration. The presence of Core on the pricing page indicates current listing, but the documentation does not identify it as a model that is always publicly available.
Strengths and trade-offs
Core’s main strength is breadth combined with a large context window. It can accept text, images, video, and audio in one model, which reduces the need to split a workflow across separate specialist systems. This is particularly useful when the answer depends on relationships between different media types, such as matching a written incident report to frames in a video or checking an image against instructions in a long document.
Its second strength is positioning for difficult analysis rather than only short conversational responses. The model was built for complex reasoning, multilingual use, coding, enterprise workloads, and multimodal question answering. Reka reported that its pretraining covered 32 languages.
The trade-offs are equally important. Core is not the obvious choice for native media generation because its documented output is text. It is also not the best fit when guaranteed function calling is central to an agent architecture. Its API pricing is higher for output than for input, and multimodal requests add separate media charges. Compared with a smaller or faster model, Core may be unnecessarily expensive for simple classification, short text transformations, or high-volume low-complexity requests.
Speed is another consideration. The supplied evaluation records give Core a moderate speed score rather than presenting it as an ultra-low-latency model. That score is an editorial assessment, not a provider benchmark or service-level guarantee. In practice, the model’s larger capability profile should be weighed against the response-time and cost requirements of the application.
When to choose Reka Core
Reka Core is a reasonable candidate when the workflow needs several of the following at once:
- Long-context analysis of documents or extended prompts.
- Questions that combine text with images, video, or audio.
- Video or image understanding that produces a written answer.
- Complex multilingual analysis.
- Code generation or code explanation alongside broader reasoning tasks.
- Enterprise or private deployment arrangements where Reka can confirm model access and configuration.
For example, a research team could provide a long report and supporting images, then ask Core to identify contradictions and explain the evidence. A media team could submit a video and ask for a description of important events. A developer could use it to analyze a codebase excerpt together with documentation and screenshots of an interface.
Another option may be more appropriate when the priority is low latency, predictable public availability, or built-in tool execution. Reka Flash is the specifically documented sibling for function calling and is identified, along with Reka Edge, as a baseline model that is always available for public access. A smaller model may also be preferable for routine text tasks where Core’s multimodal and long-context capabilities would not justify the additional cost.
Limitations to check before deployment
Before relying on Core in production, confirm the exact model identifier, account access, supported input formats, and current pricing in Reka’s documentation. The supplied research does not verify a maximum output-token limit, knowledge cutoff, fine-tuning, caching, batch processing, structured outputs, or a distinct JSON mode for Core.
Function calling should be treated as unavailable at the documented API level. Developers should also test media-heavy requests separately from text-only requests because image, video, and audio usage is priced independently. Finally, benchmark results from the 2024 launch period should be treated as historical context rather than a substitute for testing the current endpoint on the application’s own documents, images, recordings, and code.
Bottom line
Reka Core is best understood as a large, text-output model for demanding multimodal analysis. Its verified differentiators are the combination of text, image, video, and audio input, a 128K-token context window, and listed API access with separate media pricing. It can be a strong fit for long-context and cross-media reasoning, but its account-dependent availability, undocumented output and operational limits, and lack of documented function calling make verification essential before deployment.

