Command R

Command R7B

by Cohere · Live

Command R7B is Cohere’s compact 7-billion-parameter text model for low-cost, high-throughput applications. It offers a 128K context window, 4K maximum output, 23-language support, tool use, structured outputs, RAG capabilities, and open-weight deployment options under a non-commercial license.

Text Reasoning Coding
Cohere Command R7B is a 7-billion-parameter language model released on December 13, 2024. It combines a 128,000-token context window with low API pricing, text-only generation, multilingual support across 23 languages, and capabilities for retrieval-augmented generation, tool use, and agentic workflows.
Outputs

What Command R7B can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Web search Streaming JSON mode Structured output
Model profile

Performance characteristics

7/10 Reasoning
7/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Command R
Model type Lightweight
Context window 128K tokens
Maximum output 4K tokens
Knowledge cutoff June 1, 2024
Release date December 13, 2024
Status Live
Knowledge cutoff notes

Cohere's exact model documentation lists June 1, 2024 as the knowledge cutoff. Web search, retrieval, connectors, or user-provided documents can supply newer information during use but do not change the underlying cutoff.

Model notes

The exact API identifier is command-r7b-12-2024. Cohere lists the model as live and supports it through Chat API endpoints. It is a text-only model with approximately 7 billion parameters, a 128,000-token context window, and a 4,000-token maximum output. The open-weight release is published by Cohere Labs/Cohere For AI under a CC-BY-NC license with additional acceptable-use requirements, so commercial self-hosting rights should be reviewed separately. Cohere documents tool use, structured outputs, JSON object responses, streaming, and the web-search connector for compatible platform usage. Editorial scores are comparative estimates rather than vendor ratings.

Cost

Model pricing

Input $0.0375 per 1 million tokens
Output $0.15 per 1 million tokens
Model guide

Command R7B: Cohere’s Efficient Model for RAG and Tool-Using Agents

Command R7B is Cohere’s compact, multilingual text-generation model for fast, cost-efficient applications involving retrieval-augmented generation, tool use, coding, chat, and multi-step agents.

What Command R7B is

Command R7B is Cohere’s smallest and fastest model in the Command R family. The model was released on December 13, 2024, for applications where response speed, throughput, operating cost, and deployment efficiency are more important than using the largest available language model.

The name refers to an approximately 7-billion-parameter model. Parameters are the learned values that allow a language model to generate and interpret text; a smaller parameter count generally makes a model less expensive and easier to run than a much larger model, although it can also limit performance on difficult reasoning or knowledge tasks.

Its exact Cohere API identifier is command-r7b-12-2024. Cohere lists the model as live, with a 128,000-token context window and a maximum output of 4,000 tokens. It is available through Cohere’s API and as open weights for research and non-commercial use.

Where it fits in Cohere’s lineup

Command R7B is positioned as an efficient member of Cohere’s Command family rather than as a frontier-scale model. Its role is to provide a practical balance between capability and resource use for language workloads such as retrieval-augmented generation, enterprise chat, coding assistance, tool calling, and agent workflows.

That positioning matters when choosing a model. A larger model may be preferable for especially difficult reasoning, complex coding, or tasks where maximum answer quality is the main priority. Command R7B is aimed at workloads that may generate many requests or require several model calls in sequence, where lower per-token cost and faster processing can have a meaningful operational effect.

Core capabilities and supported tasks

Command R7B is a text-generation model designed to respond to prompts, supplied documents, retrieved passages, and tool results. Retrieval-augmented generation, or RAG, means that an application searches a knowledge base or document collection and gives the relevant results to the model as context. This can help an application answer questions about current or private information without relying only on the model’s original training.

Cohere specifically positions Command R7B for:

  • RAG over supplied documents or external data
  • Conversational applications and enterprise chatbots
  • Tool and function calling
  • Multi-step agents and research workflows
  • Code generation and code-assistance tasks
  • Summarization, question answering, and instruction following

Tool use allows the model to request an operation defined by the surrounding application, such as searching a database, retrieving a document, or calling a business system. The model does not independently operate those systems; the application must execute the requested tool and return the result. This makes Command R7B suitable for agents that break a task into several information-seeking or action-oriented steps.

The model supports 23 languages, including English, French, Spanish, Italian, German, Portuguese, Japanese, Korean, Arabic, Chinese, Russian, Polish, Turkish, Vietnamese, Dutch, Czech, Indonesian, Ukrainian, Romanian, Greek, Hindi, Hebrew, and Persian. This makes it more suitable for multilingual enterprise applications than a model intended only for English-language use, although the supplied research does not provide language-specific benchmark results.

Technical specifications and limits

SpecificationVerified information
Model familyCommand R
Approximate size7 billion parameters
Input modalityText
Output modalityText
Context window128,000 tokens
Maximum output4,000 tokens
Knowledge cutoffJune 1, 2024
Release dateDecember 13, 2024
API model IDcommand-r7b-12-2024

The 128,000-token context window is the maximum amount of text that can be supplied in a request, including the prompt, retrieved material, conversation history, and other input content, subject to the provider’s request handling. It is large enough for long documents or substantial collections of retrieved passages, but applications still need to select relevant context carefully because unnecessary text can increase cost and reduce focus.

The maximum generated response is 4,000 tokens. This is an output limit, not a guarantee that every request will produce a response of that length. Applications that need longer documents may need to generate them in multiple stages.

The documented knowledge cutoff is June 1, 2024. The model therefore should not be treated as inherently current. Retrieval, tools, web-search connectors, or user-provided documents can supply newer information during use, but they do not change the underlying cutoff.

Modalities, coding, and structured responses

Command R7B is text-only. It accepts text and produces text; it does not natively understand images, audio, or video, and it does not generate images, audio, or video. An application can describe an image or transcribe audio elsewhere and then send the resulting text to Command R7B, but that would be an external preprocessing workflow rather than native multimodal support.

The model supports coding assistance, including code generation, explanations, and code-oriented responses. The supplied research does not establish that it is optimized for a particular programming language or that it matches larger models on difficult software-engineering tasks, so coding quality should be evaluated against the application’s own requirements.

Cohere documents structured outputs and JSON object responses for compatible API usage. These features can make it easier for an application to parse fields returned by the model, but the model’s output remains text. Structured output does not turn Command R7B into a non-text or multimodal model.

Pricing and access

Cohere lists Command R7B at $0.0375 per million input tokens and $0.15 per million output tokens. Input tokens are the text sent to the model, while output tokens are the text it generates. The separate rates matter for applications that provide large retrieved documents or conversation histories, as well as for applications that produce long answers.

At these listed rates, Command R7B is positioned as a low-cost option for high-volume text workloads. Its economics are particularly relevant to RAG and agent systems, which may make multiple model calls for one user request. Actual application cost will also depend on how much context is sent, how many tool or agent steps are used, and the deployment arrangement.

The model is available through Cohere’s Chat API, including streaming interfaces. Cohere also documents tool use, structured outputs, JSON object responses, and a web-search connector for compatible platform usage.

Command R7B is also available as an open-weight release for local or private deployment. The supplied model-card information identifies a CC-BY-NC license with additional acceptable-use requirements. This is a significant limitation for businesses: open weights do not automatically provide unrestricted commercial self-hosting rights, so organizations should review the license and applicable terms before deployment.

Strengths and trade-offs

Where Command R7B is strong

  • Cost: The listed token rates are low relative to larger enterprise language models.
  • Speed and throughput: Its compact size is intended to support fast responses and efficient high-volume processing.
  • Long context: A 128,000-token window supports lengthy documents, retrieved passages, and extended conversations.
  • Agent workflows: Tool use and multi-step operation support make it suitable for applications that need more than one generation step.
  • Multilingual coverage: Support for 23 languages broadens its usefulness in international and multilingual applications.
  • Deployment flexibility: API access and open-weight availability provide alternatives for hosted and private experimentation, subject to licensing terms.

Important limitations

  • It is a compact model and should not be assumed to match the strongest large models on difficult reasoning, advanced coding, or broad knowledge tasks.
  • It is text-only, with no native image, audio, or video input or output.
  • The June 1, 2024 knowledge cutoff means current information must come from retrieval, tools, or user-provided context.
  • The 4,000-token maximum output can be restrictive for very long single-pass documents.
  • The open-weight license is non-commercial according to the supplied model-card information, which may restrict commercial self-hosting.
  • Fine-tuning availability should not be assumed. Cohere has retired or restricted fine-tuning capabilities across parts of its older model catalog, and the supplied research does not verify a current fine-tuning option specifically for Command R7B.

Reasoning and coding expectations

Command R7B is intended for practical reasoning within RAG, tool-use, and agentic workflows. It can help interpret retrieved information, decide which available tool to request, summarize results, and produce a response based on multiple steps. These capabilities are useful for enterprise assistants, but they should not be confused with a guarantee of frontier-level reasoning.

For coding, the model can generate code and support code-assistance scenarios. Its low cost may make it useful for routine transformations, documentation, query generation, and developer assistants that can validate outputs automatically. For complex architecture decisions, difficult debugging, or high-stakes code generation, a larger or more specialized model may be more appropriate.

When to choose Command R7B

Choose Command R7B when the application needs a fast, inexpensive text model with a long context window and support for retrieval, tools, and multilingual interactions. It is a practical candidate for internal knowledge assistants, document question answering, enterprise chat, RAG pipelines, code helpers, and agents that need to make several relatively small model calls.

It is especially attractive when throughput and operating cost matter more than achieving the best possible answer on every difficult prompt. A team can use retrieved company documents and tool results to provide current or private information, while Command R7B handles the language-generation part of the workflow.

Another option may be more appropriate when the application requires native image or audio understanding, media generation, unrestricted commercial self-hosting, very advanced reasoning, or consistently stronger performance on complex coding and mathematical tasks. In those cases, the trade-off may favor a larger model, a multimodal model, or a model with a more permissive deployment license.

Overall, Command R7B’s main distinction is not maximum model capability. It is the combination of a 7-billion-parameter design, long context, multilingual text generation, tool and RAG support, low listed token pricing, and an efficiency-oriented deployment profile.


Answers to Frequently Asked Questions

What are the main limitations of Command R7B?
Command R7B is text-only, has a June 1, 2024 knowledge cutoff, and may be less capable than larger models on difficult reasoning, advanced coding, and complex mathematical tasks. Its maximum output is 4,000 tokens, and its open-weight release is subject to a CC-BY-NC license and additional acceptable-use requirements, which may restrict commercial self-hosting.
Is Command R7B suitable for RAG and tool-using agents?
Yes. Command R7B is designed for RAG, tool and function calling, and multi-step agent workflows. Applications can provide retrieved documents or tool results as context, while the application itself executes any requested tools.
How much does Command R7B cost and what is its API model ID?
Cohere lists Command R7B at $0.0375 per million input tokens and $0.15 per million output tokens. Its API model ID is command-r7b-12-2024.
What is Command R7B?
Command R7B is Cohere’s compact, fast language model in the Command R family. It has approximately 7 billion parameters and is designed for cost-efficient text generation, retrieval-augmented generation, enterprise chat, coding assistance, tool calling, and multi-step agent workflows.
What is the context window and maximum output of Command R7B?
Command R7B supports a 128,000-token context window and a maximum generated output of 4,000 tokens. The context window can include prompts, retrieved documents, conversation history, and tool results.


Sources 7
Provider

About Cohere