DBRX

DBRX Instruct

by Databricks Data + AI Platform · Open-weight model; Databricks Foundation Model API serving retired

DBRX Instruct is Databricks’ instruction-tuned open-weight language model with 132 billion total parameters, approximately 36 billion active parameters, and a 32,768-token context window. It supports text-only question answering, generation, summarization, and coding assistance. Databricks retired its Foundation Model API serving in 2025, so the model is now primarily relevant for self-hosted or independently managed deployments.

Text Reasoning Coding
DBRX Instruct is the instruction-following version of Databricks’ DBRX foundation model. It accepts text and produces text for tasks such as answering questions, summarizing information, generating prose, and writing code. The model remains relevant as an open-weight checkpoint for independently managed inference, but it should not be treated as a currently available Databricks-hosted endpoint: its Foundation Model API serving was retired during 2025.
Outputs

What DBRX Instruct can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Fine-tuning
Model profile

Performance characteristics

7/10 Reasoning
8/10 Coding
5/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family DBRX
Model type General Purpose
Context window 33K tokens
Release date 2024-03-27
Status Open-weight model; Databricks Foundation Model API serving retired
Deprecation date 2025-04-30
Shutdown date 2025-12-19
Knowledge cutoff notes

No authoritative model-specific knowledge cutoff was identified in the reviewed Databricks model documentation.

Model notes

DBRX Instruct is the instruction-tuned member of the DBRX family and is distinct from DBRX Base. It has 132 billion total parameters with approximately 36 billion active per token, uses 16 experts with 4 selected per token, and was trained on 12 trillion tokens of text and code. The canonical checkpoint was released under the Databricks Open Model License. Databricks retired DBRX and DBRX Instruct from Foundation Model APIs pay-per-token serving on April 30, 2025 and from provisioned-throughput serving on December 19, 2025. The open-weight model can still be independently deployed, subject to its license and hardware requirements. Editorial scores are comparative estimates rather than vendor specifications.

Cost

Model pricing

Input 10.714 DBU per 1 million input tokens when historically available through Databricks pay-per-token serving; endpoint retired April 30, 2025
Output 32.143 DBU per 1 million output tokens when historically available through Databricks pay-per-token serving; endpoint retired April 30, 2025
Model guide

DBRX Instruct: Databricks’ Open-Weight Mixture-of-Experts Model

DBRX Instruct is Databricks’ instruction-tuned, open-weight language model. Its mixture-of-experts architecture has 132 billion total parameters, approximately 36 billion active parameters per token, and a 32,768-token context window. It is designed for English-language text generation, question answering, coding assistance, and self-hosted deployment, although Databricks’ hosted Foundation Model API availability ended in 2025.

What is DBRX Instruct?

DBRX Instruct is an instruction-tuned large language model from Databricks. Instruction tuning means the model was adapted to respond to user requests rather than simply predict the next token in a document. In practical terms, it is intended to follow prompts, answer questions, generate text, and assist with programming tasks.

It is part of the DBRX family and is distinct from DBRX Base, which is the corresponding base language model. Databricks released DBRX Instruct as an open-weight model in March 2024 under the Databricks Open Model License, subject to the associated acceptable-use policy. The model can therefore be downloaded and deployed with compatible infrastructure instead of being limited to a proprietary hosted interface.

DBRX Instruct is aimed primarily at technical teams, researchers, and organizations that want more control over model hosting, data handling, fine-tuning, or serving infrastructure. It is not a consumer chatbot product and is not currently a new Databricks-managed pay-per-token endpoint.

Architecture and context window

DBRX Instruct uses a decoder-only transformer architecture with a mixture-of-experts, or MoE, design. It contains 132 billion total parameters, but approximately 36 billion are active for each input token. This allows the model to use a large parameter pool without running every parameter for every token, although the complete checkpoint and inference system still require substantial resources.

The model has 16 experts and selects four experts for each token. Its architecture also includes grouped-query attention, rotary position encodings, and gated linear units. Databricks reports that the model was trained on 12 trillion tokens of text and code. Its maximum context length is 32,768 tokens, allowing a prompt and conversation history of roughly 32K tokens in supported deployments.

The supplied model documentation does not identify a model-specific knowledge cutoff, and no maximum output-token limit is established in the available research. Output capacity can therefore depend on the serving implementation and remaining context space rather than on a confirmed universal limit.

Capabilities and supported modalities

DBRX Instruct is text-only. It accepts text input and returns text output; it does not natively accept images, audio, or video, and it does not generate images, speech, music, or video. This makes it suitable for language and programming workflows but unsuitable for multimodal document interpretation or media generation without additional models and processing systems.

Its intended tasks include general English-language question answering, instruction following, summarization, drafting, text transformation, and code generation. Databricks also identifies coding assistance as a supported use case. The model was trained on text and code, but that does not mean it can execute code or verify that generated programs work.

Databricks’ model documentation states that DBRX models do not provide native function calling or code execution. A deployment can potentially be connected to external tools through application code, but the model itself should not be treated as a tool-using agent with built-in function support. Native structured-output or JSON-mode support is also not established by the supplied research.

Reasoning, coding, speed, and cost

DBRX Instruct can perform ordinary multi-step language reasoning and is useful for breaking down questions, producing explanations, and transforming requirements into text or code. However, it is not documented as a dedicated reasoning model, and the available research does not provide a specialized reasoning mode or verified benchmark results. Its reasoning performance should therefore be evaluated through the specific prompts and workloads that matter to the buyer.

Coding is one of its stronger practical areas. It can generate software snippets, explain code, help draft implementations, and assist with debugging conversations. Its lack of native code execution means that generated code must be tested separately. Teams requiring an integrated coding agent should add an execution and tool layer or choose a model and platform with verified built-in tool support.

The editorial assessment supplied for this page rates coding more favorably than reasoning, while assigning a middle-range estimate for speed and cost. These are comparative editorial judgments, not Databricks-published scores. The main practical trade-off is scale: DBRX Instruct’s MoE design reduces the number of active parameters per token, but the 132-billion-parameter checkpoint remains demanding to host compared with smaller open-weight language models. Quantized community versions may reduce hardware requirements, but they are derivative distributions rather than the canonical Databricks checkpoint.

Deployment and current availability

DBRX Instruct can be used as a self-hosted or independently managed model where the operator provides compatible inference software, hardware, monitoring, and storage. Full-precision deployment requires substantial memory and multi-GPU infrastructure according to Databricks’ documentation. The exact requirements vary with precision, batching, sequence length, quantization, and the serving stack, so the research does not establish one universal hardware specification.

Databricks previously made DBRX Instruct available through Model Serving and Foundation Model APIs. Its pay-per-token Foundation Model API availability ended on April 30, 2025. Provisioned-throughput serving ended on December 19, 2025. As a result, new users should not plan around a current Databricks-hosted DBRX Instruct endpoint without independently verifying a replacement or an alternative deployment route.

Historically, the listed pay-per-token rates were 10.714 DBU per one million input tokens and 32.143 DBU per one million output tokens. These were usage rates rather than a subscription price, and they should not be interpreted as current availability or current pricing because the relevant hosted endpoint has been retired. Self-hosting introduces different costs, including GPUs, storage, networking, engineering, and operations.

Main strengths and limitations

  • Open-weight flexibility: Organizations can manage deployment independently and can assess self-hosting or domain-specific adaptation under the Databricks Open Model License.
  • Large context: The 32,768-token context window supports relatively long instructions, documents, and coding conversations where the serving setup exposes the full limit.
  • Strong coding orientation: Text-and-code training makes code generation and software-development assistance central use cases.
  • Large MoE architecture: Approximately 36 billion active parameters per token provide a more selective computation pattern than activating the entire 132-billion-parameter model for every token.
  • High infrastructure demands: The full model is considerably more difficult to host than smaller open-weight alternatives, especially without quantization.
  • No native tools or execution: Function calling and code execution are not built-in capabilities according to the model documentation.
  • Text-only operation: It cannot directly analyze images, audio, or video or create non-text media.
  • Retired hosted serving: Databricks’ documented Foundation Model API endpoints are no longer a current route for new hosted inference.

Best use cases for DBRX Instruct

DBRX Instruct is most appropriate when a team specifically values an open-weight, text-based model and has the infrastructure or deployment partner needed to operate it. Suitable uses include:

  • Self-hosted English-language question answering and text generation
  • Code drafting, explanation, transformation, and software-development assistance
  • Summarization and long-form text processing within the 32K context limit
  • Private or independently managed inference environments
  • Domain-specific experimentation or fine-tuning where the license and hardware requirements are acceptable
  • Research into mixture-of-experts language-model deployment

It is a less suitable choice for a lightweight personal chatbot, image or document-vision workflows, speech applications, real-time media generation, or applications that require built-in function calling and reliable code execution. A smaller contemporary open-weight model may be more practical when low latency, lower GPU cost, or simple deployment matters more than DBRX Instruct’s scale. A hosted model with active provider support may also be preferable when the team does not want to operate inference infrastructure.

When to choose DBRX Instruct

Choose DBRX Instruct when the primary requirement is a capable, text-only, open-weight model for self-managed language or coding workloads, and the organization can support a demanding deployment. Its 132-billion-parameter MoE structure, 36-billion active-parameter profile, and 32K context make it technically substantial, but those characteristics do not automatically make it the fastest or cheapest option for every application.

Choose another option when you need a currently supported Databricks-hosted endpoint, multimodal input, native tools, verified structured output, code execution, or a smaller operational footprint. DBRX Instruct’s continued value is mainly in its downloadable checkpoint and the control that independently managed deployment can provide, not in current first-party API convenience.


Answers to Frequently Asked Questions

Who should choose DBRX Instruct?
DBRX Instruct is best suited to technical teams, researchers, and organizations that need an open-weight, text-based model for self-managed question answering, long-context processing, coding assistance, or domain-specific experimentation. It is less suitable for lightweight deployments, multimodal applications, built-in tool use, code execution, or teams seeking a currently supported managed endpoint.
How can DBRX Instruct be deployed, and is Databricks hosted serving still available?
DBRX Instruct can be self-hosted or independently managed using compatible inference software and substantial hardware, typically involving multi-GPU infrastructure for full-precision deployment. Databricks’ pay-per-token Foundation Model API availability ended on April 30, 2025, and provisioned-throughput serving ended on December 19, 2025, so new users should verify any replacement service before relying on hosted inference.
Does DBRX Instruct support multimodal inputs, function calling, or code execution?
No. DBRX Instruct is text-only and does not natively accept images, audio, or video. Its documentation also does not provide native function calling or code execution. External application and tool layers can be added, but generated code must be tested separately.
What is DBRX Instruct?
DBRX Instruct is an instruction-tuned, open-weight large language model from Databricks for text generation, question answering, summarization, and coding assistance. It is part of the DBRX family and is distinct from the base model DBRX Base.
What are the main technical specifications of DBRX Instruct?
DBRX Instruct uses a decoder-only mixture-of-experts architecture with 132 billion total parameters and approximately 36 billion active parameters per token. It has 16 experts, selects four experts per token, was trained on 12 trillion text and code tokens, and supports a maximum context length of 32,768 tokens.


Sources 7
Provider

About Databricks Data + AI Platform