What is DBRX Base?
DBRX Base is a pretrained, decoder-only large language model provided by Databricks. “Base” means that it is the general pretrained model rather than a version primarily optimized for following conversational instructions. Its core job is next-token prediction: given a sequence of text, it generates a continuation. That makes it suitable for text completion, code completion, research, and as a starting point for additional fine-tuning.
The model uses a mixture-of-experts architecture. Instead of using every parameter for every token, DBRX Base contains 132 billion total parameters and activates approximately 36 billion parameters per input token. It uses 16 experts and activates four experts per token. This design gives the model a very large overall capacity while limiting the amount of computation used for each individual token compared with a dense model containing the same total number of parameters.
DBRX Base was released by Databricks on March 27, 2024. Its training data has a stated knowledge cutoff of December 2023. That cutoff describes the information learned during pretraining; it does not give the model web access or automatically provide current information.
Current availability and positioning
DBRX Base now occupies a different position from a current hosted commercial model. Databricks' research identifies the open weights as remaining available for self-hosted use, while DBRX access through Databricks Foundation Model APIs had a retirement date of December 19, 2025. The DBRX fine-tuning family was retired on April 30, 2025. The model was also previously available through limited provisioned-throughput Model Serving.
As a result, there is no current official hosted token price to quote for DBRX Base after retirement. Users who deploy the open weights themselves must account for infrastructure, storage, operations, and engineering costs rather than paying a simple per-token API rate. Those costs depend on the chosen hardware and serving setup, and are not specified in the supplied research.
This makes DBRX Base most relevant to organizations and researchers that value control over model execution, customization, or reproducibility. It is less suitable for someone looking for the simplest way to call a maintained, low-latency language model through a current hosted API.
Verified specifications
| Specification | DBRX Base |
|---|---|
| Provider | Databricks |
| Model type | Decoder-only mixture-of-experts language model |
| Total parameters | 132 billion |
| Active parameters per token | Approximately 36 billion |
| Experts | 16 total; 4 activated per token |
| Context window | 32,768 tokens |
| Training knowledge cutoff | December 2023 |
| Input | Text |
| Output | Text |
| Open-weight use | Available for self-hosted use |
| Fine-tuning | Full-parameter and LoRA support documented for the open model; Databricks-hosted fine-tuning retired |
The 32,768-token context window is the maximum documented context length in the supplied research. A token is a fragment of text rather than necessarily a complete word, so the practical amount of readable text varies by language and formatting. The research does not provide a separate maximum output-token limit. The context window should therefore not be interpreted as a guaranteed output length.
What DBRX Base does well
Large capacity with sparse activation
DBRX Base combines a 132-billion-parameter model with mixture-of-experts routing. The model does not activate all 132 billion parameters for every token; it routes each token through four of its 16 experts. This is the central architectural distinction and helps explain why the model can offer substantial capacity without performing the full dense computation on every token.
For users with the infrastructure and expertise to serve it, this design is useful for experimenting with large-scale language-model behavior, custom inference systems, and fine-tuning approaches. It is not, however, a guarantee of lower total deployment cost: a model with 132 billion total parameters still has significant memory and operational requirements.
Useful as a customization and research base
The official repository documents both full-parameter and LoRA fine-tuning. LoRA, or low-rank adaptation, lets a user train a smaller set of additional parameters instead of changing the entire base model. That can make experimentation more manageable, although the underlying model still has substantial serving and storage requirements.
DBRX Base is therefore a practical candidate for teams investigating domain adaptation, model behavior, code completion, or custom text-generation workflows. Its open-weight availability also gives users more control than a closed hosted endpoint over where inference runs and how the model is integrated.
A substantial text context
With a 32K-token context window, DBRX Base can process considerably more text than short-context models. Potential applications include long code files, technical documents, sizeable prompts, and multi-part completion tasks. The model remains text-only: the supplied specifications do not support images, audio, or video as inputs or outputs.
Important limitations
Self-hosting is a significant commitment
DBRX Base is not a lightweight model intended for casual local use. Its total parameter count and mixture-of-experts design imply substantial infrastructure requirements, even though only part of the network is active for each token. The supplied research does not specify minimum GPU memory, a recommended hardware configuration, or a guaranteed tokens-per-second rate, so those details should be tested against the intended serving stack rather than assumed.
The model's editorial speed score is 4 out of 10, while its editorial cost score is 7 out of 10. These are comparative editorial estimates, not Databricks-published benchmarks or prices. They indicate a model that may offer reasonable value for its capabilities in the right deployment, but is unlikely to match smaller models for latency or simplicity.
Text-only and without native tool support
DBRX Base accepts text and produces text. It does not natively process images, audio, or video according to the supplied specifications. Its tool-use and function-calling field is also marked unsupported. A surrounding application could parse generated text and connect it to tools, but that would be an application-level integration rather than a verified native model capability.
Not a current Databricks-hosted production choice
The retirement of DBRX Foundation Model API access and hosted fine-tuning changes the practical recommendation. A new team seeking a managed endpoint, current service-level expectations, or straightforward usage-based billing should evaluate another currently supported option. DBRX Base can still be valuable when open weights and self-managed deployment matter more than managed availability.
Reasoning, coding, and output behavior
DBRX Base is a general-purpose pretrained completion model, not a reasoning-specialized model in the supplied data. Its editorial reasoning score is 6 out of 10, but this is an internal comparative assessment rather than a provider-published reasoning benchmark. It should not be described as having a dedicated chain-of-thought mode or advanced reasoning feature.
The editorial coding score is 7 out of 10. That assessment, together with the model's text-completion design, makes code completion and programming-oriented generation reasonable use cases. However, the score is subjective and does not establish performance on a particular coding benchmark. Users should validate the model on their own programming languages, repository style, and completion format.
Streaming is listed as supported, which can allow generated text to be delivered incrementally by a compatible serving implementation. Structured output is not listed as supported, and JSON mode is marked unavailable in the supplied data. Generated JSON may still be attempted through prompting or post-processing, but reliable schema-constrained generation should not be assumed.
Best use cases
- Self-hosted English text completion: organizations that need to run a large language model in infrastructure they control.
- Code completion experiments: teams evaluating a large pretrained model for code generation or continuation.
- Fine-tuning research: researchers using full-parameter or LoRA adaptation to explore a domain or task.
- Model architecture research: work involving sparse mixture-of-experts models and routing behavior.
- Long text prompts: tasks that benefit from a context window of up to 32,768 tokens and do not require non-text inputs.
These uses assume that the operator can manage model deployment and accepts the absence of current Databricks-hosted API access.
When to choose DBRX Base
Choose DBRX Base when open weights, self-hosting, customization, or research access are central requirements. It is particularly defensible when a team has existing infrastructure and wants a large mixture-of-experts base model that can be adapted with LoRA or full-parameter fine-tuning.
Choose a different type of model when deployment simplicity is more important. A smaller hosted language model is likely to be more appropriate for low-latency applications, modest infrastructure budgets, or straightforward API integration. A current managed model is also a better fit when the project needs supported production availability rather than an open-weight artifact. A multimodal model is necessary for image, audio, or video inputs, and a tool-oriented model or platform is preferable when native function calling and structured action execution are essential.
DBRX Base can also be the wrong choice for applications that need reliable JSON schemas, native tool use, current world knowledge, or a clearly supported commercial endpoint. Its December 2023 training cutoff and lack of web search mean that current information must come from an external retrieval or application layer, if such a layer is added.
Bottom line
DBRX Base is best understood as an open-weight, large-scale text-completion model for self-hosting and experimentation, not as a current general-purpose Databricks API product. Its 132-billion-parameter mixture-of-experts architecture, 32K context window, and documented fine-tuning support give it meaningful value for model research and custom deployments. The trade-off is operational: it is text-only, has no verified native tool or structured-output support, lacks a current official hosted price, and requires users to take responsibility for infrastructure and maintenance.

