What is Dolly v2 12B?
Dolly v2 12B is a 12-billion-parameter, open-weight causal language model created by Databricks. A causal language model generates text by predicting the next token—the next small unit of text—based on the content that came before it. The model can therefore be used for ordinary text-generation tasks such as answering questions, summarizing passages, drafting text, and following written instructions.
The “12B” designation refers to approximately 12 billion model parameters. Parameters are the learned numerical values used by the model to recognize patterns and produce outputs. A larger parameter count does not automatically guarantee better results, but it gives a useful indication of the model’s scale and the computing resources that may be required to run it.
Databricks released Dolly v2 12B on April 12, 2023. It is derived from EleutherAI Pythia-12B and was instruction-tuned using approximately 15,000 human-generated instruction-response records created by Databricks employees. Instruction tuning is an additional training step that teaches a pretrained model to respond more directly to user requests instead of merely continuing text.
Provider and current position
The provider is Databricks, and the canonical model identifier is databricks/dolly-v2-12b on Hugging Face. Dolly v2 12B is not a current hosted Databricks API model in the supplied research. It is better understood as a legacy, downloadable model that users can run independently.
That distinction matters when evaluating availability. The model remains publicly downloadable through its Hugging Face repository, and no official shutdown date was found. However, downloading the weights does not provide managed inference, automatic scaling, an uptime commitment, or a standard hosted-model price. Users are responsible for the environment and computing resources needed for inference.
Key specifications
| Specification | Verified information |
|---|---|
| Model type | 12-billion-parameter causal language model |
| Model family | Dolly v2 |
| Base model | EleutherAI Pythia-12B |
| Instruction tuning | Approximately 15,000 human-generated examples |
| Release date | April 12, 2023 |
| Context length | 2,048 tokens |
| Maximum output tokens | No separate authoritative exact value found |
| Input and output | Text input and text output |
| Weights | Publicly downloadable for self-hosted use |
| Fine-tuning | Supported according to the supplied model data |
The 2,048-token context value is based on the Pythia-derived model configuration. A token is a small piece of text, so the limit is not identical to a 2,048-word limit. The input prompt and generated completion share the model’s available context in practical use, although the supplied sources do not establish a separate exact maximum-generation setting.
What Dolly v2 12B can do
Dolly v2 12B is designed for instruction-following text generation. Suitable tasks include drafting short passages, answering straightforward questions, producing summaries, rewriting text, and exploring conversational or assistant-like behavior in a controlled environment. Because the weights are available, developers can also study the model internally, adapt it to a particular dataset, or experiment with additional instruction tuning.
Its open-weight status is especially useful for organizations that want direct control over model execution. A self-hosted deployment can be examined and customized more directly than a closed hosted service. The released training code and dataset also make Dolly useful for educational work involving instruction tuning and open language-model development.
These advantages should not be confused with frontier-level capability. Databricks’ own documentation identifies weaknesses in complex prompts, programming, mathematical operations, factual accuracy, dates and times, open-ended question answering, list enumeration, stylistic imitation, and hallucination control. A hallucination is an output that sounds plausible but is unsupported or incorrect. Outputs should therefore be checked before they are used for factual, financial, operational, or safety-sensitive decisions.
Modalities, tools, and reasoning
Dolly v2 12B is text-only. It accepts text input and produces text output; the supplied specifications do not indicate image, audio, video, music, embedding, or other non-text output. It also does not have documented built-in web search, function calling, or tool-use support.
The model can produce text that resembles reasoning or step-by-step explanation, but it is not documented as having a specialized reasoning mode. The supplied editorial evaluation rates its reasoning capability at 3 out of 10 and its coding capability at 2 out of 10. Those are comparative editorial scores, not provider-published benchmarks. They reflect the model’s practical limitations rather than a formal Databricks performance guarantee.
Coding is possible as a text-generation task, but the research specifically warns against relying on Dolly v2 12B for serious programming. It also lacks the documented structured-output, JSON-mode, action, and tool-execution features commonly associated with newer hosted models.
Pricing and running costs
Dolly v2 12B has no official hosted API price in the supplied research. The weights are available for self-hosted use, so there is no per-token provider charge listed for calling the model directly. That does not mean running it is cost-free: users may incur costs for suitable hardware, cloud compute, storage, deployment, monitoring, and engineering time.
This creates a different cost profile from a hosted model. Self-hosting can be attractive when a team needs control over execution or wants to experiment without sending prompts to an external inference service. For occasional use, however, the operational work and hardware requirements may outweigh the absence of an API fee. The supplied research does not provide a minimum hardware configuration or a reliable per-request operating cost, so those figures should not be inferred.
Main strengths and limitations
Strengths
- Open availability: The model weights, training code, and dataset were released for commercial use, subject to the applicable licensing terms.
- Self-hosting: Users can run the model independently rather than relying on a current Databricks-hosted endpoint.
- Instruction-following focus: Fine-tuning on human-generated examples makes it more suitable for direct requests than an untuned base language model.
- Research value: Its open artifacts support experimentation, educational projects, and instruction-tuning studies.
- Customization: The supplied specifications identify fine-tuning as supported, allowing users to investigate adaptation to specialized data.
Limitations
- Short context: The 2,048-token context is restrictive for long documents, extended conversations, and large code files.
- Outdated capability level: Databricks describes the model as not state of the art, and it lacks many capabilities associated with modern managed models.
- Reliability concerns: Factual mistakes, hallucinations, weak handling of dates and times, and poor performance on complex prompts require human review.
- Weak programming and mathematics: The model is not a strong choice for dependable code generation, numerical work, or technical problem solving.
- No managed API features: There is no supplied evidence of official hosted pricing, web search, function calling, streaming, batch inference, or structured JSON output.
- Operational responsibility: Self-hosting transfers infrastructure, security, scaling, and maintenance responsibilities to the user.
When to choose this model
Choose Dolly v2 12B when the primary requirement is an openly downloadable instruction-tuned model for local experimentation, research, teaching, or customization. It can also make sense when commercial-use availability and direct control over the model are more important than leading accuracy, long context, or managed-service convenience.
For example, a developer studying instruction tuning could use the released model and dataset as a practical starting point. A research team could compare local adaptation techniques without building its work around a proprietary endpoint. An organization with suitable infrastructure could evaluate a self-hosted text model for low-risk internal experiments.
Another option is more appropriate when the task requires reliable factual answers, advanced reasoning, strong mathematical or programming performance, image or audio understanding, long documents, structured outputs, tool calling, or a supported production API. Newer hosted models generally offer more convenient deployment and stronger capabilities, while smaller local models may provide a better speed-and-cost trade-off when Dolly’s 12-billion-parameter scale is unnecessary.
Overall assessment
Dolly v2 12B remains a useful example of an openly released instruction-tuned language model, but it should be evaluated as a legacy research and self-hosting option rather than as a current general-purpose assistant. Its central benefit is control: users can download, inspect, run, and adapt the model. Its central trade-off is capability and reliability. The limited context, lack of built-in tools, absence of a hosted API price, and documented weaknesses in factual accuracy, programming, mathematics, and complex instructions make careful scope selection essential.
For learning, experimentation, and model-customization work, those trade-offs may be acceptable. For dependable production assistance or demanding reasoning tasks, a newer model with stronger evaluation results, managed inference, and modern tool or structured-output support is likely to be a better fit.

