Dolly v2

Dolly v2 12B

by Databricks Data + AI Platform · Legacy open-weight model; publicly downloadable; no official shutdown date found

Dolly v2 12B is a commercially usable, open-weight causal language model released by Databricks in April 2023. Based on Pythia-12B and fine-tuned on the Databricks Dolly 15K dataset, it is intended for local inference, research, and customization rather than state-of-the-art general-purpose performance.

Text Reasoning Coding
Dolly v2 12B is an openly downloadable language model from Databricks that can follow instructions, answer questions, summarize text, and support other text-generation tasks. Its main distinction is accessibility: the model weights, training code, and dataset were released for commercial use, allowing organizations and researchers to run or adapt it themselves instead of depending on a hosted API. It is now a legacy model and is not presented as state of the art, but it remains relevant for experimentation and instruction-tuning studies.
Outputs

What Dolly v2 12B can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Fine-tuning
Model profile

Performance characteristics

3/10 Reasoning
2/10 Coding
4/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Dolly v2
Model type General Purpose
Context window 2K tokens
Release date 2023-04-12
Status Legacy open-weight model; publicly downloadable; no official shutdown date found
Knowledge cutoff notes

No authoritative provider-published knowledge-cutoff date was found for Dolly v2 12B. The instruction-tuning examples were generated during March and April 2023, but that does not establish the pretrained model's knowledge cutoff.

Model notes

Canonical Hugging Face identifier: databricks/dolly-v2-12b. The model is a 12-billion-parameter causal language model derived from EleutherAI Pythia-12B and fine-tuned on approximately 15,000 human-generated instruction-response records from Databricks employees. Databricks released the model weights, training code, and dataset for commercial use. The model is self-hosted rather than a current Databricks hosted API offering. Official documentation identifies weaknesses in complex prompts, programming, mathematical operations, factual accuracy, dates and times, open-ended question answering, hallucination, list enumeration, and stylistic imitation. The 2,048-token context value is based on the Pythia-derived model configuration; a separate authoritative exact maximum-generation value was not found.

Cost

Model pricing

Input No official hosted API price; model weights are available for self-hosted use
Output No official hosted API price; model weights are available for self-hosted use
Model guide

Dolly v2 12B: Databricks’ Open-Weight Instruction-Tuned Model

Dolly v2 12B is a 12-billion-parameter open-weight causal language model released by Databricks in April 2023. Derived from EleutherAI Pythia-12B and fine-tuned on approximately 15,000 human-generated instruction-response examples, it is intended for local inference, research, customization, and commercial use rather than frontier-level general-purpose performance.

What is Dolly v2 12B?

Dolly v2 12B is a 12-billion-parameter, open-weight causal language model created by Databricks. A causal language model generates text by predicting the next token—the next small unit of text—based on the content that came before it. The model can therefore be used for ordinary text-generation tasks such as answering questions, summarizing passages, drafting text, and following written instructions.

The “12B” designation refers to approximately 12 billion model parameters. Parameters are the learned numerical values used by the model to recognize patterns and produce outputs. A larger parameter count does not automatically guarantee better results, but it gives a useful indication of the model’s scale and the computing resources that may be required to run it.

Databricks released Dolly v2 12B on April 12, 2023. It is derived from EleutherAI Pythia-12B and was instruction-tuned using approximately 15,000 human-generated instruction-response records created by Databricks employees. Instruction tuning is an additional training step that teaches a pretrained model to respond more directly to user requests instead of merely continuing text.

Provider and current position

The provider is Databricks, and the canonical model identifier is databricks/dolly-v2-12b on Hugging Face. Dolly v2 12B is not a current hosted Databricks API model in the supplied research. It is better understood as a legacy, downloadable model that users can run independently.

That distinction matters when evaluating availability. The model remains publicly downloadable through its Hugging Face repository, and no official shutdown date was found. However, downloading the weights does not provide managed inference, automatic scaling, an uptime commitment, or a standard hosted-model price. Users are responsible for the environment and computing resources needed for inference.

Key specifications

SpecificationVerified information
Model type12-billion-parameter causal language model
Model familyDolly v2
Base modelEleutherAI Pythia-12B
Instruction tuningApproximately 15,000 human-generated examples
Release dateApril 12, 2023
Context length2,048 tokens
Maximum output tokensNo separate authoritative exact value found
Input and outputText input and text output
WeightsPublicly downloadable for self-hosted use
Fine-tuningSupported according to the supplied model data

The 2,048-token context value is based on the Pythia-derived model configuration. A token is a small piece of text, so the limit is not identical to a 2,048-word limit. The input prompt and generated completion share the model’s available context in practical use, although the supplied sources do not establish a separate exact maximum-generation setting.

What Dolly v2 12B can do

Dolly v2 12B is designed for instruction-following text generation. Suitable tasks include drafting short passages, answering straightforward questions, producing summaries, rewriting text, and exploring conversational or assistant-like behavior in a controlled environment. Because the weights are available, developers can also study the model internally, adapt it to a particular dataset, or experiment with additional instruction tuning.

Its open-weight status is especially useful for organizations that want direct control over model execution. A self-hosted deployment can be examined and customized more directly than a closed hosted service. The released training code and dataset also make Dolly useful for educational work involving instruction tuning and open language-model development.

These advantages should not be confused with frontier-level capability. Databricks’ own documentation identifies weaknesses in complex prompts, programming, mathematical operations, factual accuracy, dates and times, open-ended question answering, list enumeration, stylistic imitation, and hallucination control. A hallucination is an output that sounds plausible but is unsupported or incorrect. Outputs should therefore be checked before they are used for factual, financial, operational, or safety-sensitive decisions.

Modalities, tools, and reasoning

Dolly v2 12B is text-only. It accepts text input and produces text output; the supplied specifications do not indicate image, audio, video, music, embedding, or other non-text output. It also does not have documented built-in web search, function calling, or tool-use support.

The model can produce text that resembles reasoning or step-by-step explanation, but it is not documented as having a specialized reasoning mode. The supplied editorial evaluation rates its reasoning capability at 3 out of 10 and its coding capability at 2 out of 10. Those are comparative editorial scores, not provider-published benchmarks. They reflect the model’s practical limitations rather than a formal Databricks performance guarantee.

Coding is possible as a text-generation task, but the research specifically warns against relying on Dolly v2 12B for serious programming. It also lacks the documented structured-output, JSON-mode, action, and tool-execution features commonly associated with newer hosted models.

Pricing and running costs

Dolly v2 12B has no official hosted API price in the supplied research. The weights are available for self-hosted use, so there is no per-token provider charge listed for calling the model directly. That does not mean running it is cost-free: users may incur costs for suitable hardware, cloud compute, storage, deployment, monitoring, and engineering time.

This creates a different cost profile from a hosted model. Self-hosting can be attractive when a team needs control over execution or wants to experiment without sending prompts to an external inference service. For occasional use, however, the operational work and hardware requirements may outweigh the absence of an API fee. The supplied research does not provide a minimum hardware configuration or a reliable per-request operating cost, so those figures should not be inferred.

Main strengths and limitations

Strengths

  • Open availability: The model weights, training code, and dataset were released for commercial use, subject to the applicable licensing terms.
  • Self-hosting: Users can run the model independently rather than relying on a current Databricks-hosted endpoint.
  • Instruction-following focus: Fine-tuning on human-generated examples makes it more suitable for direct requests than an untuned base language model.
  • Research value: Its open artifacts support experimentation, educational projects, and instruction-tuning studies.
  • Customization: The supplied specifications identify fine-tuning as supported, allowing users to investigate adaptation to specialized data.

Limitations

  • Short context: The 2,048-token context is restrictive for long documents, extended conversations, and large code files.
  • Outdated capability level: Databricks describes the model as not state of the art, and it lacks many capabilities associated with modern managed models.
  • Reliability concerns: Factual mistakes, hallucinations, weak handling of dates and times, and poor performance on complex prompts require human review.
  • Weak programming and mathematics: The model is not a strong choice for dependable code generation, numerical work, or technical problem solving.
  • No managed API features: There is no supplied evidence of official hosted pricing, web search, function calling, streaming, batch inference, or structured JSON output.
  • Operational responsibility: Self-hosting transfers infrastructure, security, scaling, and maintenance responsibilities to the user.

When to choose this model

Choose Dolly v2 12B when the primary requirement is an openly downloadable instruction-tuned model for local experimentation, research, teaching, or customization. It can also make sense when commercial-use availability and direct control over the model are more important than leading accuracy, long context, or managed-service convenience.

For example, a developer studying instruction tuning could use the released model and dataset as a practical starting point. A research team could compare local adaptation techniques without building its work around a proprietary endpoint. An organization with suitable infrastructure could evaluate a self-hosted text model for low-risk internal experiments.

Another option is more appropriate when the task requires reliable factual answers, advanced reasoning, strong mathematical or programming performance, image or audio understanding, long documents, structured outputs, tool calling, or a supported production API. Newer hosted models generally offer more convenient deployment and stronger capabilities, while smaller local models may provide a better speed-and-cost trade-off when Dolly’s 12-billion-parameter scale is unnecessary.

Overall assessment

Dolly v2 12B remains a useful example of an openly released instruction-tuned language model, but it should be evaluated as a legacy research and self-hosting option rather than as a current general-purpose assistant. Its central benefit is control: users can download, inspect, run, and adapt the model. Its central trade-off is capability and reliability. The limited context, lack of built-in tools, absence of a hosted API price, and documented weaknesses in factual accuracy, programming, mathematics, and complex instructions make careful scope selection essential.

For learning, experimentation, and model-customization work, those trade-offs may be acceptable. For dependable production assistance or demanding reasoning tasks, a newer model with stronger evaluation results, managed inference, and modern tool or structured-output support is likely to be a better fit.


Answers to Frequently Asked Questions

How much does it cost to run Dolly v2 12B?
There is no official hosted API price listed for Dolly v2 12B in the supplied research. Because the model is self-hosted, users do not pay a provider per-token fee for direct inference, but they may incur costs for hardware or cloud compute, storage, deployment, monitoring, security, and engineering.
What are the main capabilities and limitations of Dolly v2 12B?
Dolly v2 12B can generate text, answer straightforward questions, summarize, rewrite content, follow instructions, and support experimentation or fine-tuning. Its limitations include a 2,048-token context, weak performance on complex prompts, programming, mathematics, factual accuracy, dates and times, and open-ended question answering. It also lacks documented built-in web search, function calling, structured JSON output, and specialized reasoning features.
What is Dolly v2 12B?
Dolly v2 12B is a 12-billion-parameter, open-weight causal language model created by Databricks. It is based on EleutherAI Pythia-12B and was instruction-tuned with approximately 15,000 human-generated instruction-response examples.
Can Dolly v2 12B be used through a hosted API?
Dolly v2 12B is not identified as a current hosted Databricks API model. Its weights are publicly downloadable from the Hugging Face repository under the identifier databricks/dolly-v2-12b, so users are responsible for hosting, computing resources, scaling, and maintenance.


Sources 3
Provider

About Databricks Data + AI Platform