Dragon-DocChat

Cerebras Dragon-DocChat

by Cerebras · Available as open-weight query and context encoder checkpoints

Cerebras Dragon-DocChat is an open-weight dual-encoder model for conversational document retrieval. It creates embeddings for questions and document passages, helping RAG systems find relevant evidence before a separate language model generates an answer. The model is specialized for retrieval rather than chat, coding, reasoning, or tool use. Cerebras reports improved top-1 recall over Dragon+ and ChatQA Dragon-Multiturn, but no hosted token price or exact context limit is supplied.

Embeddings Reasoning Coding
Cerebras Dragon-DocChat is not a general-purpose chatbot or text-generation model. It is a specialized retrieval component built for conversational document question answering and retrieval-augmented generation (RAG). Cerebras provides it as separate query-encoder and context-encoder checkpoints: one encodes a user’s question, while the other encodes document passages. A search system compares those embeddings to identify the most relevant passages. The retrieved text can then be passed to a separate answer-generation model, such as the associated Cerebras Llama3-DocChat system.
Outputs

What Cerebras Dragon-DocChat can produce

Embeddings
Inputs

What it can understand

Text
Capabilities

Supported features

Fine-tuning
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Dragon-DocChat
Model type Other
Release date 2024-08-21
Status Available as open-weight query and context encoder checkpoints
Model notes

Cerebras Dragon-DocChat is a dual-encoder embedding retriever built on Dragon+ and fine-tuned with ChatQA conversational question-answering data using contrastive loss with hard negatives. The model is distributed as separate query-encoder and context-encoder checkpoints, which are used together for document retrieval. Cerebras reported absolute top-1 recall improvements of 8.9% over Dragon+ and 3.5% over ChatQA Dragon-Multiturn. It is an open-weight research model rather than a Cerebras hosted token-priced API model. The exact model does not generate natural-language answers; the associated Cerebras Llama3-DocChat model performs answer generation.

Model guide

Cerebras Dragon-DocChat: A Dual-Encoder Retriever for Conversational Document Search

Cerebras Dragon-DocChat is an open-weight dual-encoder retrieval model for multi-turn document question answering. It converts questions and document passages into embeddings so a retrieval system can find relevant evidence before a separate language model generates an answer.

What is Cerebras Dragon-DocChat?

Cerebras Dragon-DocChat is an open-weight retrieval model released by Cerebras on August 21, 2024. Its purpose is to improve document retrieval for conversational question-answering systems, especially when a user’s question depends on earlier turns in the conversation.

The model is a dual encoder. Rather than generating a response directly, it converts two types of text into numerical representations called embeddings:

  • A query encoder represents the user’s question, including the conversational context supplied by the application.
  • A context encoder represents passages or documents that may contain the answer.

A retrieval pipeline compares the query embedding with document embeddings and ranks passages by similarity. The highest-ranked passages can then be supplied to a separate text-generation model. This division of labor makes Dragon-DocChat a search and retrieval model, not a standalone conversational assistant.

How it fits into the Cerebras catalog

Dragon-DocChat belongs to Cerebras’s DocChat research work and is distributed through open model checkpoints rather than as a token-priced hosted model in the supplied materials. Cerebras describes it as a model for conversational question answering over documents. The official release includes separate query-encoder and context-encoder checkpoints, available through Cerebras’s model resources and GitHub repository.

Its role is narrower than Cerebras’s hosted inference offerings, which provide access to language models for chat completion and other generation tasks. Dragon-DocChat is instead intended to be assembled into a retrieval-augmented generation architecture. In a typical implementation, an application indexes documents with the context encoder, encodes each new user question with the query encoder, retrieves the most relevant passages, and sends those passages to an answer-generation model.

Training and retrieval design

Cerebras built Dragon-DocChat on Dragon+ and fine-tuned it with conversational question-answering data from ChatQA. The training approach uses contrastive learning with hard negatives. In practical terms, the model is trained to place a question close to passages that answer it and farther away from plausible but incorrect passages.

Hard negatives are particularly useful for document search because an incorrect passage may share many words with the question while still failing to answer it. Training against these difficult alternatives is intended to improve ranking quality rather than simple keyword overlap.

According to Cerebras, Dragon-DocChat produced an absolute top-1 recall improvement of 8.9% over Dragon+ and 3.5% over ChatQA Dragon-Multiturn. These are provider-reported benchmark claims, and they should not be treated as a guarantee for every document collection, language, chunking strategy, or retrieval implementation.

Inputs, outputs, and modalities

Dragon-DocChat accepts text input and produces embeddings. It does not directly produce natural-language answers, images, audio, or video. The model therefore has no conventional chat response, image generation, speech, or multimodal output capability in the supplied specifications.

CapabilityDragon-DocChat
Primary outputText embeddings for retrieval
Text inputYes
Image, audio, or video inputNo verified support
Natural-language text generationNo
Tool or function callingNo
Structured response generationNo
Web searchNo

The exact context length and maximum output-token limit are not specified in the supplied research. Since the output is an embedding rather than generated prose, a conventional maximum output-token figure is not applicable. Implementers should check the model card and checkpoint configuration before selecting document chunk sizes or designing an indexing pipeline.

Main strengths and trade-offs

The model’s main strength is specialization. A dedicated conversational retriever can be a better fit for multi-turn document search than a general language model asked to perform retrieval through prompting. Dragon-DocChat is designed to represent both questions and passages in a way that supports ranking, and it can be incorporated into systems where the document index and answer-generation layer are controlled by the developer.

Its open-weight availability is another practical advantage for teams that need to run retrieval in their own environment or customize the surrounding pipeline. The supplied research also rates its speed and cost characteristics favorably in editorial scoring, with a speed score of 8 out of 10 and a cost score of 9 out of 10. Those scores are editorial evaluations, not Cerebras-published guarantees, and actual performance depends on hardware, indexing design, batch size, and the volume of documents.

The trade-off is that Dragon-DocChat cannot complete the full question-answering workflow by itself. It retrieves evidence but does not explain that evidence to the user. A separate generation model, prompt, orchestration layer, and usually a document-processing pipeline are required. This adds system complexity compared with using one general-purpose model for direct question answering.

Reasoning, coding, and tool support

Dragon-DocChat is not evaluated or presented as a reasoning model, coding assistant, or autonomous agent. Its task is relevance estimation: determining which document passages best match a conversational query. It does not expose a chat-style reasoning process, generate code, call tools, or execute functions according to the supplied specifications.

It can still support applications that involve technical documentation or code repositories because those are document-retrieval use cases. However, the model should be responsible for finding relevant passages, not for interpreting code, writing a patch, or deciding which external action to perform. Those responsibilities belong to downstream models and application logic.

Pricing and availability

No input or output token price is provided for Dragon-DocChat. The supplied research identifies it as an open-weight research model rather than a Cerebras-hosted, token-priced API model. Consequently, there is no verified per-request, per-million-token, or subscription price to report for this model.

Using it may still involve infrastructure costs. A team may need compute for creating document embeddings, storing the index, serving retrieval requests, and running a separate answer-generation model. The financial advantage of an open checkpoint depends on the deployment environment and request volume; it should not be confused with free end-to-end question answering.

When to choose Dragon-DocChat

Dragon-DocChat is a good candidate when the main problem is retrieving evidence from a changing or specialized document collection, particularly when questions refer to earlier turns in a conversation. Examples include internal knowledge bases, technical documentation, policy collections, product manuals, and research archives used in a RAG system.

  • Choose it when you need a dedicated retriever rather than a model that directly writes answers.
  • Choose it when open-weight deployment and control over the indexing pipeline are important.
  • Choose it when multi-turn question context matters and keyword search alone is insufficient.
  • Choose it when retrieval quality and infrastructure cost matter more than having a single all-in-one assistant.

Another option may be more appropriate when the application needs direct conversational responses, tool calling, code generation, web access, or multimodal understanding. A general-purpose language model is also simpler for small collections or prototypes where building and maintaining a separate retrieval layer would not justify the added complexity. Conversely, the associated Llama3-DocChat answer-generation model is the relevant component when the system needs to turn retrieved evidence into natural-language responses; Dragon-DocChat itself remains the retrieval component.

Bottom line

Cerebras Dragon-DocChat is best understood as an open-weight search model for conversational RAG, not as a chatbot. Its dual-encoder design separates question understanding from document representation, allowing applications to retrieve passages before invoking a separate generator. Cerebras reports meaningful top-1 recall improvements over the cited Dragon+ and ChatQA Dragon-Multiturn baselines, but the model has no verified hosted price, context limit, or direct answer-generation capability in the supplied information. Its value is greatest for developers building controlled, document-grounded retrieval systems rather than users looking for a standalone assistant.


Answers to Frequently Asked Questions

Is Cerebras Dragon-DocChat available through a paid API, and what does it cost?
The supplied information describes Dragon-DocChat as an open-weight research model rather than a token-priced Cerebras-hosted API model. No verified input or output token price is provided, although deployment may still require infrastructure, indexing, storage, and separate answer-generation costs.
What are the reported retrieval improvements of Dragon-DocChat?
Cerebras reports an absolute top-1 recall improvement of 8.9% over Dragon+ and 3.5% over ChatQA Dragon-Multiturn. These are provider-reported benchmark results and may vary depending on the documents, language, chunking strategy, and retrieval implementation.
Can Dragon-DocChat generate natural-language answers or function as a chatbot?
No. Dragon-DocChat produces text embeddings for retrieval and does not directly generate answers, call tools, write code, or provide web or multimodal capabilities. A separate answer-generation model and application pipeline are required.
What is Cerebras Dragon-DocChat used for?
Cerebras Dragon-DocChat is an open-weight dual-encoder retrieval model for conversational document search and retrieval-augmented generation (RAG). It finds relevant passages from documents so a separate language model can use them to generate an answer.
How does the Dragon-DocChat dual-encoder architecture work?
Dragon-DocChat uses a query encoder to represent a user’s question and conversational context, and a context encoder to represent document passages. The system compares their embeddings and ranks passages by relevance.


Sources 5
Provider

About Cerebras