DeepSeek-OCR

DeepSeek-OCR

by DeepSeek · Current open-weight model; publicly available for local and self-hosted inference

DeepSeek-OCR is an approximately 3B-parameter MIT-licensed vision-language model specialized for image-to-text OCR, document parsing, layout-aware Markdown, figure understanding, and optical context compression. It supports multiple resolution modes and local deployment through Transformers or vLLM, but no official hosted price or native API endpoint is verified for this exact checkpoint.

Text Reasoning Coding
DeepSeek-OCR is a specialized image-to-text model for turning document images into usable text and layout information. It supports OCR, document-to-Markdown conversion, figure parsing, image description, and text-region localization. The checkpoint is available under the MIT license for local or self-hosted inference through Transformers or vLLM, with documented resolution modes that trade image detail against memory use and processing cost.
Outputs

What DeepSeek-OCR can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Streaming
Model profile

Performance characteristics

2/10 Reasoning
2/10 Coding
7/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family DeepSeek-OCR
Model type Multimodal
Context window 8K tokens
Maximum output 8K tokens
Release date 2025-10-20
Status Current open-weight model; publicly available for local and self-hosted inference
Knowledge cutoff notes

No authoritative knowledge-cutoff date was published for the exact DeepSeek-OCR checkpoint. It is an OCR-focused vision-language model whose behavior is primarily determined by the supplied image and prompt.

Model notes

DeepSeek-OCR is a specialized image-to-text model rather than a general-purpose DeepSeek API model. The official checkpoint is approximately 3B parameters, uses BF16 weights, and is distributed under the MIT license. Official examples document Tiny, Small, Base, Large, and dynamic-resolution Gundam modes. The model requires custom model code and is commonly run with Transformers or vLLM. The vLLM integration documents an 8,192-token maximum generation setting and a custom n-gram logits processor for OCR quality. No official hosted DeepSeek API pricing or native DeepSeek API endpoint was found for this exact model. The repository later announced DeepSeek-OCR 2, but the original DeepSeek-OCR checkpoint remains publicly downloadable and usable.

Model guide

DeepSeek-OCR: Open-Weight OCR for Document-to-Text Processing

DeepSeek-OCR is an open-weight, approximately 3-billion-parameter vision-language model from DeepSeek that converts images and documents into text, Markdown, structured layout descriptions, and visual explanations. It is designed for OCR and optical context-compression research rather than general-purpose hosted chat.

What is DeepSeek-OCR?

DeepSeek-OCR is an open-weight vision-language model released by DeepSeek on October 20, 2025. Its main job is to read visual documents and produce text. Rather than serving as a broad conversational assistant, it focuses on optical character recognition (OCR), document parsing, layout-aware Markdown generation, figure interpretation, image description, and locating specified text within an image.

The published checkpoint contains approximately 3 billion parameters and is distributed under the MIT license through DeepSeek's official GitHub repository and Hugging Face model page. This makes it suitable for researchers and developers who want to run the model locally, inspect its deployment behavior, or integrate OCR into a self-managed document pipeline.

DeepSeek-OCR belongs to DeepSeek's open-model ecosystem, but it is distinct from a general-purpose hosted DeepSeek chat or API model. The supplied documentation does not identify an official hosted DeepSeek API endpoint or API price for this exact checkpoint.

How its optical context-compression approach works

DeepSeek-OCR explores an LLM-centric method for visual-text compression. A vision encoder first converts a document image into a relatively compact sequence of visual tokens. The language model then uses that representation to generate OCR text, Markdown, descriptions, or other requested outputs.

In practical terms, the approach attempts to preserve useful written content and layout information without representing every visual detail as a large number of language-model tokens. This is important for document workloads, where pages can contain dense text, tables, headings, figures, and mixed formatting.

Optical context compression is a research and architectural focus, not a guarantee that every document will be compressed without information loss. Results can vary with image quality, page layout, resolution mode, and the prompt used to describe the desired output.

Supported inputs and tasks

The model accepts text prompts together with image input and returns text. Official examples cover several practical workflows:

  • Free-form OCR from document or image pages.
  • Conversion of documents into layout-aware Markdown.
  • OCR with layout grounding.
  • Figure parsing and visual interpretation.
  • Detailed image description.
  • Locating a requested text span inside an image.

These capabilities make the model useful for scanned documents, image-based archives, PDF page extraction, table and layout transcription, and experiments that require both recognized text and an explanation of visual elements. Its output is text; it does not generate images, audio, video, embeddings, or direct machine actions.

Resolution modes and documented limits

DeepSeek-OCR provides several documented image-resolution modes. They are intended to let users choose between visual detail, token usage, memory consumption, and processing speed:

ModeDocumented image configurationPractical role
Tiny512 × 512Lower-detail processing with reduced resource requirements
Small640 × 640Moderate resolution for lighter workloads
Base1024 × 1024Higher-detail general document processing
Large1280 × 1280More visual detail at higher resource cost
GundamMultiple 640 × 640 crops combined with a 1024 × 1024 global viewDynamic-resolution processing for larger or more detailed content

Deployment metadata documents an 8,192-token context length and an 8,192-token maximum generation setting for the vLLM integration. These values should be treated as documented serving limits for the referenced deployment configuration, not as a promise that every image or document will fit comfortably within those limits. Large pages, detailed prompts, and long generated Markdown can consume the available context or output budget.

Local deployment with Transformers or vLLM

DeepSeek-OCR is primarily intended for local or self-hosted inference. The official examples use NVIDIA GPUs with CUDA 11.8, PyTorch 2.6.0, Transformers, custom model code, and optional FlashAttention. The model is also supported by vLLM, including batched image inputs and streaming generation workflows.

Because the checkpoint uses custom model code, deployments generally require trust_remote_code or an equivalent trusted-execution setting. That setting allows downloaded repository code to run as part of model loading, so operators should pin trusted versions and review their deployment process before using it in a production environment.

The official guidance recommends disabling prefix caching and multimodal processor caching for typical OCR workloads. The rationale is that OCR jobs commonly process different images rather than repeatedly reusing the same image and multi-turn prefix. This is a workload-specific recommendation, not a universal requirement for every vLLM deployment.

The model is approximately 3B parameters and uses BF16 weights according to the supplied model information. Actual memory requirements and throughput depend on the selected resolution mode, batching, GPU hardware, attention implementation, image dimensions, and generation settings. No verified universal latency or hardware minimum is provided.

Capabilities, reasoning, coding, and tools

DeepSeek-OCR can perform visual interpretation and follow text instructions about an image, but it is not presented as a general-purpose reasoning model. Its strongest reasoning-like behavior is task-focused: identifying text, preserving document structure, interpreting figures, and locating visual content. It should not be selected primarily for extended mathematical reasoning, open-ended analysis, or general conversational problem solving.

Coding is not a primary capability. Developers can use code to deploy and integrate the checkpoint, but the model itself is designed to read images and produce text rather than generate or execute software. The supplied specifications do not document function calling, tool use, web search, code execution, or agent orchestration. Streaming generation is documented in the vLLM workflow, but streaming only changes how generated text is delivered; it does not add external tools or actions.

Editorially, DeepSeek-OCR has a favorable cost profile because the checkpoint is open-weight and MIT-licensed, while its speed and resource use depend heavily on image resolution and serving hardware. It is less operationally simple than a managed OCR API because users must supply the infrastructure, model runtime, GPU capacity, updates, monitoring, and security controls.

Pricing and API availability

No official hosted price was verified for DeepSeek-OCR itself. The supplied research also does not identify a native DeepSeek API endpoint for this exact checkpoint. Therefore, there is no verified per-token or per-image price to report.

Using the model locally can avoid hosted inference charges, but it is not cost-free. Operators still pay for GPU hardware or rented compute, storage, electricity, engineering time, and maintenance. Whether self-hosting is cheaper depends on document volume, batch size, hardware utilization, and the value of keeping documents inside a controlled environment.

Strengths and limitations

Key strengths

  • Specialized for OCR and document understanding instead of requiring a general chat model to interpret every page.
  • Open-weight checkpoint distributed under the MIT license.
  • Approximately 3B parameters, making it more approachable for self-hosted experimentation than much larger multimodal models.
  • Supports OCR, Markdown conversion, layout grounding, figure parsing, image description, and text-region localization.
  • Offers multiple fixed-resolution modes plus the dynamic-resolution Gundam workflow.
  • Can be deployed with Transformers or vLLM and can support batched and streaming inference through documented serving workflows.
  • Investigates visual-text compression that can make document content more manageable for language-model processing.

Important limitations

  • It requires a GPU-oriented setup and custom model code for the documented deployment paths.
  • OCR accuracy can vary with image quality, typography, page structure, resolution, and prompt selection.
  • It is not a managed commercial OCR service with a verified hosted endpoint or enterprise service-level agreement.
  • It is not designed as a broad chat, advanced reasoning, coding, or tool-using agent model.
  • It produces text rather than images, audio, video, embeddings, or direct actions.
  • The 8,192-token context and generation settings can constrain long documents or verbose Markdown output.
  • Higher-resolution and dynamic-resolution modes can increase memory use and processing cost.

When to choose DeepSeek-OCR

Choose DeepSeek-OCR when the central problem is converting images or documents into searchable, structured, or descriptive text and you are prepared to operate the model yourself. Good fits include local document digitization, searchable-archive creation, scanned PDF extraction, table and layout transcription, Markdown conversion, figure parsing, OCR evaluation, and research into multimodal context efficiency.

It is particularly attractive when an open-weight MIT-licensed model and local processing matter more than a polished hosted workflow. Self-hosting may also be useful for organizations that need to keep source documents within their own infrastructure, although the supplied research does not provide an independent privacy or security certification for the checkpoint.

A managed OCR or multimodal API may be more appropriate when the priority is quick integration, predictable operations, automatic scaling, vendor support, or an enterprise SLA. A general-purpose vision-language model may be a better option when OCR is only one part of a larger workflow involving broad reasoning, coding, web research, tool calls, or conversational assistance. DeepSeek-OCR is best understood as a focused document-processing component, not a replacement for every type of multimodal model.

Bottom line

DeepSeek-OCR is a focused, open-weight model for turning visual documents into text and structured descriptions. Its combination of OCR-specific tasks, layout awareness, resolution choices, MIT licensing, and local deployment support makes it relevant for document pipelines and multimodal research. Its trade-off is operational: users must manage GPU-backed inference and should not expect the hosted API convenience, general reasoning breadth, or tool ecosystem of a full-service multimodal platform.


Answers to Frequently Asked Questions

What are the main limitations of DeepSeek-OCR?
DeepSeek-OCR requires a GPU-oriented deployment and custom model code, and its accuracy depends on image quality, typography, layout, resolution, and prompting. It is focused on document processing rather than general conversation, advanced reasoning, coding, tool use, or agent workflows.
Does DeepSeek-OCR have an official hosted API or pricing?
No verified hosted price or native DeepSeek API endpoint was identified for this specific checkpoint. Local use avoids hosted inference fees but still requires GPU infrastructure, storage, electricity, engineering, and maintenance costs.
What license does DeepSeek-OCR use?
DeepSeek-OCR is distributed under the MIT license. Its approximately 3-billion-parameter checkpoint is available through DeepSeek's official GitHub repository and Hugging Face model page.
What is DeepSeek-OCR used for?
DeepSeek-OCR is an open-weight vision-language model designed to convert visual documents into text. It supports OCR, layout-aware Markdown generation, layout grounding, figure parsing, image description, and locating specified text within an image.
Can DeepSeek-OCR be run locally?
Yes. DeepSeek-OCR is primarily intended for local or self-hosted inference using NVIDIA GPUs, CUDA, PyTorch, Transformers, and optional FlashAttention. It can also be deployed with vLLM for batched and streaming workflows.


Sources 5
Provider

About DeepSeek