Nemotron Parse

NVIDIA Nemotron Parse 2.0

by NVIDIA AI · current

NVIDIA Nemotron Parse 2.0 is a specialized open-weight vision-language model that converts document images into structured text with semantic classes, bounding boxes, reading order, tables, charts, and multilingual OCR support. It is intended for document-intelligence pipelines rather than general conversation, coding, or media generation.

Text Reasoning Coding
NVIDIA Nemotron Parse 2.0 is built for document intelligence rather than general conversation. Given an RGB image of a scanned page or rendered document, it can extract text while identifying where that text and other elements appear on the page. Its output is intended for OCR augmentation, document indexing, retrieval pipelines, table and chart extraction, data curation, and agentic workflows that need a machine-readable representation of complex documents.
Outputs

What NVIDIA Nemotron Parse 2.0 can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Nemotron Parse
Model type Multimodal
Release date 2026-08-03
Status current
Knowledge cutoff notes

No authoritative knowledge-cutoff date was identified for this document-parsing model. It is designed for visual document analysis rather than general world-knowledge question answering.

Model notes

NVIDIA Nemotron Parse 2.0 is the current Parse model listed by NVIDIA; NVIDIA identifies v1.2 as the prior version. The hosted NVIDIA Build endpoint uses the lowercase identifier nvidia/nemotron-parse, while the open-weight repository uses nvidia/NVIDIA-Nemotron-Parse-2.0. The model converts document images into structured text with semantic classes, bounding boxes, reading order, tables, and charts. The model card specifies OpenMDW-1.1 for the model and configuration files and CC-BY-4.0 for the tokenizer. Public model-specific token pricing was not verified.

Model guide

NVIDIA Nemotron Parse 2.0: Open-Weight Document Parsing for OCR and Layout Understanding

NVIDIA Nemotron Parse 2.0 is an open-weight vision-language model specialized in turning document images into structured, machine-readable text. It combines OCR with semantic element classification, bounding boxes, reading order, table structure, chart information, and multilingual document support.

What is NVIDIA Nemotron Parse 2.0?

NVIDIA Nemotron Parse 2.0 is an open-weight vision-language model for document understanding and parsing. A vision-language model accepts visual input and uses learned language and document representations to describe or structure what it sees. In this case, the model's purpose is not to answer broad questions about an image, but to convert a document page into organized text and layout information.

The model is designed for scanned documents, rendered PDFs, forms, presentations, tables, charts, and pages containing mixed content. Alongside extracted text, it can identify semantic elements such as titles, paragraphs, captions, tables, charts, page headers, page footers, footnotes, pictures, and bibliography entries. It also returns spatial annotations, including bounding boxes, so downstream software can connect each extracted element to its position on the original page.

NVIDIA lists Nemotron Parse 2.0 as the current Parse model, with NVIDIA Nemotron Parse v1.2 identified as the preceding version. The model is available through NVIDIA resources, including the official model listing and NVIDIA's Hugging Face organization.

What problem does it solve?

Traditional OCR can turn pixels into text, but plain OCR output often loses the relationships that make a document useful. A table may become a sequence of disconnected words, a chart may be treated as an ordinary image, and headers or footnotes may be mixed into the body text. Nemotron Parse 2.0 is intended to preserve more of this document structure.

A typical workflow starts with an RGB image of a document page. The model analyzes the page, extracts readable text, classifies visible elements, and provides layout annotations and reading-order information. A downstream application can then use the result to create a searchable index, populate an extraction system, prepare content for retrieval-augmented generation, or send uncertain fields to a human reviewer.

This makes the model particularly relevant when the location and type of content matter. For example, a document-processing system may need to distinguish a table from the paragraph above it, retain a chart as a separate object, or avoid treating a page footer as part of the main answer text.

Core capabilities

  • Document OCR: extracts text from scanned or rendered document images.
  • Semantic classification: identifies document components such as paragraphs, titles, tables, charts, captions, headers, footers, footnotes, pictures, and bibliography entries.
  • Spatial grounding: associates recognized elements with bounding boxes on the source page.
  • Reading order: represents the natural sequence in which page elements should be read.
  • Table and chart handling: supports documents where tabular and chart-based information is central to the page.
  • Multilingual parsing: includes expanded support for CJK and Indic scripts, according to the supplied model research.
  • Handwriting extraction: NVIDIA highlights improved handwritten-text extraction compared with the earlier Parse version.

These capabilities make Parse 2.0 more than a text-recognition component. It is better understood as a structured document parser that uses visual context to organize the text it finds.

Input and output modalities

The primary input is an RGB document image. This can represent a scanned page, a rendered PDF page, a form, a presentation slide, or another document image. The model is categorized as multimodal because it accepts visual input and produces text based on that visual content.

Its output is text containing formatted content and spatial annotations. The model does not generate images, audio, or video. It should therefore not be evaluated as a media-generation system or as a general-purpose conversational assistant.

Although its output is structured and may include predictable document fields, the supplied research does not verify a separate general-purpose JSON-schema or JSON-mode capability. Applications that require strict schemas should validate and transform the model's output in their own processing layer rather than assuming that structured document output is equivalent to a provider-guaranteed JSON mode.

What changed from the earlier Parse version?

Compared with NVIDIA Nemotron Parse v1.2, version 2.0 adds an approximately 20,000-token vocabulary expansion intended to improve multilingual support. NVIDIA also describes a chart-aware document-parsing class and updated training coverage for chart- and table-heavy documents.

The supplied research also identifies improved handwritten-text extraction and table-structure recovery as changes in the newer version. These improvements are relevant to document collections that contain more than clean, digitally generated paragraphs. However, they should be treated as provider or model-card claims rather than as a guarantee of accuracy on every handwriting style, scan quality, chart format, or table layout.

Technical limits and unavailable specifications

Several conventional language-model specifications have not been verified for Nemotron Parse 2.0. The supplied research does not provide a context-window length, maximum output-token limit, model-specific token pricing, or a separate hosted-service quota. These values should not be inferred from the model's vocabulary expansion or from the limits of a particular inference server.

SpecificationVerified information
Model typeVision-language document-parsing model
Primary inputRGB document images
Text outputYes, including formatted content and spatial annotations
Image, audio, or video outputNo
Context lengthNot verified in the supplied research
Maximum output tokensNot verified in the supplied research
Model-specific pricingNot publicly verified in the supplied research
General JSON modeNot verified; structured output should not be treated as a guaranteed JSON-schema mode

Deployment constraints will also depend on the selected inference stack, hardware, image resolution, batching configuration, and service endpoint. Those operational limits are separate from the model's intrinsic capabilities and should be measured in the environment where the model will run.

Reasoning, coding, and tool support

Nemotron Parse 2.0 is optimized for visual document analysis rather than open-ended reasoning. It can interpret relationships between page elements, reading order, tables, and charts, but that should not be confused with a general reasoning model intended for complex planning or broad knowledge questions.

Its primary output is extracted and organized document content, not software code. The supplied research therefore rates coding usefulness as low and does not identify it as a coding model. It is also not documented here as having built-in function calling, browsing, external tools, or agent actions. A larger pipeline can place tools around the model—for example, an indexer, database, validator, or human-review queue—but those are application-level integrations rather than verified native model features.

Deployment and licensing

The model is available as an NVIDIA-published open-weight model through NVIDIA resources and the NVIDIA Hugging Face organization. The supplied research identifies compatible deployment paths including Transformers, vLLM, SGLang, and containerized NVIDIA NIM workflows. The exact setup, performance, hardware requirements, and production support will vary by implementation.

The Hugging Face model card identifies the OpenMDW License Agreement, version 1.1, for the model and configuration files. The included tokenizer is identified as being under CC-BY-4.0. Teams should review the applicable license documents before using the model in a commercial product, redistributing weights, or combining it with other components.

The hosted NVIDIA Build endpoint uses the identifier nvidia/nemotron-parse, while the open-weight repository uses nvidia/NVIDIA-Nemotron-Parse-2.0. These identifiers refer to access and packaging conventions; they should not be assumed to represent different primary models without checking the relevant documentation.

Strengths and trade-offs

The model's main strength is specialization. A general vision-language model may be able to describe a page, but Nemotron Parse 2.0 is specifically designed to produce document-oriented structure: semantic classes, bounding boxes, reading order, tables, charts, and formatted text. That specialization can reduce the amount of layout reconstruction required in a document pipeline.

Its open-weight availability is another practical advantage for teams that need more control over deployment, data handling, or integration with their own inference infrastructure. It can be considered for local or managed deployment rather than being limited to a single consumer chat interface.

The trade-off is scope. This is not a general assistant, a text-only language model, an embedding model, or a content-generation system. It may also require careful validation on low-quality scans, unusual layouts, handwriting, and specialized tables. Open weights do not remove the need for suitable hardware, image preprocessing, output validation, monitoring, and licensing review.

Best use cases

  • Extracting text and layout from scanned records and rendered PDFs.
  • Building searchable document indexes that preserve headings, tables, captions, and page locations.
  • Preparing documents for retrieval-augmented generation while reducing confusion between body text, footnotes, headers, and tables.
  • Recovering table and chart information from reports, presentations, and business documents.
  • Creating training or evaluation data for document-intelligence systems.
  • Supporting human-in-the-loop extraction, where uncertain fields or pages are routed for review.
  • Parsing multilingual collections that include CJK or Indic scripts.

When should you choose Nemotron Parse 2.0?

Choose Nemotron Parse 2.0 when the central problem is converting visually complex documents into structured text and layout metadata. It is a particularly logical candidate when tables, charts, reading order, bounding boxes, multilingual pages, or handwriting matter more than conversational fluency.

A general-purpose vision-language model may be more appropriate when the application needs open-ended image questions, broad reasoning, visual explanation, or interactive conversation. A conventional OCR engine may be preferable for a narrowly defined, high-volume text-recognition task if semantic layout and chart awareness are unnecessary. A dedicated embedding model is the more suitable component when the immediate requirement is vector search rather than document parsing.

Nemotron Parse 2.0 should also be tested against representative pages before production adoption. Compare extraction quality on clean scans, difficult scans, multi-column layouts, tables, charts, handwritten content, and the languages in the target collection. Because no verified public model-specific pricing or universal context and output limits are supplied here, cost and throughput should be measured using the intended deployment path rather than estimated from generic NVIDIA or language-model pricing.

Bottom line

NVIDIA Nemotron Parse 2.0 is a focused document-intelligence model. Its value lies in combining OCR with page structure, semantic labels, spatial grounding, reading order, table handling, chart awareness, and multilingual improvements. It is a strong candidate for document ingestion and analysis pipelines, but it should not be treated as a general chatbot, coding model, media generator, or automatically schema-compliant API. The right evaluation question is whether it preserves the structure and meaning needed by the downstream document workflow.


Answers to Frequently Asked Questions

When should you choose NVIDIA Nemotron Parse 2.0?
Choose NVIDIA Nemotron Parse 2.0 when document structure, tables, charts, reading order, bounding boxes, multilingual content, or handwriting are important. It is suitable for searchable indexes, retrieval-augmented generation pipelines, document extraction, and human-in-the-loop review. A general-purpose vision-language model may be better for open-ended visual questions, while conventional OCR may be preferable for simple high-volume text recognition.
What input and output formats does NVIDIA Nemotron Parse 2.0 support?
The primary input is an RGB image of a document page, such as a scanned page or rendered PDF page. The output is text containing formatted content and spatial annotations. The model does not generate images, audio, or video, and a general-purpose guaranteed JSON-schema mode has not been verified.
How does NVIDIA Nemotron Parse 2.0 differ from traditional OCR?
Traditional OCR primarily extracts text, while NVIDIA Nemotron Parse 2.0 is designed to preserve document structure. It can distinguish elements such as tables, charts, headers, footnotes, and body paragraphs, while also returning their locations and intended reading order.
What is NVIDIA Nemotron Parse 2.0 used for?
NVIDIA Nemotron Parse 2.0 is used to convert scanned documents, rendered PDFs, forms, presentations, tables, charts, and mixed-content pages into structured text and layout information. It can identify semantic elements, preserve reading order, and associate extracted content with bounding boxes.
What are the main capabilities of NVIDIA Nemotron Parse 2.0?
Its main capabilities include document OCR, semantic classification of titles, paragraphs, tables, charts, captions, headers, footers, footnotes, pictures, and bibliography entries, spatial grounding with bounding boxes, reading-order detection, table and chart handling, multilingual parsing, and improved handwriting extraction.


Sources 5
Provider

About NVIDIA AI