Nemotron OCR

Nemotron OCR v1

by NVIDIA AI · Available; English-only OCR model; newer Nemotron OCR v2 is available for updated English and multilingual OCR deployments

NVIDIA Nemotron OCR v1 is an English-only image OCR model that combines text detection, Transformer recognition, and relational layout analysis. It processes PNG and JPEG images, supports batches, and returns recognized text, confidence scores, bounding boxes, and reading-order information. It is designed for GPU-accelerated document ingestion, enterprise search, retrieval-augmented generation, and document intelligence. No fixed public usage price, universal context limit, or maximum output limit is identified in the supplied research. Nemotron OCR v2 is the more suitable NVIDIA option for multilingual OCR.

Text Reasoning Coding
Nemotron OCR v1 is an image-to-text model in NVIDIA’s NeMo Retriever collection. Rather than generating general-purpose prose or answering questions, it identifies English text in images and returns structured OCR results that can preserve useful information about regions, confidence, layout, and reading order. Its combination of a RegNetY-8GF detector, Transformer recognizer, and relational analysis module makes it suitable for document ingestion, enterprise search, retrieval-augmented generation, and other pipelines that need image-based content converted into searchable data.
Outputs

What Nemotron OCR v1 can produce

Text
Inputs

What it can understand

Images
Capabilities

Supported features

Structured output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Nemotron OCR
Model type Other
Release date 2025-10-23
Status Available; English-only OCR model; newer Nemotron OCR v2 is available for updated English and multilingual OCR deployments
Knowledge cutoff notes

NVIDIA does not publish a conventional textual knowledge cutoff for this image OCR model. Its behavior is based on trained OCR, recognition, and layout-analysis data rather than a general-purpose language-model knowledge boundary.

Model notes

Canonical model repository is nvidia/nemotron-ocr-v1. The model combines a RegNetY-8GF text detector, Transformer recognizer, and relational model for layout and reading-order analysis. NVIDIA reports 52,467,240 total parameters. It accepts RGB PNG/JPEG images and can process single images or batches, with word, sentence, and paragraph aggregation levels. Outputs contain recognized text, confidence scores, and bounding-box coordinates. The model card states that it supports English only. The model is commercially usable under the NVIDIA Open Model License Agreement, while post-processing scripts use Apache 2.0. NVIDIA documentation continues to list v1 as an accessible OCR model, while Nemotron OCR v2 is the newer English and multilingual option. Editorial scores are relative to the broader AI-model market and are not vendor-provided benchmarks.

Model guide

NVIDIA Nemotron OCR v1: English Document OCR with Layout and Reading-Order Analysis

Nemotron OCR v1 is NVIDIA’s commercially usable, English-only optical character recognition model for extracting text, bounding boxes, confidence scores, and document-layout relationships from scanned documents, photographs, charts, tables, receipts, infographics, and natural-scene images.

What is Nemotron OCR v1?

Nemotron OCR v1 is an optical character recognition model from NVIDIA’s NeMo Retriever collection. OCR converts text shown in an image into machine-readable text. In this model’s case, the result is not limited to a plain string: the pipeline can also return confidence scores, bounding-box coordinates, and information useful for understanding how text regions relate to one another on a page.

The model is intended for English OCR rather than general-purpose language generation. It can process scanned documents, photographed pages, receipts, business records, charts, tables, infographics, and natural-scene images containing text. Its outputs can then be indexed for search, passed into a retrieval-augmented generation system, or used by document-processing software.

NVIDIA released Nemotron OCR v1 on October 23, 2025. It remains an available English OCR option in NVIDIA’s model and deployment catalog, although Nemotron OCR v2 is the newer choice when multilingual support or a more recent OCR implementation is required.

How the model works

Nemotron OCR v1 combines three main functions in one end-to-end pipeline. First, a text detector finds likely text regions in the image. NVIDIA identifies the detector backbone as RegNetY-8GF, a convolutional neural-network architecture used to analyze visual patterns and locate text.

Next, a Transformer-based recognizer reads the detected regions and transcribes their contents. Transformers are neural-network components that can model relationships across a sequence, which is useful when recognizing words and lines rather than treating each character as an isolated mark.

A relational module then analyzes connections among text regions. This helps the system reason about logical groupings, reading order, blocks, columns, tables, charts, and other page structures. NVIDIA reports approximately 52.5 million total parameters across the detector, recognizer, and relational components, with the model components trained jointly as an OCR pipeline.

Supported inputs and outputs

The documented input is an RGB image in PNG or JPEG format. Nemotron OCR v1 accepts either uint8 pixel values, commonly used for standard image data, or float32 values. A single image can be represented as a 3 × H × W tensor, while batched inference uses B × 3 × H × W, where H and W represent image height and width and B represents the batch size.

The pipeline supports single-image and batched processing. Results can be aggregated at word, sentence, or paragraph level, allowing an implementation to select a granularity appropriate to its downstream task. For example, word-level results may be useful for precise coordinate-based highlighting, while paragraph-level aggregation can be more convenient for search indexing or retrieval.

Outputs include recognized text, confidence scores, and bounding-box coordinates. The layout-related processing can additionally support reading-order and block analysis. This is important for pages where simply reading pixels from left to right would produce incorrect results, such as multi-column reports, tables, forms, and infographics.

CapabilityVerified information
Primary inputRGB PNG or JPEG images
Input formatsuint8 or float32 image values
ProcessingSingle images or batches
Text coverageEnglish only
OutputRecognized text, confidence scores, and bounding boxes
AggregationWord, sentence, or paragraph levels
Layout informationReading order and relationships among text blocks

Where it fits in NVIDIA’s lineup

Nemotron OCR v1 is a specialized model rather than a general-purpose Nemotron conversational model. It is positioned within NVIDIA’s document and retrieval tooling, including NeMo Retriever and NVIDIA Inference Microservices. NVIDIA provides downloadable model artifacts and deployment options for GPU-accelerated serving.

The model is also listed in NVIDIA’s OCR documentation and can be used in document-ingestion and retrieval workflows. Its role is to turn visual documents into structured text and layout data before another system performs search, retrieval, summarization, classification, or question answering.

Nemotron OCR v2 is the newer related option identified in the supplied documentation. The important practical distinction is language coverage: v1 is English-only, while v2 provides English and multilingual OCR variants. Existing systems may still choose v1 for compatibility, established pipelines, or a specific deployment requirement, but new projects that need languages beyond English should evaluate v2 instead.

Strengths and practical use cases

Nemotron OCR v1’s main strength is that it goes beyond basic image-to-text transcription. It is designed to retain information about where text appears and how separate regions relate to one another. That makes it more useful for structured document processing than an OCR system that returns only an undifferentiated text string.

  • Document ingestion: Convert scanned reports, invoices, receipts, and business records into searchable text while retaining region coordinates.
  • Enterprise search: Extract text from image-based files so that documents can be indexed alongside digitally generated content.
  • RAG preprocessing: Prepare visual documents for retrieval-augmented generation systems by producing text and layout-aware metadata before indexing.
  • Tables and charts: Process text distributed across structured visual elements where reading order and block relationships matter.
  • Agentic document workflows: Supply downstream software with recognized text, confidence values, and coordinates for validation or human review.
  • Natural-scene text: Recognize English text in photographs and other images outside conventional office documents.

The relational output is particularly useful when document structure affects meaning. A report with two columns, a receipt with separate fields, or an infographic with multiple labeled regions can require more than simply concatenating every detected word in coordinate order.

Accuracy and performance information

NVIDIA’s model card reports several internal evaluation results: a character error rate of 0.1633, bag-of-character error rate of 0.0453, bag-of-word error rate of 0.1203, table-extraction TEDS of 0.781, and multimodal retrieval Recall@5 scores of 0.779 on public earnings data and 0.901 on digital corpora.

These are provider-reported internal evaluations, not universal guarantees. OCR results vary with image resolution, lighting, page quality, typography, occlusion, language, and document design. The figures should therefore be treated as reference points from NVIDIA’s testing rather than as directly comparable results for every OCR workload.

Nemotron OCR v1 is designed for NVIDIA GPU-accelerated environments. A local implementation requires a suitable CUDA software stack and compatible hardware configuration. The supplied research does not specify a universal inference speed, minimum GPU, maximum image dimensions, context window, or maximum output-token limit. Those values should be verified against the particular NVIDIA deployment path and version being used.

Limitations and unsupported tasks

The clearest limitation is language coverage. Nemotron OCR v1 supports English only, so it is not an appropriate standalone solution for multilingual archives, documents containing substantial non-English text, or deployments that must automatically recognize many writing systems. Nemotron OCR v2 is the more relevant NVIDIA option when multilingual support is a requirement.

Image quality also matters. Blurred photographs, low contrast, unusual fonts, severe perspective distortion, partially hidden text, and highly stylized layouts can reduce recognition quality. Confidence scores can help identify uncertain results, but they do not remove the need for validation in high-stakes workflows such as financial records, legal documents, or identity-related processing.

This is not a general-purpose text-generation model. It is not documented as a conversational assistant, coding model, visual question-answering system, speech model, image generator, video model, or reasoning model. It also has no verified tool or function-calling capability in the supplied research. Its structured output describes OCR findings; it does not indicate a general JSON-mode or agent-tool interface.

Pricing, licensing, and access

No public per-image, per-token, or subscription price is provided in the supplied research. The model can be accessed through NVIDIA model artifacts and NVIDIA deployment surfaces, but the cost of using it will depend on the chosen infrastructure, GPU capacity, serving method, and any applicable NVIDIA platform or enterprise terms. It would be inaccurate to describe Nemotron OCR v1 as having a fixed consumer API price based on the available information.

NVIDIA states that the model is commercially usable under the NVIDIA Open Model License Agreement. The accompanying post-processing scripts are licensed under Apache 2.0. Organizations should review the applicable license terms and deployment documentation before incorporating the model into a commercial product.

When to choose Nemotron OCR v1

Choose Nemotron OCR v1 when the primary requirement is English text extraction from images and the workflow benefits from bounding boxes, confidence scores, reading order, or layout relationships. It is a good fit for teams already operating NVIDIA GPU infrastructure and for pipelines that need OCR as a preprocessing stage for search, retrieval, RAG, or document intelligence.

Another OCR system may be more appropriate when multilingual recognition is essential, when deployment must run outside an NVIDIA-oriented environment, or when a managed service with clearly published usage pricing is preferred. A general-purpose vision-language model may be a better choice when the task requires answering questions about an image, summarizing visual content, or combining OCR with open-ended reasoning rather than producing dedicated OCR structures.

Within NVIDIA’s current OCR family, Nemotron OCR v2 deserves evaluation for new multilingual deployments. Nemotron OCR v1 remains a practical choice for English-only workloads where its existing model artifacts, layout-aware outputs, or compatibility with an established NeMo Retriever or NIM pipeline are more important than adopting the newer version.

Bottom line

Nemotron OCR v1 is a specialized, English-only OCR model focused on turning complex images into structured text and document-layout information. Its detector-recognizer-relational design makes it more suitable for document ingestion and retrieval pipelines than for conversational use. The most important selection questions are whether English-only coverage is sufficient, whether NVIDIA GPU deployment fits the environment, and whether the application needs layout-aware OCR rather than general visual reasoning.


Answers to Frequently Asked Questions

What are the main limitations and deployment requirements of Nemotron OCR v1?
Nemotron OCR v1 is English-only, can be affected by poor image quality or complex typography, and is not a general-purpose conversational or visual reasoning model. It is designed for NVIDIA GPU-accelerated environments and requires a compatible CUDA software stack and hardware configuration. No fixed public usage price is specified in the supplied information.
What are the main use cases for Nemotron OCR v1?
The model is designed for document ingestion, enterprise search, retrieval-augmented generation preprocessing, table and chart processing, agentic document workflows, and recognition of English text in natural-scene images. Its layout-aware outputs are useful for multi-column documents, forms, receipts, and infographics.
Is Nemotron OCR v1 multilingual?
No. Nemotron OCR v1 supports English only. NVIDIA Nemotron OCR v2 is the newer option for English and multilingual OCR workloads.
What is NVIDIA Nemotron OCR v1?
NVIDIA Nemotron OCR v1 is an English-only optical character recognition model from the NeMo Retriever collection. It extracts text from images and can also return confidence scores, bounding-box coordinates, reading order, and relationships among text regions.
What image formats and outputs does Nemotron OCR v1 support?
Nemotron OCR v1 accepts RGB images in PNG or JPEG format with uint8 or float32 values. It supports single-image and batched processing and can produce recognized text, confidence scores, bounding boxes, and layout information at word, sentence, or paragraph level.


Sources 5
Provider

About NVIDIA AI