What is Nemotron OCR v1?
Nemotron OCR v1 is an optical character recognition model from NVIDIA’s NeMo Retriever collection. OCR converts text shown in an image into machine-readable text. In this model’s case, the result is not limited to a plain string: the pipeline can also return confidence scores, bounding-box coordinates, and information useful for understanding how text regions relate to one another on a page.
The model is intended for English OCR rather than general-purpose language generation. It can process scanned documents, photographed pages, receipts, business records, charts, tables, infographics, and natural-scene images containing text. Its outputs can then be indexed for search, passed into a retrieval-augmented generation system, or used by document-processing software.
NVIDIA released Nemotron OCR v1 on October 23, 2025. It remains an available English OCR option in NVIDIA’s model and deployment catalog, although Nemotron OCR v2 is the newer choice when multilingual support or a more recent OCR implementation is required.
How the model works
Nemotron OCR v1 combines three main functions in one end-to-end pipeline. First, a text detector finds likely text regions in the image. NVIDIA identifies the detector backbone as RegNetY-8GF, a convolutional neural-network architecture used to analyze visual patterns and locate text.
Next, a Transformer-based recognizer reads the detected regions and transcribes their contents. Transformers are neural-network components that can model relationships across a sequence, which is useful when recognizing words and lines rather than treating each character as an isolated mark.
A relational module then analyzes connections among text regions. This helps the system reason about logical groupings, reading order, blocks, columns, tables, charts, and other page structures. NVIDIA reports approximately 52.5 million total parameters across the detector, recognizer, and relational components, with the model components trained jointly as an OCR pipeline.
Supported inputs and outputs
The documented input is an RGB image in PNG or JPEG format. Nemotron OCR v1 accepts either uint8 pixel values, commonly used for standard image data, or float32 values. A single image can be represented as a 3 × H × W tensor, while batched inference uses B × 3 × H × W, where H and W represent image height and width and B represents the batch size.
The pipeline supports single-image and batched processing. Results can be aggregated at word, sentence, or paragraph level, allowing an implementation to select a granularity appropriate to its downstream task. For example, word-level results may be useful for precise coordinate-based highlighting, while paragraph-level aggregation can be more convenient for search indexing or retrieval.
Outputs include recognized text, confidence scores, and bounding-box coordinates. The layout-related processing can additionally support reading-order and block analysis. This is important for pages where simply reading pixels from left to right would produce incorrect results, such as multi-column reports, tables, forms, and infographics.
| Capability | Verified information |
|---|---|
| Primary input | RGB PNG or JPEG images |
| Input formats | uint8 or float32 image values |
| Processing | Single images or batches |
| Text coverage | English only |
| Output | Recognized text, confidence scores, and bounding boxes |
| Aggregation | Word, sentence, or paragraph levels |
| Layout information | Reading order and relationships among text blocks |
Where it fits in NVIDIA’s lineup
Nemotron OCR v1 is a specialized model rather than a general-purpose Nemotron conversational model. It is positioned within NVIDIA’s document and retrieval tooling, including NeMo Retriever and NVIDIA Inference Microservices. NVIDIA provides downloadable model artifacts and deployment options for GPU-accelerated serving.
The model is also listed in NVIDIA’s OCR documentation and can be used in document-ingestion and retrieval workflows. Its role is to turn visual documents into structured text and layout data before another system performs search, retrieval, summarization, classification, or question answering.
Nemotron OCR v2 is the newer related option identified in the supplied documentation. The important practical distinction is language coverage: v1 is English-only, while v2 provides English and multilingual OCR variants. Existing systems may still choose v1 for compatibility, established pipelines, or a specific deployment requirement, but new projects that need languages beyond English should evaluate v2 instead.
Strengths and practical use cases
Nemotron OCR v1’s main strength is that it goes beyond basic image-to-text transcription. It is designed to retain information about where text appears and how separate regions relate to one another. That makes it more useful for structured document processing than an OCR system that returns only an undifferentiated text string.
- Document ingestion: Convert scanned reports, invoices, receipts, and business records into searchable text while retaining region coordinates.
- Enterprise search: Extract text from image-based files so that documents can be indexed alongside digitally generated content.
- RAG preprocessing: Prepare visual documents for retrieval-augmented generation systems by producing text and layout-aware metadata before indexing.
- Tables and charts: Process text distributed across structured visual elements where reading order and block relationships matter.
- Agentic document workflows: Supply downstream software with recognized text, confidence values, and coordinates for validation or human review.
- Natural-scene text: Recognize English text in photographs and other images outside conventional office documents.
The relational output is particularly useful when document structure affects meaning. A report with two columns, a receipt with separate fields, or an infographic with multiple labeled regions can require more than simply concatenating every detected word in coordinate order.
Accuracy and performance information
NVIDIA’s model card reports several internal evaluation results: a character error rate of 0.1633, bag-of-character error rate of 0.0453, bag-of-word error rate of 0.1203, table-extraction TEDS of 0.781, and multimodal retrieval Recall@5 scores of 0.779 on public earnings data and 0.901 on digital corpora.
These are provider-reported internal evaluations, not universal guarantees. OCR results vary with image resolution, lighting, page quality, typography, occlusion, language, and document design. The figures should therefore be treated as reference points from NVIDIA’s testing rather than as directly comparable results for every OCR workload.
Nemotron OCR v1 is designed for NVIDIA GPU-accelerated environments. A local implementation requires a suitable CUDA software stack and compatible hardware configuration. The supplied research does not specify a universal inference speed, minimum GPU, maximum image dimensions, context window, or maximum output-token limit. Those values should be verified against the particular NVIDIA deployment path and version being used.
Limitations and unsupported tasks
The clearest limitation is language coverage. Nemotron OCR v1 supports English only, so it is not an appropriate standalone solution for multilingual archives, documents containing substantial non-English text, or deployments that must automatically recognize many writing systems. Nemotron OCR v2 is the more relevant NVIDIA option when multilingual support is a requirement.
Image quality also matters. Blurred photographs, low contrast, unusual fonts, severe perspective distortion, partially hidden text, and highly stylized layouts can reduce recognition quality. Confidence scores can help identify uncertain results, but they do not remove the need for validation in high-stakes workflows such as financial records, legal documents, or identity-related processing.
This is not a general-purpose text-generation model. It is not documented as a conversational assistant, coding model, visual question-answering system, speech model, image generator, video model, or reasoning model. It also has no verified tool or function-calling capability in the supplied research. Its structured output describes OCR findings; it does not indicate a general JSON-mode or agent-tool interface.
Pricing, licensing, and access
No public per-image, per-token, or subscription price is provided in the supplied research. The model can be accessed through NVIDIA model artifacts and NVIDIA deployment surfaces, but the cost of using it will depend on the chosen infrastructure, GPU capacity, serving method, and any applicable NVIDIA platform or enterprise terms. It would be inaccurate to describe Nemotron OCR v1 as having a fixed consumer API price based on the available information.
NVIDIA states that the model is commercially usable under the NVIDIA Open Model License Agreement. The accompanying post-processing scripts are licensed under Apache 2.0. Organizations should review the applicable license terms and deployment documentation before incorporating the model into a commercial product.
When to choose Nemotron OCR v1
Choose Nemotron OCR v1 when the primary requirement is English text extraction from images and the workflow benefits from bounding boxes, confidence scores, reading order, or layout relationships. It is a good fit for teams already operating NVIDIA GPU infrastructure and for pipelines that need OCR as a preprocessing stage for search, retrieval, RAG, or document intelligence.
Another OCR system may be more appropriate when multilingual recognition is essential, when deployment must run outside an NVIDIA-oriented environment, or when a managed service with clearly published usage pricing is preferred. A general-purpose vision-language model may be a better choice when the task requires answering questions about an image, summarizing visual content, or combining OCR with open-ended reasoning rather than producing dedicated OCR structures.
Within NVIDIA’s current OCR family, Nemotron OCR v2 deserves evaluation for new multilingual deployments. Nemotron OCR v1 remains a practical choice for English-only workloads where its existing model artifacts, layout-aware outputs, or compatibility with an established NeMo Retriever or NIM pipeline are more important than adopting the newer version.
Bottom line
Nemotron OCR v1 is a specialized, English-only OCR model focused on turning complex images into structured text and document-layout information. Its detector-recognizer-relational design makes it more suitable for document ingestion and retrieval pipelines than for conversational use. The most important selection questions are whether English-only coverage is sufficient, whether NVIDIA GPU deployment fits the environment, and whether the application needs layout-aware OCR rather than general visual reasoning.

