What is NVIDIA Nemotron OCR v2?
NVIDIA Nemotron OCR v2 is an optical character recognition (OCR) model for extracting text from images. Unlike a general-purpose language model, it is specialized for locating and transcribing text rather than writing essays, answering questions, generating code, or conducting open-ended conversations.
The model accepts RGB images in PNG or JPEG format and returns structured OCR detections. A detection can include the recognized text, a confidence score, and the coordinates of the corresponding text region. Nemotron OCR v2 also analyzes relationships between detected regions, helping downstream systems reconstruct document structure and reading order.
NVIDIA positions the model within its NeMo Retriever collection and makes it available as a downloadable model through Hugging Face. It can also run through NVIDIA NIM, a containerized inference service, or NVIDIA-hosted services. This gives users several deployment choices, from integrating the model into their own GPU environment to testing it through a hosted endpoint.
Architecture and model variants
Nemotron OCR v2 combines three jointly trained components:
- Text detector: A RegNetX-8GF convolutional network locates text regions in the image.
- Text recognizer: A Transformer-based component converts the visual text regions into characters and words.
- Relational model: A layout-oriented component estimates grouping, document structure, and reading order.
NVIDIA provides English and multilingual variants. The English recognizer has three Transformer layers, a 256-dimensional hidden size, a maximum sequence length of 32, and an 855-character set. The multilingual recognizer has six Transformer layers, a 512-dimensional hidden size, a maximum sequence length of 128, and a 14,244-character set.
The model card reports approximately 53.8 million parameters for the English variant and 83.9 million for the multilingual variant. These are model-architecture details, not a guarantee of a particular deployment footprint or speed on every GPU.
Supported languages, inputs, and outputs
The multilingual variant supports 11 languages according to NVIDIA's model listing. The training mixture covers real-world and synthetic examples involving documents, natural scenes, tables, charts, infographics, and handwritten pages. The model supports word-, sentence-, and paragraph-level aggregation and performs automatic multi-scale resizing.
Inputs are image-based: individual RGB PNG or JPEG files, or batches of images. The supplied specifications do not identify text, audio, or video as direct input modalities. Nemotron OCR v2 produces text and structured OCR information rather than images, audio, video, or speech.
Its relational processing is important when an application needs more than a plain text transcription. For example, a document pipeline can use bounding boxes and reading-order information to preserve the relationship between a heading, its paragraph, a table, and nearby annotations. The same outputs can support chart-to-text extraction, table-to-text extraction, infographic processing, and image-based search.
Deployment options and performance
Nemotron OCR v2 can run with PyTorch on supported NVIDIA GPUs, including Ampere, Hopper, Lovelace, and Blackwell architectures. NVIDIA NIM provides a CUDA-based runtime for image OCR, with English and multilingual variants, latency-oriented and throughput-oriented configurations, and a /v1/ocr inference endpoint.
NVIDIA's NIM performance documentation reports up to 129.0 images per second for the multilingual variant on a GB200 under the listed benchmark configuration. The English variant is reported at up to 147.8 images per second under its corresponding configuration. These are provider-reported benchmark results, not universal performance guarantees. Results can change with GPU type, image encoding, resolution, batch size, concurrency, and runtime settings.
In practical terms, the model is most attractive when OCR throughput matters and the organization already operates compatible NVIDIA infrastructure. A hosted endpoint can reduce deployment work, while a self-managed model or NIM deployment can provide more control over processing location and scaling. The research does not establish one universal cost for all of these deployment paths.
Pricing and access
No official public per-token or per-image price was identified for Nemotron OCR v2. The model is downloadable, and NVIDIA lists a free hosted endpoint. Hosted availability, quotas, authentication requirements, and usage conditions can change, so the current NVIDIA model page should be checked before using it in production.
Self-hosting is not necessarily cost-free: the model requires compatible GPU infrastructure, and the cost of GPUs, storage, operations, and any applicable NVIDIA software or service licensing is separate from the model's downloadable status. NIM is a deployment mechanism rather than a single universal pricing plan for every Nemotron OCR v2 use case.
Main strengths and limitations
Strengths
- Structured OCR results: Text, confidence scores, and bounding boxes are more useful for document pipelines than an unlocated text string.
- Layout awareness: Relational analysis can help preserve grouping and reading order across complex pages.
- Multilingual coverage: The multilingual variant supports 11 languages according to NVIDIA's listing and uses a substantially larger character set than the English variant.
- Broad document coverage: The documented training mixture includes forms, reports, tables, charts, infographics, natural scenes, and handwritten pages.
- Deployment flexibility: Users can work with Hugging Face, PyTorch, NVIDIA NIM, or NVIDIA-hosted services.
- GPU-oriented throughput: NVIDIA's published NIM results indicate high image-processing throughput under a specified GB200 configuration.
Limitations
- It is not a general-purpose language model: It is not designed for conversational generation, broad reasoning, coding, or free-form text creation.
- GPU requirements: The documented accelerated deployments require compatible NVIDIA GPU infrastructure. Performance and feasibility depend on the selected hardware.
- No conventional language-model context window: The supplied research does not publish a general context length or maximum output-token limit. The model works on images and produces OCR detections rather than a conventional prompt-and-completion response.
- No general tool calling: The specifications do not identify function calling, web search, streaming, or agent tools as model capabilities.
- OCR accuracy remains input-dependent: Image quality, resolution, unusual typography, handwriting, page layout, and language can affect results. Confidence scores should be used as signals for validation rather than treated as absolute truth.
Reasoning, coding, and tool support
Nemotron OCR v2 performs task-specific visual and layout inference, but it should not be evaluated as a reasoning model. Its relational component can infer relationships between text regions and help determine reading order, yet this is different from multi-step reasoning over arbitrary questions or documents.
It has no documented coding capability and does not generate software. It also does not provide built-in web browsing, function calling, external tools, or general agent actions. An application can place the OCR output into a separate language model or document-processing workflow, but that additional reasoning or tool use comes from the surrounding system, not from Nemotron OCR v2 itself.
Best use cases
Nemotron OCR v2 is a strong fit when an application needs to convert visual content into structured text before another system processes it. Suitable examples include:
- Ingesting scanned PDFs, forms, reports, and photographed documents.
- Extracting text and coordinates for enterprise search or retrieval-augmented generation pipelines.
- Preprocessing tables, charts, and infographics for document-intelligence systems.
- Building image-based search with text-region metadata.
- Preserving reading order and layout relationships during document conversion.
- Processing multilingual document collections on NVIDIA GPU infrastructure.
For production workflows, applications should retain the source image, store confidence scores, and define review rules for low-confidence detections or document types that are especially sensitive to OCR errors.
When to choose Nemotron OCR v2
Choose Nemotron OCR v2 when the primary problem is multilingual OCR with structured locations and layout relationships, particularly when GPU-accelerated processing or NVIDIA deployment tooling is already part of the environment. It is more appropriate than a general vision-language assistant when predictable OCR fields, bounding boxes, and throughput matter more than open-ended explanation.
A general-purpose vision-language model may be more suitable when the application must answer questions about an image, summarize a document, interpret ambiguous visual content, or combine OCR with broad reasoning in one response. A lightweight traditional OCR system may be preferable for simple, narrow-language scans where the relational layout features and GPU deployment do not justify the additional infrastructure. Nemotron OCR v2 should therefore be selected for its OCR-specific structured output, multilingual coverage, and deployment options—not as a replacement for a conversational or coding model.

