Nemotron Page Elements

nemotron-page-elements-v3

by NVIDIA AI · Current; downloadable and available through NVIDIA NIM

A specialized NVIDIA YOLOX object detector that maps document pages into labeled regions with bounding boxes and confidence scores, helping downstream systems route content to OCR, table extraction, indexing, and multimodal retrieval workflows.

Reasoning Coding
NVIDIA Nemotron Page Elements v3 is built for the stage before OCR, table extraction, indexing, or multimodal retrieval. Given an RGB image of a document page, it identifies important visual regions and returns their locations and labels. The downloadable model uses TensorRT on NVIDIA GPU hardware and is also exposed through NVIDIA's hosted NIM catalog.
Inputs

What it can understand

Images
Capabilities

Supported features

Structured output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Nemotron Page Elements
Model type Other
Release date 2026-03-02
Status Current; downloadable and available through NVIDIA NIM
Knowledge cutoff notes

A knowledge cutoff is not applicable or publicly specified because this is a specialized object detection model rather than a generative language model.

Model notes

Specialized YOLOX object detector with approximately 54 million parameters. Accepts RGB images resized to 1024 × 1024 pixels and returns bounding boxes, confidence scores, and labels for tables, charts, infographics, titles, text, and header/footer regions. Uses TensorRT and is documented for Linux on NVIDIA Ampere, Hopper, and Lovelace hardware. The model supersedes nemoretriever-page-elements-v2. Hosted NIM trial access may be rate limited. NVIDIA reports AP values of 44.643% for tables, 54.191% for charts, 38.529% for titles, 66.863% for infographics, 45.418% for text, and 53.895% for header/footer detection.

Model guide

NVIDIA Nemotron Page Elements v3: Document Layout Detection for OCR and RAG Pipelines

NVIDIA Nemotron Page Elements v3 is a specialized YOLOX-based object detection model that locates tables, charts, infographics, titles, text, and header or footer regions in document page images. It provides structured bounding boxes, class labels, and confidence scores for document-processing workflows, rather than generating text or interpreting document contents.

What NVIDIA Nemotron Page Elements v3 does

NVIDIA Nemotron Page Elements v3 is a document-layout object detection model. Its job is to find meaningful regions on a page, not to read, summarize, or generate the content inside those regions. For example, it can identify the approximate location of a table in a report, a chart in a presentation, a title above a section, or a repeated header and footer.

The model is useful when a document-processing system needs to decide what should happen next. A detected table can be routed to table-structure extraction, a text region to OCR, and a chart or infographic to a specialized visual-analysis component. This makes the model an early routing and layout-analysis step in larger document AI pipelines.

NVIDIA lists the model as current, downloadable, and available through its hosted NIM interface. It supersedes the earlier nemoretriever-page-elements-v2 model. The model is part of NVIDIA's broader Nemotron and NeMo Retriever ecosystem, but the model itself is focused specifically on page-element detection.

Which document elements it detects

Nemotron Page Elements v3 produces detections for six documented categories:

  • Table: Data arranged in rows and columns.
  • Chart: Visualizations such as bar, line, or pie charts.
  • Infographic: Diagrams, flowcharts, and other complex visual displays.
  • Title: Section titles or titles associated with tables, charts, and infographics.
  • Header or footer: Repeated page-level regions, such as document headers, footers, or page furniture.
  • Text: Text paragraphs or standalone text that does not fit another class.

Each detection includes a bounding box showing where the element appears, a confidence score, and a class label. The output therefore describes the page's visual structure in a form that downstream software can process.

Architecture, input, and output

According to NVIDIA's model documentation, the detector uses the YOLOX object-detection architecture with a DarkNet53 backbone and a feature pyramid network with a decoupled head. YOLOX is designed to locate and classify objects efficiently in an image. In this model, the objects are document regions rather than everyday photographic objects.

NVIDIA reports approximately 54 million parameters. Input pages are RGB images resized to 1024 × 1024 pixels before inference. The documented runtime engine is TensorRT, and the supported hardware includes NVIDIA Ampere, Hopper, and Lovelace GPU architectures. The supplied documentation identifies Linux as the operating system for the documented deployment.

SpecificationDocumented detail
Primary taskDocument page-element object detection
ArchitectureYOLOX with DarkNet53 backbone and FPN decoupled head
Approximate parameters54 million
Image inputRGB page image
Preprocessing size1024 × 1024 pixels
OutputBounding boxes, confidence scores, and element labels
RuntimeTensorRT
Documented hardwareNVIDIA Ampere, Hopper, and Lovelace GPUs
Documented operating systemLinux

The 1024 × 1024 image size is an input-processing specification, not a context window or token limit. This model does not expose a language-model context length, maximum generated-token limit, or text-generation output budget in the supplied documentation.

Training and reported evaluation

NVIDIA reports that the model was pretrained on 118,287 COCO train2017 images and then fine-tuned on 36,093 Digital Corpora images. The fine-tuning annotations were produced with Azure AI Document Intelligence and a human data-annotation team.

The reported class-level average precision values are:

ClassReported average precision
Infographic66.863%
Chart54.191%
Header or footer53.895%
Text45.418%
Table44.643%
Title38.529%

These figures are provider-reported evaluation results, not a guarantee for every document collection. Performance can vary with page design, scan quality, image resolution, typography, language, and the document domain. The lower reported title score compared with the infographic score also suggests that individual classes should not be treated as equally reliable without testing on representative material.

Where it fits in NVIDIA's catalog

Nemotron Page Elements v3 is a specialized model rather than a general-purpose Nemotron language model. NVIDIA makes it available as a downloadable model and through a hosted NIM interface. NIM is relevant here as a way to serve the model, but the underlying capability remains page-layout detection.

The model can therefore sit near the front of an enterprise document pipeline. A typical workflow might convert a PDF page into an image, send the image to Nemotron Page Elements v3, use the returned boxes to crop or route regions, and then apply OCR, table parsing, chart analysis, or retrieval indexing. A multimodal retrieval-augmented generation system can use the detections to preserve layout information before sending selected regions to later components.

This positioning distinguishes the model from a generative vision-language model. A vision-language model may describe an image or answer questions about it, while Nemotron Page Elements v3 supplies a more constrained and predictable map of page regions. Conversely, the detector does not provide the semantic interpretation that a generative model might provide.

Main strengths and limitations

Strengths

  • Clear specialization: The model focuses on a defined set of document elements instead of attempting broad visual or language reasoning.
  • Useful structured output: Bounding boxes, labels, and confidence scores can be consumed directly by downstream processing software.
  • Pipeline compatibility: It can help route different regions to OCR, table extraction, indexing, or multimodal retrieval components.
  • NVIDIA deployment path: TensorRT support and documented compatibility with Ampere, Hopper, and Lovelace hardware provide a deployment route for NVIDIA GPU environments.
  • Downloadable availability: Organizations can evaluate or deploy the model under the NVIDIA Open Model License Agreement, subject to that license's terms.

Limitations

  • Detection is not extraction: The model locates a table but does not reconstruct its rows and columns or return the table's values.
  • Detection is not OCR: A text bounding box does not contain a transcription of the text.
  • No generation or general reasoning: The model is not intended for conversation, summarization, question answering, coding, or free-form document interpretation.
  • Limited class vocabulary: Its documented output categories cover common page elements but do not represent every possible document object.
  • Hardware and operating-system requirements: The documented deployment targets Linux and NVIDIA GPU architectures supported by TensorRT.
  • Quality varies by document: Unusual layouts, low-quality scans, dense pages, and domain-specific formats may require validation and fallback logic.

Modalities, reasoning, and tool support

The verified input modality is an RGB image. The model does not accept audio, video, or a conversational text prompt as its primary documented input. Its output is structured detection data rather than text, an image, audio, or video.

Reasoning and coding capabilities are not applicable in the way they are for language models. The detector does not independently explain a document, write code, call functions, browse the web, or use external tools. Any OCR, parsing, retrieval, or language reasoning must be supplied by other components in the surrounding application.

The structured output is valuable because it gives software an explicit representation of detected regions. However, structured detections should not be confused with a general JSON-mode capability for arbitrary user requests. The supplied research does not document a separate JSON mode, streaming behavior, fine-tuning interface, caching feature, or batch API for this model.

Pricing, access, and license

No input-token, output-token, subscription, or per-image price is supplied for Nemotron Page Elements v3. The downloadable model is identified as available for commercial and non-commercial use under the NVIDIA Open Model License Agreement. Hosted trial access through NVIDIA NIM may be rate limited, but the research does not provide a price for hosted production usage.

In practical terms, downloadable use shifts the main cost considerations toward NVIDIA GPU infrastructure, storage, deployment, and operations rather than a documented per-request model fee. Hosted NIM access may reduce infrastructure work, but its applicable limits and commercial pricing should be checked in NVIDIA's current service documentation before production planning.

When to choose Nemotron Page Elements v3

Choose this model when the immediate problem is locating regions on document pages and the rest of the processing pipeline is handled by separate components. It is a good fit for:

  • Preprocessing enterprise reports, forms, presentations, and scanned documents.
  • Routing text regions to OCR while sending tables to table-structure extraction.
  • Detecting charts and infographics before applying specialized visual analysis.
  • Removing or separately handling repeated headers and footers during indexing.
  • Adding page-layout metadata to document search or multimodal RAG systems.
  • Running a focused detector in an NVIDIA GPU-accelerated environment.

A different option may be more appropriate if the application needs transcription, table reconstruction, document question answering, summarization, image captioning, or broad visual reasoning in one model. In those cases, a dedicated OCR or table parser, a vision-language model, or a combined document-understanding system may be needed after or instead of this detector. Nemotron Page Elements v3 is most useful when predictable region detection is more important than having one model perform every document task.

Bottom line

NVIDIA Nemotron Page Elements v3 is a focused layout-analysis component for document AI. Its core value is converting a page image into labeled regions that downstream systems can process. The model's documented architecture, TensorRT runtime, NVIDIA GPU support, and structured detections make it suitable for OCR preparation, table and chart routing, indexing, and layout-aware retrieval pipelines. It should not be evaluated as a chatbot or general-purpose multimodal model: it identifies where page elements are, while other tools must read, extract, interpret, or generate content from them.


Answers to Frequently Asked Questions

When should organizations choose Nemotron Page Elements v3?
Organizations should choose it when they need predictable document-region detection as an early step in an OCR, document search, or multimodal RAG pipeline. A different or additional model is needed for transcription, table reconstruction, summarization, document question answering, or broad visual reasoning.
What are the input, output, and hardware requirements for Nemotron Page Elements v3?
The documented input is an RGB page image resized to 1024 × 1024 pixels. The output consists of bounding boxes, confidence scores, and element labels. The model uses TensorRT and is documented for Linux deployments on NVIDIA Ampere, Hopper, and Lovelace GPUs.
Does Nemotron Page Elements v3 perform OCR or extract table data?
No. The model performs page-element detection only. It identifies where a text region or table is located, but separate OCR and table-structure extraction tools are required to transcribe text or reconstruct table rows, columns, and values.
What is NVIDIA Nemotron Page Elements v3 used for?
NVIDIA Nemotron Page Elements v3 detects meaningful regions on document pages, such as tables, charts, infographics, titles, headers, footers, and text. Its structured output helps route regions to OCR, table extraction, visual analysis, indexing, or retrieval components.
What document elements can Nemotron Page Elements v3 detect?
The model detects six documented categories: Table, Chart, Infographic, Title, Header or footer, and Text. Each detection includes a bounding box, confidence score, and class label.


Sources 4
Provider

About NVIDIA AI