What NVIDIA Nemotron Page Elements v3 does
NVIDIA Nemotron Page Elements v3 is a document-layout object detection model. Its job is to find meaningful regions on a page, not to read, summarize, or generate the content inside those regions. For example, it can identify the approximate location of a table in a report, a chart in a presentation, a title above a section, or a repeated header and footer.
The model is useful when a document-processing system needs to decide what should happen next. A detected table can be routed to table-structure extraction, a text region to OCR, and a chart or infographic to a specialized visual-analysis component. This makes the model an early routing and layout-analysis step in larger document AI pipelines.
NVIDIA lists the model as current, downloadable, and available through its hosted NIM interface. It supersedes the earlier nemoretriever-page-elements-v2 model. The model is part of NVIDIA's broader Nemotron and NeMo Retriever ecosystem, but the model itself is focused specifically on page-element detection.
Which document elements it detects
Nemotron Page Elements v3 produces detections for six documented categories:
- Table: Data arranged in rows and columns.
- Chart: Visualizations such as bar, line, or pie charts.
- Infographic: Diagrams, flowcharts, and other complex visual displays.
- Title: Section titles or titles associated with tables, charts, and infographics.
- Header or footer: Repeated page-level regions, such as document headers, footers, or page furniture.
- Text: Text paragraphs or standalone text that does not fit another class.
Each detection includes a bounding box showing where the element appears, a confidence score, and a class label. The output therefore describes the page's visual structure in a form that downstream software can process.
Architecture, input, and output
According to NVIDIA's model documentation, the detector uses the YOLOX object-detection architecture with a DarkNet53 backbone and a feature pyramid network with a decoupled head. YOLOX is designed to locate and classify objects efficiently in an image. In this model, the objects are document regions rather than everyday photographic objects.
NVIDIA reports approximately 54 million parameters. Input pages are RGB images resized to 1024 × 1024 pixels before inference. The documented runtime engine is TensorRT, and the supported hardware includes NVIDIA Ampere, Hopper, and Lovelace GPU architectures. The supplied documentation identifies Linux as the operating system for the documented deployment.
| Specification | Documented detail |
|---|---|
| Primary task | Document page-element object detection |
| Architecture | YOLOX with DarkNet53 backbone and FPN decoupled head |
| Approximate parameters | 54 million |
| Image input | RGB page image |
| Preprocessing size | 1024 × 1024 pixels |
| Output | Bounding boxes, confidence scores, and element labels |
| Runtime | TensorRT |
| Documented hardware | NVIDIA Ampere, Hopper, and Lovelace GPUs |
| Documented operating system | Linux |
The 1024 × 1024 image size is an input-processing specification, not a context window or token limit. This model does not expose a language-model context length, maximum generated-token limit, or text-generation output budget in the supplied documentation.
Training and reported evaluation
NVIDIA reports that the model was pretrained on 118,287 COCO train2017 images and then fine-tuned on 36,093 Digital Corpora images. The fine-tuning annotations were produced with Azure AI Document Intelligence and a human data-annotation team.
The reported class-level average precision values are:
| Class | Reported average precision |
|---|---|
| Infographic | 66.863% |
| Chart | 54.191% |
| Header or footer | 53.895% |
| Text | 45.418% |
| Table | 44.643% |
| Title | 38.529% |
These figures are provider-reported evaluation results, not a guarantee for every document collection. Performance can vary with page design, scan quality, image resolution, typography, language, and the document domain. The lower reported title score compared with the infographic score also suggests that individual classes should not be treated as equally reliable without testing on representative material.
Where it fits in NVIDIA's catalog
Nemotron Page Elements v3 is a specialized model rather than a general-purpose Nemotron language model. NVIDIA makes it available as a downloadable model and through a hosted NIM interface. NIM is relevant here as a way to serve the model, but the underlying capability remains page-layout detection.
The model can therefore sit near the front of an enterprise document pipeline. A typical workflow might convert a PDF page into an image, send the image to Nemotron Page Elements v3, use the returned boxes to crop or route regions, and then apply OCR, table parsing, chart analysis, or retrieval indexing. A multimodal retrieval-augmented generation system can use the detections to preserve layout information before sending selected regions to later components.
This positioning distinguishes the model from a generative vision-language model. A vision-language model may describe an image or answer questions about it, while Nemotron Page Elements v3 supplies a more constrained and predictable map of page regions. Conversely, the detector does not provide the semantic interpretation that a generative model might provide.
Main strengths and limitations
Strengths
- Clear specialization: The model focuses on a defined set of document elements instead of attempting broad visual or language reasoning.
- Useful structured output: Bounding boxes, labels, and confidence scores can be consumed directly by downstream processing software.
- Pipeline compatibility: It can help route different regions to OCR, table extraction, indexing, or multimodal retrieval components.
- NVIDIA deployment path: TensorRT support and documented compatibility with Ampere, Hopper, and Lovelace hardware provide a deployment route for NVIDIA GPU environments.
- Downloadable availability: Organizations can evaluate or deploy the model under the NVIDIA Open Model License Agreement, subject to that license's terms.
Limitations
- Detection is not extraction: The model locates a table but does not reconstruct its rows and columns or return the table's values.
- Detection is not OCR: A text bounding box does not contain a transcription of the text.
- No generation or general reasoning: The model is not intended for conversation, summarization, question answering, coding, or free-form document interpretation.
- Limited class vocabulary: Its documented output categories cover common page elements but do not represent every possible document object.
- Hardware and operating-system requirements: The documented deployment targets Linux and NVIDIA GPU architectures supported by TensorRT.
- Quality varies by document: Unusual layouts, low-quality scans, dense pages, and domain-specific formats may require validation and fallback logic.
Modalities, reasoning, and tool support
The verified input modality is an RGB image. The model does not accept audio, video, or a conversational text prompt as its primary documented input. Its output is structured detection data rather than text, an image, audio, or video.
Reasoning and coding capabilities are not applicable in the way they are for language models. The detector does not independently explain a document, write code, call functions, browse the web, or use external tools. Any OCR, parsing, retrieval, or language reasoning must be supplied by other components in the surrounding application.
The structured output is valuable because it gives software an explicit representation of detected regions. However, structured detections should not be confused with a general JSON-mode capability for arbitrary user requests. The supplied research does not document a separate JSON mode, streaming behavior, fine-tuning interface, caching feature, or batch API for this model.
Pricing, access, and license
No input-token, output-token, subscription, or per-image price is supplied for Nemotron Page Elements v3. The downloadable model is identified as available for commercial and non-commercial use under the NVIDIA Open Model License Agreement. Hosted trial access through NVIDIA NIM may be rate limited, but the research does not provide a price for hosted production usage.
In practical terms, downloadable use shifts the main cost considerations toward NVIDIA GPU infrastructure, storage, deployment, and operations rather than a documented per-request model fee. Hosted NIM access may reduce infrastructure work, but its applicable limits and commercial pricing should be checked in NVIDIA's current service documentation before production planning.
When to choose Nemotron Page Elements v3
Choose this model when the immediate problem is locating regions on document pages and the rest of the processing pipeline is handled by separate components. It is a good fit for:
- Preprocessing enterprise reports, forms, presentations, and scanned documents.
- Routing text regions to OCR while sending tables to table-structure extraction.
- Detecting charts and infographics before applying specialized visual analysis.
- Removing or separately handling repeated headers and footers during indexing.
- Adding page-layout metadata to document search or multimodal RAG systems.
- Running a focused detector in an NVIDIA GPU-accelerated environment.
A different option may be more appropriate if the application needs transcription, table reconstruction, document question answering, summarization, image captioning, or broad visual reasoning in one model. In those cases, a dedicated OCR or table parser, a vision-language model, or a combined document-understanding system may be needed after or instead of this detector. Nemotron Page Elements v3 is most useful when predictable region detection is more important than having one model perform every document task.
Bottom line
NVIDIA Nemotron Page Elements v3 is a focused layout-analysis component for document AI. Its core value is converting a page image into labeled regions that downstream systems can process. The model's documented architecture, TensorRT runtime, NVIDIA GPU support, and structured detections make it suitable for OCR preparation, table and chart routing, indexing, and layout-aware retrieval pipelines. It should not be evaluated as a chatbot or general-purpose multimodal model: it identifies where page elements are, while other tools must read, extract, interpret, or generate content from them.

