Nemotron

Nemotron Table Structure v1

by NVIDIA AI · Current; downloadable and available through NVIDIA NIM/Build

A focused NVIDIA computer-vision model that detects table cells, rows, columns, and merged cells in document images. Its structured detections support OCR alignment, table reconstruction, document ingestion, and retrieval pipelines, but it does not perform OCR or natural-language generation by itself.

Reasoning Coding
NVIDIA Nemotron Table Structure v1 is a focused computer-vision model for one difficult document-processing task: finding the structural parts of tables in images. It detects cells, rows, columns, and merged cells, then returns bounding boxes, class labels, and confidence scores in a structured format. The results can be combined with OCR so that recognized text is assigned to the correct cells and the original table layout can be reconstructed.
Inputs

What it can understand

Images
Capabilities

Supported features

Structured output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Nemotron
Model type Other
Release date 2026-03-02
Status Current; downloadable and available through NVIDIA NIM/Build
Model notes

This is a specialized YOLOX-based object-detection model rather than a language model. It detects cells, rows, and columns and returns bounding boxes, class labels, and confidence scores in a JSON-compatible structure. NVIDIA documents approximately 54 million parameters, 1024 x 1024 image resizing, TensorRT runtime support, and compatibility with Ampere, Hopper, and Lovelace GPUs on Linux. It is intended to work with OCR systems and does not independently transcribe table text. NVIDIA lists commercial and non-commercial use under its applicable model-license terms. The NVIDIA NeMo Retriever support matrix identifies it as the core table-structure service and lists the hosted endpoint https://ai.api.nvidia.com/v1/cv/nvidia/nemotron-table-structure-v1.

Model guide

NVIDIA Nemotron Table Structure v1: Table Detection for Document AI

NVIDIA Nemotron Table Structure v1 is a specialized YOLOX-based object-detection model that identifies table cells, rows, columns, and merged cells in document images. It is designed to support OCR alignment, table reconstruction, document ingestion, and retrieval workflows rather than generate text or perform OCR on its own.

What NVIDIA Nemotron Table Structure v1 does

NVIDIA Nemotron Table Structure v1 is a specialized object-detection model for tables contained in document images. Instead of answering questions, writing text, or converting an entire document into a finished spreadsheet, it identifies where important table components appear in an image.

Its main detection classes are individual cells, rows, and columns. Cells may include merged cells, which are common in reports, invoices, forms, presentations, and other documents. For each detected element, the model provides location information in the form of bounding boxes, along with a class label and confidence score.

This makes the model a structural component in a larger document-AI pipeline. An OCR engine can recognize the words inside the detected regions, while post-processing can use the detected geometry to determine which text belongs to which cell and how rows and columns should be arranged.

Where it fits in NVIDIA's catalog

NVIDIA provides the model through the Build.NVIDIA.com catalog and as a downloadable model. It is also identified as the table-structure component in NVIDIA NeMo Retriever workflows. The model is therefore best understood as a focused vision service within NVIDIA's document-processing and retrieval ecosystem, not as a general-purpose Nemotron language model.

NVIDIA documents both hosted and self-hosted access paths. The hosted trial endpoint and a self-hosted NVIDIA NIM deployment are separate options, with their own operational and access requirements. The model is also associated with an NGC container and an official model repository, giving teams more than one route for evaluation or deployment.

Its documented deployment environment includes TensorRT runtime support, Linux operation, and compatibility with NVIDIA Ampere, Hopper, and Lovelace GPU architectures. These requirements make it particularly relevant to organizations already operating NVIDIA GPU infrastructure.

Technical profile and input

CharacteristicDocumented information
ProviderNVIDIA
Model familyNemotron
Model typeSpecialized object detection
ArchitectureYOLOX with a DarkNet53 backbone, feature pyramid network, and decoupled detection head
Approximate parametersApproximately 54 million
Input modalityRGB document imagery
Documented image resizing1024 × 1024 pixels
Detected componentsCells, merged cells, rows, and columns
OutputBounding boxes, class labels, and confidence scores in a JSON-compatible structure
Runtime and deploymentTensorRT, Linux, and supported NVIDIA GPU architectures

The RGB-image input can represent a rendered PDF page, scanned document, presentation slide, or another visual source containing a table. The supplied specifications do not identify a text context window, maximum output-token limit, audio input, video input, or language-model-style context length. Those concepts do not directly describe this model's primary operation.

How it supports table extraction

A typical workflow begins by converting a PDF page, scan, or slide into an image. Nemotron Table Structure v1 then analyzes the image and identifies the geometry of the table. The resulting boxes can distinguish a row boundary from a column boundary and can identify individual cells, including cells that span multiple rows or columns.

An OCR system can process the same image or the detected cell regions to produce text. A separate orchestration or post-processing layer can then associate each recognized text segment with a detected cell, order the cells, and produce a structured representation such as a spreadsheet-like table, Markdown, JSON, or a database record.

This separation is important. The model supplies visual layout information, but it does not independently transcribe the contents of a cell. It also does not by itself complete table question answering, data validation, semantic interpretation, or business-rule processing.

Main strengths

  • Focused table geometry detection: The model is optimized for locating the structural parts of tables rather than attempting to solve every document-understanding task at once.
  • Support for merged cells: Detecting merged cells can help preserve layouts that do not follow a simple one-cell-per-column pattern.
  • Useful structured output: Bounding boxes, labels, and confidence scores are practical inputs for downstream OCR alignment and reconstruction systems.
  • Integration with NVIDIA tooling: Build.NVIDIA.com, NGC, TensorRT, and NeMo Retriever provide deployment paths for teams using NVIDIA infrastructure.
  • Suitable for retrieval pipelines: Preserving table structure can improve the quality of indexing and retrieval when important information is arranged spatially rather than in ordinary paragraphs.

These strengths are especially relevant when the problem is not simply finding text, but retaining the relationships among headers, rows, columns, and values.

Limitations and boundaries

Nemotron Table Structure v1 is not an OCR engine. It detects where table elements are located, but applications still need a separate method to read the text inside those elements. It is also not a general-purpose vision-language model: the supplied specifications do not describe document question answering, unrestricted visual reasoning, image generation, or natural-language response generation.

Additional post-processing is normally required to turn detections into a usable table. For example, an application may need to resolve overlapping boxes, order cells spatially, infer relationships between detected rows and columns, associate OCR text with the correct region, and handle ambiguous confidence scores.

Performance may vary with image quality, unusual layouts, skew, handwriting, complex merged cells, and tables that differ substantially from the training data. The supplied research does not provide a benchmark score or a universal accuracy guarantee, so production teams should evaluate representative documents before relying on the model for automated extraction.

The model also has infrastructure constraints. NVIDIA documents Linux and TensorRT support together with compatibility for Ampere, Hopper, and Lovelace GPUs. Organizations without suitable NVIDIA hardware or deployment experience may find a hosted service or a different document-processing platform easier to operate.

Pricing and access

No input or output price is specified in the supplied research. The model is described as available through NVIDIA's Build.NVIDIA.com catalog and as a downloadable model, with separate hosted and self-hosted NIM deployment paths. Consequently, a definitive per-image, per-page, subscription, or infrastructure price should not be inferred from the available information.

Licensing is described as allowing commercial and non-commercial use under NVIDIA's applicable model-license terms. Teams should review the current license and the conditions associated with the selected hosted, downloadable, or self-hosted deployment route before using the model in a production system.

Capabilities that do not apply or are not documented

This item should not be evaluated like a conversational language model. It does not produce direct text, image, video, audio, music, speech, or embedding output according to the supplied model record. Its primary output is structured detection data. The supplied research also does not verify tool or function calling, streaming, fine-tuning, caching, batch APIs, web search, or a separate JSON mode. The JSON-compatible detection structure should not be confused with a general language-model JSON-output feature.

Reasoning and coding are likewise not its purpose. It can contribute to a reasoning or coding application as one component in a larger pipeline, but the model itself is intended to detect visual table structure. A language model, OCR system, or application layer would be needed for interpretation, code generation, or natural-language answers.

When to choose this model

Choose Nemotron Table Structure v1 when the central problem is locating table structure in document images and you need machine-readable detections for a downstream workflow. It is a good candidate for:

  • Extracting tables from scanned documents or PDF-rendered pages
  • Aligning OCR text with the correct cells
  • Preserving rows, columns, and merged-cell relationships
  • Preparing tables for Markdown, JSON, spreadsheet, or database conversion
  • Improving document ingestion and retrieval-augmented generation pipelines
  • Building document-processing services on supported NVIDIA GPU infrastructure

Its focused design may be preferable to using a general-purpose vision-language model when predictable table geometry is the required output and the application already has OCR and post-processing components.

When another option may be more appropriate

Use an OCR-focused system instead when the main requirement is accurate transcription and table layout is secondary. Use a broader document-understanding or vision-language system when the application must answer questions about a page, summarize content, interpret charts, or combine visual evidence with natural-language reasoning. A general language model is more appropriate for generating explanations, writing transformation code, or answering users after the table has been extracted.

A different deployment option may also be preferable if the workload cannot use Linux, TensorRT, or compatible NVIDIA GPUs. Likewise, organizations that need a clearly published usage price should confirm the current commercial terms for NVIDIA's hosted or self-managed route, or compare services that publish per-page or per-request pricing.

Bottom line

NVIDIA Nemotron Table Structure v1 is a narrowly targeted building block for document AI. Its value comes from identifying the geometry of cells, rows, columns, and merged cells in RGB document images, then exposing that information in a structured form that other systems can use. It does not replace OCR or a complete table-understanding pipeline, but it can provide the layout foundation needed to reconstruct tables and make their contents more useful for extraction and retrieval.


Answers to Frequently Asked Questions

When should teams choose NVIDIA Nemotron Table Structure v1 instead of a general vision-language model?
Teams should choose it when they need reliable, machine-readable table geometry for workflows such as OCR alignment, table reconstruction, spreadsheet or JSON conversion, and document retrieval. A broader vision-language model is more appropriate for document question answering, summarization, visual reasoning, or natural-language generation.
What hardware and deployment environment does NVIDIA Nemotron Table Structure v1 require?
NVIDIA documents deployment with TensorRT on Linux and compatibility with NVIDIA Ampere, Hopper, and Lovelace GPU architectures. It is available through Build.NVIDIA.com, downloadable model resources, hosted access, and self-hosted NVIDIA NIM deployment paths.
What input and output does NVIDIA Nemotron Table Structure v1 support?
The model accepts RGB document imagery, such as rendered PDF pages, scans, or presentation slides. Its output is a JSON-compatible structure containing detected table components, bounding boxes, class labels, and confidence scores.
What is NVIDIA Nemotron Table Structure v1 used for?
NVIDIA Nemotron Table Structure v1 is a specialized object-detection model that identifies table components in document images, including cells, merged cells, rows, and columns. It returns bounding boxes, class labels, and confidence scores for use in downstream document-AI workflows.
Does NVIDIA Nemotron Table Structure v1 perform OCR or read text inside table cells?
No. The model detects the visual structure and location of table elements but does not transcribe cell contents. An OCR engine is required to recognize text, while additional post-processing is needed to associate that text with the correct cells.


Sources 5
Provider

About NVIDIA AI