Nemotron Graphic Elements

nemotron-graphic-elements-v1

by NVIDIA AI · Current and downloadable

A specialized NVIDIA YOLOX-based object detector that locates ten chart and graph elements in RGB document images, returning bounding boxes, labels, and confidence scores for downstream OCR, extraction, indexing, retrieval, and RAG workflows.

Reasoning Coding
NVIDIA Nemotron Graphic Elements v1 is a downloadable object-detection model for finding important components inside charts and graphs. Rather than generating text or images, it identifies regions such as chart titles, axis labels, legends, mark labels, and value labels, then returns their locations and confidence scores. The model is intended as a focused building block for document-processing systems, especially workflows that need to locate chart elements before applying OCR, chart parsing, retrieval, or other downstream analysis.
Inputs

What it can understand

Images
Capabilities

Supported features

Structured output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Nemotron Graphic Elements
Model type Other
Release date 2026-03-02
Status Current and downloadable
Knowledge cutoff notes

The model is a trained object detector and NVIDIA does not publish a knowledge-cutoff date for it. Its behavior depends on its training and evaluation datasets rather than a language-model knowledge cutoff.

Model notes

Specialized YOLOX object detector with approximately 54 million parameters. Detects ten classes: chart_title, x_title, y_title, xlabel, ylabel, other, legend_label, legend_title, mark_label, and value_label. Accepts single images or image batches as RGB NumPy arrays and returns structured detections containing bounding boxes, labels, and confidence scores. Uses TensorRT and is documented for Linux systems with NVIDIA Ampere, Hopper, or Lovelace GPUs. The model supersedes CACHED and is available under the NVIDIA Open Model License Agreement. NVIDIA provides a hosted trial experience, while downloadable deployment is available through NIM. Official documentation does not specify token context, maximum output tokens, or per-request pricing because this is an image object-detection model rather than a language model.

Model guide

NVIDIA Nemotron Graphic Elements v1: Chart-Element Detection for Document AI

NVIDIA Nemotron Graphic Elements v1 is a specialized YOLOX-based computer-vision model that detects and localizes ten types of chart components in RGB document images. It returns structured bounding boxes, class labels, and confidence scores for use in chart extraction, OCR preparation, document indexing, multimodal retrieval, and RAG pipelines.

What NVIDIA Nemotron Graphic Elements v1 does

NVIDIA Nemotron Graphic Elements v1 is a specialized computer-vision model for detecting graphic elements in charts and graphs embedded in documents. Its job is to answer a focused question: where are the meaningful chart components in this image, and what type of component is each one?

For example, given an RGB image containing a bar chart, the model may identify the chart title at the top, the x- and y-axis titles, tick labels, legend text, labels attached to marks, and numeric value labels. Each detection includes a bounding box describing the element's location, a class label, and a confidence score. These results can then be passed to OCR or a larger document-understanding pipeline.

This makes the model different from a general-purpose vision-language model. It does not provide a natural-language explanation of a chart, reconstruct the underlying data table, or answer questions about trends. Its primary output is structured detection data that helps other systems process a chart more reliably.

Where it fits in NVIDIA's catalog

NVIDIA provides Nemotron Graphic Elements v1 as a downloadable model within its broader NIM and model ecosystem. The supplied model information identifies it as a current, downloadable model, with a release date of March 2, 2026. NVIDIA describes it as superseding the earlier CACHED model.

Its position in the catalog is therefore narrower than that of a general language, vision-language, or generative model. It is a specialized inference component for document AI and chart processing. NVIDIA also provides a hosted trial experience, while downloadable deployment is available through NIM and related model repositories. The hosted experience and self-managed deployment should be treated as separate access routes with potentially different operational and licensing conditions.

Inputs, outputs, and detected classes

The model accepts two-dimensional RGB images. Documentation supports processing either a single image or a batch of images. The available research does not specify a fixed pixel resolution, maximum image dimensions, token context window, or maximum output-token limit. Those language-model limits are not directly applicable to this object detector.

Nemotron Graphic Elements v1 recognizes ten chart-element classes:

  • chart_title
  • x_title
  • y_title
  • xlabel
  • ylabel
  • other
  • legend_label
  • legend_title
  • mark_label
  • value_label

The output is structured detection information containing bounding boxes, labels, and confidence scores for each input image. The model does not directly output text, audio, video, or generated images. Although it processes visual input, its output is best understood as machine-readable localization data rather than multimodal conversation.

Architecture and hardware requirements

According to NVIDIA's model information, the detector uses a YOLOX architecture with a DarkNet53 backbone and a feature pyramid network with a decoupled head. YOLOX is an object-detection design intended to locate and classify multiple objects within an image. The feature pyramid helps the detector work across elements of different sizes, such as a large chart title and small value labels.

NVIDIA reports approximately 54 million parameters. The model is optimized for TensorRT and is documented as compatible with NVIDIA Ampere, Hopper, and Lovelace hardware on Linux. This hardware and operating-system scope is important when estimating deployment effort: the model is not presented as a cross-platform consumer application or as a browser-only chart assistant.

TensorRT optimization is relevant for production inference because it targets NVIDIA GPU execution, but actual throughput will depend on the selected GPU, image dimensions, batch size, preprocessing, and the surrounding pipeline. The supplied research does not provide a universal images-per-second figure, so deployment teams should benchmark their own workload rather than treating the model as having a guaranteed speed.

Performance and reliability considerations

NVIDIA reports evaluation results on the PMC Chart dataset, with average precision varying by detected class. The available research indicates that performance is strongest for chart and axis-title detection, while classes including mark_label, other, legend_title, and value_label are more challenging.

These are provider-reported evaluation observations, not a guarantee for every document collection. Results can change with chart style, image quality, font size, layout density, scan artifacts, and the visual conventions used in a company's reports. A chart that is clean and high resolution may be easier to process than a compressed screenshot, an unusual infographic, or a dense legacy PDF.

The model's confidence scores can support filtering and review policies. For example, a document pipeline might automatically accept high-confidence title detections while sending uncertain value-label or legend detections to a second processing stage. That kind of workflow design is an editorial implementation recommendation, not a behavior that NVIDIA guarantees.

Practical use cases

Nemotron Graphic Elements v1 is most useful when chart structure must be located before another system performs interpretation. Suitable applications include:

  • Document extraction: locate chart regions and their component labels in reports, presentations, and other business documents.
  • OCR preparation: provide targeted image regions to OCR instead of sending an entire page or chart to a text-recognition system.
  • Chart-element indexing: store detected regions and categories as searchable metadata.
  • Multimodal retrieval: improve retrieval pipelines by identifying chart components that can be linked to surrounding page text or extracted labels.
  • RAG augmentation: prepare chart regions for a later retrieval-augmented generation workflow.
  • Legacy-report processing: add structure to older visual documents where charts are available as images rather than native data objects.

A typical pipeline might first render a PDF page to an RGB image, run the detector, crop the detected regions, apply OCR to relevant labels, and then combine those results with page text and document metadata. A separate chart-understanding or data-reconstruction component would still be needed to infer relationships among labels, marks, axes, and values.

What the model does not do

The model is not an OCR engine by itself. Its detections indicate where elements are located, but the available research does not say that it reads every detected word or number. OCR or another recognition component is needed when the workflow requires the actual text inside a bounding box.

It is also not a complete chart-understanding system. It does not, by itself, recover the underlying table, determine every data series, explain a trend, answer questions about the visualization, or generate a natural-language summary. The other class also indicates that not every visible chart feature necessarily receives a more specific category.

No token context limit or maximum output-token limit is documented because this is an image detector rather than a conversational language model. The supplied information also does not document tool calling, function calling, streaming generation, fine-tuning, caching, or a general JSON mode. Its structured detection response should not be confused with a general-purpose language-model JSON mode.

Pricing, licensing, and access

No per-request or recurring price is specified in the supplied research for Nemotron Graphic Elements v1. The model is available as a downloadable deployment option, and NVIDIA provides a hosted trial experience that may have separate access conditions or rate limits. Because no verified price is available, the cost of a production deployment should be evaluated through the chosen NVIDIA hosting or infrastructure route rather than inferred from a language-model API price.

The model card states that the model is available for commercial and non-commercial use under the NVIDIA Open Model License Agreement. Users should review the current license text and the terms associated with hosted access before incorporating the model into a commercial workflow. Downloadable deployment may also introduce infrastructure costs for NVIDIA GPUs, storage, monitoring, and operational support, even when no model-specific usage price is listed.

When to choose Nemotron Graphic Elements v1

Choose this model when the main requirement is fast, focused localization of chart components and you control, or can build, the surrounding document-processing pipeline. It is a good fit for teams that need structured detections rather than a conversational answer, especially when charts must be processed in batches and passed to OCR, indexing, retrieval, or RAG stages.

Its specialization can be an advantage over using a larger general-purpose vision-language model for every page. A detector can provide a predictable set of chart-element categories and avoid asking a generative model to infer bounding boxes through free-form output. The TensorRT-oriented design may also suit NVIDIA GPU deployments where inference cost and throughput matter.

Another option may be more appropriate when the desired result is chart explanation, question answering, OCR, table reconstruction, or natural-language reporting in one step. A general vision-language model may be more suitable for those tasks, although it may offer less specialized or less consistently structured localization. A conventional OCR system remains necessary when accurate text transcription is the goal. For highly unusual charts, very low-resolution images, or workflows requiring complete semantic reconstruction, this detector should be treated as one stage rather than the entire solution.

Bottom line

NVIDIA Nemotron Graphic Elements v1 is a narrowly focused model for locating chart components in RGB document images. Its verified strengths are its defined ten-class detection vocabulary, structured bounding-box output, YOLOX-based architecture, and suitability for NVIDIA TensorRT deployment on supported Linux GPU systems. Its main limitation is equally clear: it identifies visual regions but does not independently perform OCR, explain charts, or reconstruct their data.

For document teams building a modular chart-processing pipeline, that narrow role can be useful. The model provides the localization stage, while OCR, retrieval, chart parsing, and language-generation components handle the interpretation stages. Production users should validate class-level accuracy on their own documents and account for the absence of published pricing, context limits, and universal throughput figures.


Answers to Frequently Asked Questions

How is NVIDIA Nemotron Graphic Elements v1 accessed and licensed?
NVIDIA provides the model as a downloadable deployment option through its model ecosystem and also offers a hosted trial experience. The model card states that it is available for commercial and non-commercial use under the NVIDIA Open Model License Agreement. No verified per-request or recurring price is specified in the supplied information.
What hardware and software does NVIDIA Nemotron Graphic Elements v1 support?
The model uses a YOLOX architecture with a DarkNet53 backbone and is optimized for TensorRT. NVIDIA documents compatibility with NVIDIA Ampere, Hopper, and Lovelace hardware running Linux.
Does NVIDIA Nemotron Graphic Elements v1 perform OCR or explain charts?
No. The model locates chart elements but does not directly read their text, reconstruct the underlying data table, explain trends, answer questions about a visualization, or generate a natural-language summary. OCR and separate chart-understanding components are required for those tasks.
What is NVIDIA Nemotron Graphic Elements v1 used for?
NVIDIA Nemotron Graphic Elements v1 detects and classifies chart components in RGB document images. It identifies elements such as chart titles, axis titles, tick labels, legend text, mark labels, and value labels, returning bounding boxes, class labels, and confidence scores.
What chart elements can NVIDIA Nemotron Graphic Elements v1 detect?
The model recognizes ten classes: chart_title, x_title, y_title, xlabel, ylabel, other, legend_label, legend_title, mark_label, and value_label.


Sources 4
Provider

About NVIDIA AI