Llama Nemotron Rerank VL

Llama Nemotron Rerank VL 1B v2

by NVIDIA AI · Current; downloadable and available through NVIDIA NIM and retrieval APIs

NVIDIA’s Llama Nemotron Rerank VL 1B v2 is a multimodal cross-encoder that reranks retrieved text passages, document images, and image-text candidates against a text query. It is designed for visual document search and multimodal RAG, returning relevance scores rather than generated answers. The model supports downloadable, NIM, and retrieval API deployment, with an 8,192-token limit listed for NIM.

Reasoning Coding
Llama Nemotron Rerank VL 1B v2 is designed for the second stage of a retrieval pipeline. An embedding or hybrid search system first finds potentially relevant documents, then this NVIDIA model examines each query-document pair in more detail and assigns a relevance score. Because it can process document-page images as well as text, it is intended for PDFs, slides, tables, charts, and other visually rich material that text-only rerankers may not fully understand.
Inputs

What it can understand

Text Images Multimodal input
Model profile

Performance characteristics

2/10 Reasoning
2/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Llama Nemotron Rerank VL
Model type Multimodal
Context window 8K tokens
Maximum output tokens
Release date 2025-12-18
Status Current; downloadable and available through NVIDIA NIM and retrieval APIs
Knowledge cutoff notes

No authoritative model-specific knowledge cutoff was identified. This is a ranking model rather than a general-purpose generative assistant, so a conventional knowledge-cutoff date is not prominently specified in the available model documentation.

Model notes

NVIDIA describes this as a multimodal cross-encoder reranker with approximately 1.7 billion parameters. It accepts text, image, or image-plus-text document candidates and returns a relevance/logit score rather than generated prose. The architecture combines a SigLIP 2 400M vision encoder with the Llama Nemotron Rerank 1B v2 language component. NVIDIA reports a maximum sequence length of 8,192 tokens for the optimized NIM. The Hugging Face usage examples may require trust_remote_code. The model is released under the NVIDIA Open Model License Agreement with additional Llama 3.2 licensing terms. No public per-token hosted price was verified for the exact model; NVIDIA provides downloadable and NIM/API deployment options.

Model guide

Llama Nemotron Rerank VL 1B v2 for Multimodal Document Retrieval

Llama Nemotron Rerank VL 1B v2 is NVIDIA’s downloadable multimodal cross-encoder reranking model. It improves search results by scoring the relevance of retrieved text passages, document images, or combined image-and-text candidates against a text query. Its main role is visual document retrieval and multimodal retrieval-augmented generation, not conversational response generation.

What Llama Nemotron Rerank VL 1B v2 is

Llama Nemotron Rerank VL 1B v2 is a multimodal cross-encoder reranker provided by NVIDIA. A reranker does not normally search an entire document collection by itself. Instead, a first-stage retrieval system uses embeddings, keyword search, or a hybrid method to produce a smaller list of candidate documents. The reranker then reads the query and each candidate together, estimates their relevance, and helps place the most useful candidates at the top.

The model is particularly targeted at question-answering retrieval over visual documents. A candidate can be an ordinary text passage, an image of a document page or slide, or a combined representation containing both an image and extracted text. This allows ranking decisions to use information contained in tables, charts, page layout, screenshots, and infographics, rather than relying only on extracted text.

NVIDIA released the model on December 18, 2025. It is available as a downloadable model through Hugging Face and through NVIDIA’s NeMo Retriever Reranking NIM and retrieval API offerings. Those are deployment routes for the same general model role, rather than evidence that it is a general-purpose chat model.

How it fits in a retrieval system

A practical retrieval-augmented generation pipeline can use Llama Nemotron Rerank VL 1B v2 in three stages:

  1. Retrieve candidates: An embedding model, keyword engine, or hybrid search system finds a manageable set of potentially relevant passages or document pages.
  2. Rerank candidates: The query and each candidate are passed to Llama Nemotron Rerank VL 1B v2. The model returns a relevance or logit score that can be used to reorder the candidates.
  3. Generate an answer: A separate language model receives the highest-ranked evidence and produces the final response, if an answer is required.

This division of labor is important. The model improves evidence selection, but it does not generate the final answer, create embeddings for the initial search, or replace the rest of a retrieval application. Applying a cross-encoder to every document in a large corpus would also be inefficient, because each query-document pair must be evaluated individually. In most systems, it should be applied only to the top candidates from the first retrieval stage.

Inputs, outputs, and architecture

The input query is text. The document candidate may be text, an image, or image combined with text. Image inputs can represent document pages, slides, screenshots, tables, charts, and other visual content. The output is a relevance score or logit used for ranking; it is not generated prose, an image, audio, or video.

SpecificationVerified detail
Model typeMultimodal cross-encoder reranker
ProviderNVIDIA
Text querySupported
Text candidateSupported
Image candidateSupported
Image plus text candidateSupported
Primary outputRelevance or logit score
Maximum sequence length8,192 tokens in the NVIDIA NIM support matrix
Generated answer outputNot supported; a separate generation model is required

NVIDIA describes the architecture as an Eagle-family vision-language system combining a SigLIP 2 400M vision encoder with the Llama Nemotron Rerank 1B v2 language component. The model has approximately 1.7 billion parameters, uses mean pooling over decoder outputs, and includes a binary classification head fine-tuned for ranking. These architectural details explain why it can compare text queries with visually represented documents, but they do not turn it into a general-purpose multimodal assistant.

Reported retrieval quality

NVIDIA reports evaluations on visual document retrieval datasets including ViDoRe V1, ViDoRe V2, ViDoRe V3, DigitalCorpora-10k, and Earnings V2. In the reported five-dataset evaluation, adding the reranker to a visual embedding baseline increased average Recall@5 from 71.04% to 76.12% for text inputs, from 71.20% to 76.12% for image inputs, and from 73.24% to 77.64% for combined image-and-text inputs.

These are provider-reported benchmark results rather than a guarantee for every corpus or application. Actual results can vary with image quality, OCR accuracy, document formatting, candidate-retrieval quality, query style, and the number of candidates passed to the reranker. NVIDIA also reports competitive results on text retrieval benchmarks covering BEIR, MIRACL, MLQA, and MLDR, suggesting that the model can be used with mixed collections containing both conventional passages and visually rich documents.

Main strengths and trade-offs

The model’s clearest strength is that it can rank visual evidence alongside ordinary text. A text-only reranker may miss information that appears mainly in a chart, table, diagram, or page layout. Llama Nemotron Rerank VL 1B v2 is intended to preserve access to that information when the retrieval system represents documents as images or as image-text pairs.

Its relatively small model size compared with large generative vision-language models can also make it a focused choice for ranking workloads. It performs a narrower task and returns a score rather than spending compute on an open-ended answer. That specialization can be useful when an application already has a separate generation model and needs better evidence ordering.

The trade-off is that it is still a cross-encoder. Ranking many candidates requires many query-document evaluations, so latency and compute usage grow as the candidate set grows. A typical design should use a fast embedding or hybrid search stage to narrow the corpus first. A text-only reranker may remain more appropriate when the collection contains no meaningful visual information and low latency is more important than image understanding.

Deployment and availability

NVIDIA documents several ways to use the model. It can be loaded from NVIDIA’s Hugging Face repository with Transformers or Sentence Transformers, and the documented usage may require trusted remote code. NVIDIA also provides a NeMo Retriever Reranking NIM and a retrieval API route for hosted or containerized inference. The NIM support matrix lists an 8,192-token maximum sequence length.

The downloadable route offers more control over infrastructure and data handling, but requires the operator to provide suitable inference hardware, software configuration, and scaling. NIM or an API can simplify deployment, but availability, operational requirements, and any applicable service charges depend on the selected NVIDIA offering. No public per-token hosted price was verified for this exact model in the supplied research, so a definitive API price should not be assumed.

The model is identified as ready for commercial use under the NVIDIA Open Model License Agreement, with additional Llama 3.2 licensing terms. Organizations should review the current license documents and the terms of the chosen hosted or containerized service before production deployment.

Reasoning, coding, and tool support

Llama Nemotron Rerank VL 1B v2 performs relevance classification rather than visible multi-step reasoning. It can make a more informed ranking decision by comparing a query with text and visual evidence, but it does not provide a chain-of-thought response or act as an autonomous reasoning agent. Its role is evidence selection.

There is no verified built-in tool or function-calling capability for this model. It does not browse the web, execute code, call external services, or perform actions. Coding is not a primary capability either. Code may be present in a document candidate and can be evaluated for relevance like other text or visual content, but the model is not intended to write, debug, or explain software.

Likewise, the model does not produce structured application answers, embeddings, speech, images, video, or audio. Its useful output is a ranking score that another system can consume.

Best use cases

  • Visual document search: Rank pages from PDFs, presentations, reports, and scanned records when tables, charts, or layout contain important evidence.
  • Multimodal RAG: Select stronger evidence before a separate language model writes an answer.
  • Question answering over slides and reports: Match natural-language questions with pages containing relevant visual or textual information.
  • Mixed text-and-image collections: Apply one reranking stage to ordinary passages, page images, and image-text representations.
  • Enterprise retrieval: Improve the ordering of a limited candidate set after an initial embedding or hybrid search stage.

When to choose this model

Choose Llama Nemotron Rerank VL 1B v2 when the quality of retrieval depends on visual document content and your application can support a two-stage search pipeline. It is a particularly strong fit when a question may be answered by a chart, table, scanned page, or slide layout that a text-only index does not represent completely.

Choose a text-only reranker instead when all relevant content is already cleanly represented as text and the additional visual processing would not improve results. Choose an embedding model when the immediate need is first-stage similarity search rather than pairwise reranking. Choose a generative vision-language model when the system must interpret evidence and produce a natural-language answer in one interaction. In that design, Llama Nemotron Rerank VL 1B v2 can still be useful upstream to select better evidence.

Limitations and practical checklist

The model does not provide a complete question-answering application. It requires a query, candidate documents, a retrieval strategy, and usually a separate answer-generation component. It also does not remove the need for OCR, document rendering, access controls, or quality checks in a production pipeline. Poor page images, incomplete text extraction, or weak first-stage retrieval can limit the value of the reranker.

Before adopting it, verify the input formatting required by the selected Transformers, NIM, or API integration; confirm how image and extracted-text candidates are represented; test latency with the intended candidate count; and check the 8,192-token NIM limit against the longest inputs. Also evaluate it on the organization’s own documents rather than relying only on NVIDIA’s reported benchmarks. The model is most valuable when visual understanding changes which candidates should appear at the top, not when the task is simple text search or answer generation.


Answers to Frequently Asked Questions

What is Llama Nemotron Rerank VL 1B v2 used for?
Llama Nemotron Rerank VL 1B v2 is a multimodal cross-encoder reranker from NVIDIA. It ranks candidate documents for relevance to a text query and is designed for visual document retrieval, including PDFs, slides, tables, charts, screenshots, and image-text document representations.
What types of inputs does Llama Nemotron Rerank VL 1B v2 support?
The model accepts a text query and a candidate represented as text, an image, or an image combined with extracted text. Images may contain document pages, slides, tables, charts, screenshots, and infographics.
Does Llama Nemotron Rerank VL 1B v2 generate answers?
No. The model outputs a relevance score or logit used to reorder retrieved candidates. A separate generative language model is required to produce a natural-language answer.
How does Llama Nemotron Rerank VL 1B v2 fit into a RAG pipeline?
A first-stage embedding, keyword, or hybrid search system retrieves a manageable set of candidates. Llama Nemotron Rerank VL 1B v2 then compares the query with each text, image, or image-text candidate and reranks them. A separate language model can use the highest-ranked evidence to generate the final response.
When should you choose Llama Nemotron Rerank VL 1B v2 over a text-only reranker?
Choose it when important retrieval evidence appears in visual document content such as charts, tables, scanned pages, or slide layouts. A text-only reranker may be more suitable when all relevant information is already cleanly represented as text and lower latency is the priority.


Sources 6
Provider

About NVIDIA AI