What Llama Nemotron Rerank VL 1B v2 is
Llama Nemotron Rerank VL 1B v2 is a multimodal cross-encoder reranker provided by NVIDIA. A reranker does not normally search an entire document collection by itself. Instead, a first-stage retrieval system uses embeddings, keyword search, or a hybrid method to produce a smaller list of candidate documents. The reranker then reads the query and each candidate together, estimates their relevance, and helps place the most useful candidates at the top.
The model is particularly targeted at question-answering retrieval over visual documents. A candidate can be an ordinary text passage, an image of a document page or slide, or a combined representation containing both an image and extracted text. This allows ranking decisions to use information contained in tables, charts, page layout, screenshots, and infographics, rather than relying only on extracted text.
NVIDIA released the model on December 18, 2025. It is available as a downloadable model through Hugging Face and through NVIDIA’s NeMo Retriever Reranking NIM and retrieval API offerings. Those are deployment routes for the same general model role, rather than evidence that it is a general-purpose chat model.
How it fits in a retrieval system
A practical retrieval-augmented generation pipeline can use Llama Nemotron Rerank VL 1B v2 in three stages:
- Retrieve candidates: An embedding model, keyword engine, or hybrid search system finds a manageable set of potentially relevant passages or document pages.
- Rerank candidates: The query and each candidate are passed to Llama Nemotron Rerank VL 1B v2. The model returns a relevance or logit score that can be used to reorder the candidates.
- Generate an answer: A separate language model receives the highest-ranked evidence and produces the final response, if an answer is required.
This division of labor is important. The model improves evidence selection, but it does not generate the final answer, create embeddings for the initial search, or replace the rest of a retrieval application. Applying a cross-encoder to every document in a large corpus would also be inefficient, because each query-document pair must be evaluated individually. In most systems, it should be applied only to the top candidates from the first retrieval stage.
Inputs, outputs, and architecture
The input query is text. The document candidate may be text, an image, or image combined with text. Image inputs can represent document pages, slides, screenshots, tables, charts, and other visual content. The output is a relevance score or logit used for ranking; it is not generated prose, an image, audio, or video.
| Specification | Verified detail |
|---|---|
| Model type | Multimodal cross-encoder reranker |
| Provider | NVIDIA |
| Text query | Supported |
| Text candidate | Supported |
| Image candidate | Supported |
| Image plus text candidate | Supported |
| Primary output | Relevance or logit score |
| Maximum sequence length | 8,192 tokens in the NVIDIA NIM support matrix |
| Generated answer output | Not supported; a separate generation model is required |
NVIDIA describes the architecture as an Eagle-family vision-language system combining a SigLIP 2 400M vision encoder with the Llama Nemotron Rerank 1B v2 language component. The model has approximately 1.7 billion parameters, uses mean pooling over decoder outputs, and includes a binary classification head fine-tuned for ranking. These architectural details explain why it can compare text queries with visually represented documents, but they do not turn it into a general-purpose multimodal assistant.
Reported retrieval quality
NVIDIA reports evaluations on visual document retrieval datasets including ViDoRe V1, ViDoRe V2, ViDoRe V3, DigitalCorpora-10k, and Earnings V2. In the reported five-dataset evaluation, adding the reranker to a visual embedding baseline increased average Recall@5 from 71.04% to 76.12% for text inputs, from 71.20% to 76.12% for image inputs, and from 73.24% to 77.64% for combined image-and-text inputs.
These are provider-reported benchmark results rather than a guarantee for every corpus or application. Actual results can vary with image quality, OCR accuracy, document formatting, candidate-retrieval quality, query style, and the number of candidates passed to the reranker. NVIDIA also reports competitive results on text retrieval benchmarks covering BEIR, MIRACL, MLQA, and MLDR, suggesting that the model can be used with mixed collections containing both conventional passages and visually rich documents.
Main strengths and trade-offs
The model’s clearest strength is that it can rank visual evidence alongside ordinary text. A text-only reranker may miss information that appears mainly in a chart, table, diagram, or page layout. Llama Nemotron Rerank VL 1B v2 is intended to preserve access to that information when the retrieval system represents documents as images or as image-text pairs.
Its relatively small model size compared with large generative vision-language models can also make it a focused choice for ranking workloads. It performs a narrower task and returns a score rather than spending compute on an open-ended answer. That specialization can be useful when an application already has a separate generation model and needs better evidence ordering.
The trade-off is that it is still a cross-encoder. Ranking many candidates requires many query-document evaluations, so latency and compute usage grow as the candidate set grows. A typical design should use a fast embedding or hybrid search stage to narrow the corpus first. A text-only reranker may remain more appropriate when the collection contains no meaningful visual information and low latency is more important than image understanding.
Deployment and availability
NVIDIA documents several ways to use the model. It can be loaded from NVIDIA’s Hugging Face repository with Transformers or Sentence Transformers, and the documented usage may require trusted remote code. NVIDIA also provides a NeMo Retriever Reranking NIM and a retrieval API route for hosted or containerized inference. The NIM support matrix lists an 8,192-token maximum sequence length.
The downloadable route offers more control over infrastructure and data handling, but requires the operator to provide suitable inference hardware, software configuration, and scaling. NIM or an API can simplify deployment, but availability, operational requirements, and any applicable service charges depend on the selected NVIDIA offering. No public per-token hosted price was verified for this exact model in the supplied research, so a definitive API price should not be assumed.
The model is identified as ready for commercial use under the NVIDIA Open Model License Agreement, with additional Llama 3.2 licensing terms. Organizations should review the current license documents and the terms of the chosen hosted or containerized service before production deployment.
Reasoning, coding, and tool support
Llama Nemotron Rerank VL 1B v2 performs relevance classification rather than visible multi-step reasoning. It can make a more informed ranking decision by comparing a query with text and visual evidence, but it does not provide a chain-of-thought response or act as an autonomous reasoning agent. Its role is evidence selection.
There is no verified built-in tool or function-calling capability for this model. It does not browse the web, execute code, call external services, or perform actions. Coding is not a primary capability either. Code may be present in a document candidate and can be evaluated for relevance like other text or visual content, but the model is not intended to write, debug, or explain software.
Likewise, the model does not produce structured application answers, embeddings, speech, images, video, or audio. Its useful output is a ranking score that another system can consume.
Best use cases
- Visual document search: Rank pages from PDFs, presentations, reports, and scanned records when tables, charts, or layout contain important evidence.
- Multimodal RAG: Select stronger evidence before a separate language model writes an answer.
- Question answering over slides and reports: Match natural-language questions with pages containing relevant visual or textual information.
- Mixed text-and-image collections: Apply one reranking stage to ordinary passages, page images, and image-text representations.
- Enterprise retrieval: Improve the ordering of a limited candidate set after an initial embedding or hybrid search stage.
When to choose this model
Choose Llama Nemotron Rerank VL 1B v2 when the quality of retrieval depends on visual document content and your application can support a two-stage search pipeline. It is a particularly strong fit when a question may be answered by a chart, table, scanned page, or slide layout that a text-only index does not represent completely.
Choose a text-only reranker instead when all relevant content is already cleanly represented as text and the additional visual processing would not improve results. Choose an embedding model when the immediate need is first-stage similarity search rather than pairwise reranking. Choose a generative vision-language model when the system must interpret evidence and produce a natural-language answer in one interaction. In that design, Llama Nemotron Rerank VL 1B v2 can still be useful upstream to select better evidence.
Limitations and practical checklist
The model does not provide a complete question-answering application. It requires a query, candidate documents, a retrieval strategy, and usually a separate answer-generation component. It also does not remove the need for OCR, document rendering, access controls, or quality checks in a production pipeline. Poor page images, incomplete text extraction, or weak first-stage retrieval can limit the value of the reranker.
Before adopting it, verify the input formatting required by the selected Transformers, NIM, or API integration; confirm how image and extracted-text candidates are represented; test latency with the intended candidate count; and check the 8,192-token NIM limit against the longest inputs. Also evaluate it on the organization’s own documents rather than relying only on NVIDIA’s reported benchmarks. The model is most valuable when visual understanding changes which candidates should appear at the top, not when the task is simple text search or answer generation.

