Llama Nemotron Embed VL

Llama Nemotron Embed VL 1B v2

by NVIDIA AI · Current and downloadable

NVIDIA Llama Nemotron Embed VL 1B v2 is a downloadable multimodal embedding encoder for retrieving text, document pages, tables, charts, and image-text content. It produces 2048-dimensional vectors, supports a 10,240-token evaluated context, and is designed for vector search, question-answer retrieval, and multimodal RAG rather than conversational generation.

Embeddings Reasoning Coding
Llama Nemotron Embed VL 1B v2 is a downloadable NVIDIA model built for retrieval rather than conversation. It can encode text, document images, or pages containing both images and text into a shared vector space, allowing a text question to locate relevant material in a document collection. The model is especially useful when important information exists in page layouts, tables, charts, or infographics that plain OCR may not capture reliably.
Outputs

What Llama Nemotron Embed VL 1B v2 can produce

Embeddings
Inputs

What it can understand

Text Images Multimodal input
Model profile

Performance characteristics

2/10 Reasoning
2/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Llama Nemotron Embed VL
Model type Embedding
Context window 10K tokens
Release date 2025-12-18
Status Current and downloadable
Knowledge cutoff notes

No explicit knowledge-cutoff date was identified in the authoritative model documentation. As an embedding encoder, the model is intended to encode supplied text and images rather than answer from a stated training knowledge cutoff.

Model notes

Approximately 1.7B parameters combining a Llama 3.2 1B language model with a SigLip2 400M image encoder. Produces fixed-size 2048-dimensional embeddings. Supports text, image, and image-plus-text document inputs. NVIDIA's NIM interface requires query or passage input_type values; vLLM uses query and document roles or prefixes. The model is downloadable and governed by the NVIDIA Open Model License, with additional Llama 3.2 license terms. No public per-token hosted price was identified for the exact model. It is intended for NVIDIA GPU-accelerated deployment and does not generate natural-language responses.

Model guide

Llama Nemotron Embed VL 1B v2 for Visual Document Retrieval

Llama Nemotron Embed VL 1B v2 is NVIDIA's multimodal embedding model for retrieving relevant text, document pages, tables, charts, and image-text content. It converts queries and documents into 2048-dimensional vectors for semantic search, question-answer retrieval, vector databases, and multimodal retrieval-augmented generation.

What Llama Nemotron Embed VL 1B v2 does

Llama Nemotron Embed VL 1B v2 is a multimodal embedding encoder from NVIDIA. Instead of generating a written answer, it converts an input into a numerical representation called an embedding. The representation contains information about the input's meaning and can be compared with other embeddings to find similar or relevant content.

The model is designed for question-answer retrieval over text and visual documents. A system can encode document pages or text passages, store the resulting vectors in a vector database, and then encode a user's question to search for the closest matches. A separate language model can use the retrieved passages or page images to produce a final natural-language answer.

NVIDIA makes the model available as an open model on Hugging Face and as a downloadable NVIDIA NIM deployment. In NVIDIA's current catalog, it belongs to the retrieval and embedding portion of the lineup rather than the general-purpose conversational-model category.

Supported inputs and multimodal retrieval

The model accepts text strings, RGB images, and combinations of images and text. An image may be a scanned page, a digitally rendered document page, or a page containing several kinds of visual information. Supported retrieval scenarios include:

  • Text-to-text retrieval: finding relevant text passages for a text query.
  • Text-to-image retrieval: finding document-page images that answer or relate to a text question.
  • Image-text retrieval: searching pages represented by both a visual image and accompanying text.
  • Question-answer retrieval: selecting evidence for a retrieval-augmented generation system.

This visual input capability matters when meaning depends on layout or graphics. A page with a table, chart, form, or infographic can be indexed as an image rather than being reduced entirely to OCR text. The model can therefore help retrieval systems search the original visual document representation.

Architecture and embedding output

According to NVIDIA's model documentation, Llama Nemotron Embed VL 1B v2 uses an Eagle VLM-style design combining a Llama 3.2 1B language model with a SigLip2 400M image encoder. NVIDIA describes the combined model as having approximately 1.7 billion parameters.

The model produces one fixed-size vector with 2048 dimensions for each text, image, or multimodal document input. Mean pooling is used to create the final representation. The model is trained as a bi-encoder: queries and relevant documents are brought closer together in the embedding space, while unrelated items are separated.

In practical terms, the embedding is not useful as a standalone answer. It is an indexable representation that a similarity-search system can compare using a metric such as cosine similarity. The application still needs to retain the original text or images so that retrieved evidence can be displayed to a user or passed to a separate generation model.

Context and image-processing limits

NVIDIA reports an evaluated maximum context length of 10,240 tokens. When an image and text are supplied together, their combined token use must remain within that limit. Tokens are units used to represent text and visual information internally; the limit is not equivalent to a fixed number of words or a fixed number of pages.

The model uses dynamic image tiling. In the tested configuration, an image can use up to six tiles together with a lower-resolution thumbnail, with each tile consuming approximately 256 visual tokens. This approach allows the encoder to process page-level detail while retaining a broader view of the document.

NVIDIA's examples use modality-specific settings below the overall evaluated maximum. Image inputs are configured at up to 2,048 tokens, text-only inputs can use up to 8,192 tokens in the Transformers example, and image-text inputs use the full 10,240-token setting. These values should be treated as documented example configurations rather than a promise that every deployment will have identical memory or performance characteristics.

How to use it in a retrieval system

A typical workflow has two stages. During indexing, the application divides a document collection into passages or pages, encodes each item with Llama Nemotron Embed VL 1B v2, and stores the vectors with references to the original content. During search, it encodes the user's question and compares that query vector with the stored document vectors.

The query and document sides must use the correct roles. NVIDIA's NIM interface requires an input_type of query or passage. In the vLLM integration, the corresponding roles are query and document. This distinction is an important implementation detail: encoding a passage as though it were a query, or applying inconsistent prefixes and preprocessing, can reduce retrieval quality.

After retrieval, a separate generation model is normally responsible for summarizing the selected evidence or answering the user. Llama Nemotron Embed VL 1B v2 itself returns numerical vectors, not natural-language responses, citations, images, audio, or video.

Deployment and hardware

The model can be loaded from Hugging Face with Transformers or Sentence Transformers using trusted remote model code. NVIDIA also provides a downloadable NIM container for Linux systems with NVIDIA GPUs. The documented NIM exposes an OpenAI-compatible embeddings endpoint, which can simplify integration with applications already designed around an embeddings API.

NVIDIA's model card lists Ampere, Hopper, Lovelace, and Blackwell GPU microarchitectures as supported hardware families. The Transformers examples use version 4.56.0 or newer and optionally support Flash Attention. For higher-throughput serving, NVIDIA documents vLLM support and recommends vLLM 0.17.0 or newer.

These deployment choices make the model a stronger fit for NVIDIA GPU environments than for small CPU-only installations. CPU execution may be possible depending on the software stack, but the supplied documentation positions the model around GPU-accelerated use. Actual throughput and memory requirements depend on image size, batching, precision, context settings, and the selected serving framework.

Pricing and access

No public per-token hosted price was identified for this exact model in the supplied NVIDIA documentation. The model is described as downloadable, and it is available through NVIDIA's model and deployment ecosystem rather than being documented here as a standard consumer subscription feature.

Downloadable availability does not mean that deployment is cost-free in every environment. Users may incur expenses for NVIDIA GPU hardware, cloud instances, storage, database infrastructure, and operational support. The exact licensing terms should also be reviewed before commercial deployment: the supplied model notes identify the NVIDIA Open Model License and additional Llama 3.2 license terms.

Strengths and limitations

The model's main strength is its focus on visual document retrieval. It can represent a page image, text, or a combined image-text input in the same general retrieval workflow. That makes it suitable for collections where tables, charts, page layout, and infographics carry information that may be lost by text extraction alone.

Its 2048-dimensional output provides a consistent representation for vector databases, while the documented query and passage roles provide a clear integration pattern for retrieval pipelines. The downloadable model and NIM option can also be useful when an organization wants more control over deployment than a hosted-only embedding service provides.

The principal limitation is scope. This is not a chat model, reasoning model, code-generation model, or answer-generation model. It does not provide native tool calling, streaming responses, structured natural-language output, or a built-in web-search function in the supplied specifications. A complete question-answering product requires additional components for document ingestion, vector search, reranking if needed, answer generation, and application orchestration.

Visual processing can also increase memory and processing demands. Large page images and image-text combinations consume context capacity and may be more expensive to process than short text passages. Retrieval quality should be tested on the organization's own documents, especially when scans are low quality, pages contain unusual layouts, or the application mixes OCR text with page images.

Capability profile

CapabilityAssessment
Text inputSupported
Image inputSupported
Audio and video inputNot identified in the supplied specifications
Text generationNot supported; the model produces embeddings
Image, audio, or video outputNot supported
Embedding outputSupported; 2048 dimensions
Context limit10,240 evaluated tokens
Maximum generated tokensNot applicable to this embedding encoder
Tool or function callingNot identified

The table separates the model's verified role from capabilities commonly associated with conversational models. Its lack of text generation is not a missing feature for its intended purpose; it is what allows the model to specialize in encoding and retrieval.

When to choose this model

Choose Llama Nemotron Embed VL 1B v2 when the main problem is finding relevant evidence across text and visual documents, particularly when you can deploy on NVIDIA GPU infrastructure. It is a strong candidate for:

  • Enterprise search across PDFs, scanned pages, reports, and presentations.
  • Retrieval systems where tables, charts, or page layout are important.
  • Multimodal RAG pipelines that retrieve page images as well as text.
  • Vector-database indexing for question-answer search.
  • Teams that want a downloadable model or NVIDIA NIM deployment option.

A conventional text-only embedding model may be more appropriate when the corpus is clean text, visual layout does not affect meaning, or infrastructure must remain especially small and inexpensive. A hosted embedding service may be preferable when avoiding GPU operations is more important than local deployment control. A generative language model is the better choice when the primary requirement is to write answers, summarize content, generate code, or carry on a conversation.

Within a complete retrieval application, this model should be evaluated on retrieval recall, latency, memory use, and total infrastructure cost rather than on conversational benchmarks. Its value is in selecting useful evidence; another model or application layer must turn that evidence into an answer.


Answers to Frequently Asked Questions

How can Llama Nemotron Embed VL 1B v2 be deployed?
The model is available through Hugging Face and as a downloadable NVIDIA NIM deployment. It can be used with Transformers or Sentence Transformers, while NVIDIA also documents vLLM support for higher-throughput serving. The documented deployment options are oriented toward NVIDIA GPU environments.
What are the context limit and embedding size of Llama Nemotron Embed VL 1B v2?
NVIDIA reports an evaluated maximum context length of 10,240 tokens. The model produces one fixed-size vector with 2048 dimensions for each text, image, or multimodal input.
Does Llama Nemotron Embed VL 1B v2 generate answers or chat responses?
No. It produces 2048-dimensional embeddings rather than natural-language answers. A separate language model is typically used to summarize retrieved passages or page images and generate the final response.
Can Llama Nemotron Embed VL 1B v2 process images and document pages?
Yes. The model accepts text strings, RGB images, and combinations of images and text. It can retrieve scanned pages, rendered document pages, tables, charts, forms, and infographics where visual layout contributes to meaning.
What is Llama Nemotron Embed VL 1B v2 used for?
Llama Nemotron Embed VL 1B v2 is a multimodal embedding encoder for retrieving relevant information from text and visual documents. It converts text, images, or image-text inputs into numerical vectors that can be indexed and compared in a vector database.


Sources 4
Provider

About NVIDIA AI