Nemotron 3 Embed

Nemotron-3-Embed-1B-BF16

by NVIDIA AI · Current; available through NVIDIA NIM and downloadable Hugging Face weights

A technical overview of NVIDIA Nemotron 3 Embed 1B, including its embedding dimensions, context limit, multilingual retrieval capabilities, deployment options, licensing, and supported use cases.

Embeddings
Nemotron-3-Embed-1B-BF16 is NVIDIA's compact Nemotron 3 retrieval model. It uses a pruned and distilled Ministral-based transformer architecture, supports 34 evaluated languages, and is available through NVIDIA NIM and as downloadable Hugging Face weights under the OpenMDW-1.1 license.
Outputs

What Nemotron-3-Embed-1B-BF16 can produce

Embeddings
Inputs

What it can understand

Text
Capabilities

Supported features

Fine-tuning
Specifications

Technical details

Model family Nemotron 3 Embed
Model type Embedding
Context window 33K tokens
Release date 2026-07-16
Status Current; available through NVIDIA NIM and downloadable Hugging Face weights
Knowledge cutoff notes

No authoritative knowledge-cutoff date was identified for this embedding model. Knowledge-cutoff terminology is generally not applicable to a retrieval encoder in the same way it is to generative language models.

Model notes

The canonical downloadable checkpoint is nvidia/Nemotron-3-Embed-1B-BF16. The NVIDIA API/NIM identifier is nvidia/nemotron-3-embed-1b. It produces 2,048-dimensional dense float vectors and supports query and passage modes for retrieval. NVIDIA documents evaluation across 34 languages and reports RTEB, ViDoRe-V3 text, and MMTEB retrieval results. The model is based on a pruned and distilled Ministral-3-3B-Instruct-2512 architecture and has approximately 1.14 billion parameters. NVIDIA NeMo documentation supports full-weight fine-tuning and merged LoRA fine-tuning. No public per-token or per-request pricing was verified for the exact model; NVIDIA Build lists a free API endpoint subject to service conditions. The OpenMDW-1.1 license governs the model.

Model guide

NVIDIA Nemotron 3 Embed 1B: Multilingual Embeddings for Retrieval and RAG

NVIDIA Nemotron 3 Embed 1B is a 1.14-billion-parameter, BF16 text-embedding model designed for multilingual semantic search, retrieval, code search, agentic retrieval, and retrieval-augmented generation. It produces 2,048-dimensional dense vectors and supports text inputs up to 32,768 tokens.

What is NVIDIA Nemotron 3 Embed 1B?

NVIDIA Nemotron-3-Embed-1B-BF16 is a text-embedding model for converting written content into numerical vectors. These vectors represent the meaning of text, allowing software to compare documents, queries, code, or passages by semantic similarity rather than relying only on exact keyword matches.

The model is provided by NVIDIA and is available as downloadable weights on Hugging Face under the OpenMDW-1.1 license. It can also be accessed through NVIDIA NIM and NVIDIA's model-serving interfaces. The downloadable checkpoint is named nvidia/Nemotron-3-Embed-1B-BF16, while the NVIDIA API and NIM identifier is nvidia/nemotron-3-embed-1b.

With approximately 1.14 billion parameters, this is a relatively compact embedding model rather than a conversational language model. It does not generate answers, paragraphs, images, audio, or video. Its role is to create representations that other applications can use for retrieval, matching, ranking, and recommendation workflows.

Core specifications

SpecificationVerified detail
ProviderNVIDIA
Model familyNemotron 3 Embed
Model typeText embedding
ParametersApproximately 1.14 billion
PrecisionBF16
Embedding size2,048 dimensions
Maximum input length32,768 tokens
Evaluated language coverage34 languages
OutputDense numerical vectors
Release dateJuly 16, 2026
LicenseOpenMDW-1.1

The 2,048-dimensional output means that every supported text input is represented by a vector containing 2,048 numerical values. Applications typically store those vectors in a vector database and compare them with vectors generated from user queries or other documents.

What the model is used for

Nemotron 3 Embed 1B is intended for retrieval-oriented systems. Its main uses include semantic search, dense retrieval, retrieval-augmented generation (RAG), code search, agentic retrieval, and vector-based document matching.

  • Semantic search: Find documents that express a similar idea even when they do not contain the same words as the query.
  • RAG: Retrieve relevant passages from a private knowledge base before passing them to a separate generative model for answer writing.
  • Code search: Match natural-language descriptions with relevant functions, files, or code examples.
  • Agentic retrieval: Help an AI agent locate relevant information from tools, documents, or memory stores.
  • Document matching: Compare support tickets, policies, product descriptions, or other text collections.

The model supports query and passage modes for retrieval. In a typical search system, a passage mode is used to encode stored documents, while query mode is used for the user's search request. The resulting vectors can then be compared using the application's chosen similarity or distance method.

Multilingual coverage and technical design

NVIDIA documents evaluation across 34 languages, making the model suitable for multilingual retrieval collections and search interfaces. The supplied research does not establish that every language has identical quality, so teams should validate performance on their own languages, terminology, and document types before deployment.

The model is based on a pruned and distilled Ministral-3-3B-Instruct-2512 architecture. In practical terms, pruning and distillation are techniques used to produce a smaller, more efficient model from a larger or more capable starting point. That design helps position Nemotron 3 Embed 1B as a retrieval encoder that can offer a smaller deployment footprint than a much larger embedding model, while retaining a long input limit and multilingual functionality.

The 32,768-token input limit is useful for long documents or large passages, but it does not mean that sending an entire long document is always the best retrieval strategy. Search systems often achieve better precision by splitting documents into meaningful sections and attaching metadata before indexing them.

Capabilities and limitations

This model produces embeddings rather than natural-language responses. Its verified output type is a 2,048-dimensional embedding vector; it has no text, image, video, audio, music, speech, or action output. It accepts text input only, with no verified image, audio, or video input support.

Because it is not a generative assistant, the model should not be selected for conversation, long-form writing, direct question answering, summarization, code generation, or general-purpose reasoning. A separate language model is required when an application must turn retrieved passages into a natural-language response.

The supplied model data records no tool or function-calling support, no streaming support, and no JSON mode. Those features are generally associated with generative model responses and are not central to an embedding encoder. The model also has no documented knowledge cutoff in the supplied sources; knowledge-cutoff terminology is less directly applicable because the model's principal task is encoding text for retrieval rather than generating factual answers.

Fine-tuning and deployment options

NVIDIA's NeMo documentation supports full-weight fine-tuning and merged LoRA fine-tuning for the model. Fine-tuning can be relevant when the default embedding behavior does not align well with a specialized domain, such as internal legal terminology, scientific documents, technical support content, or a proprietary codebase. Any fine-tuning decision should be evaluated against the quality of the base model, the size and cleanliness of the training data, and the operational cost of maintaining a custom checkpoint.

There are two main supplied deployment routes. Developers can use the downloadable Hugging Face checkpoint, which provides more control over hosting and integration. Alternatively, NVIDIA NIM provides a serving path intended to simplify model deployment and inference through NVIDIA's infrastructure and APIs. The exact hardware, hosting, throughput, and operational requirements depend on the selected deployment configuration and are not specified in the supplied research.

Pricing and availability

No public per-token or per-request price was verified for the exact model. NVIDIA Build lists a free API endpoint subject to service conditions, but that should not be treated as an unlimited or permanently free production guarantee. Hosted usage, infrastructure, rate limits, and commercial support may have separate terms.

The downloadable Hugging Face weights are available under the OpenMDW-1.1 license. Before commercial deployment, users should review the license and the current terms for the chosen NVIDIA service. Self-hosting may avoid hosted inference charges, but it introduces hardware, scaling, monitoring, and maintenance costs.

When to choose Nemotron 3 Embed 1B

Choose Nemotron 3 Embed 1B when the main task is converting multilingual text into dense vectors for retrieval. It is especially relevant when a project needs a long 32,768-token input limit, 2,048-dimensional embeddings, query-and-passage retrieval modes, downloadable weights, or NVIDIA-oriented deployment options.

  • Choose it for multilingual semantic search across a large document collection.
  • Choose it as the retrieval component of a RAG pipeline, paired with a separate text-generation model.
  • Choose it for code or technical-document search when vector similarity is more useful than exact keyword matching.
  • Choose it when self-hosting or fine-tuning is important and the OpenMDW-1.1 licensing terms fit the project.
  • Choose it when a compact embedding model with a long context limit is preferable to a larger, potentially more expensive encoder.

Another embedding model may be more appropriate when an application requires a different vector dimension, a different language or domain specialization, a hosted pricing model with clearly published production rates, or independently verified benchmarks that better match the target workload. A generative language model is more appropriate when the system must answer users directly, reason over retrieved material, call tools, or produce structured text.

Performance and cost trade-offs

The supplied research does not provide verified speed, latency, or cost scores, so performance should not be described with numerical claims. The model's approximately 1.14-billion-parameter size and retrieval-focused design suggest a positioning between very small embedding encoders and substantially larger models, but actual throughput depends on hardware, batching, sequence length, quantization or precision choices, serving configuration, and workload shape.

Longer inputs can improve coverage for some documents but usually require more computation than short passages. A practical evaluation should therefore measure both retrieval quality and total indexing and query cost. Useful tests include multilingual queries, long and short documents, code-related searches, difficult near-duplicate queries, and RAG answer quality after retrieval.

Bottom line

NVIDIA Nemotron 3 Embed 1B is a specialized multilingual embedding model, not a chatbot. Its main distinguishing characteristics are a 2,048-dimensional dense-vector output, a 32,768-token input limit, support for 34 evaluated languages, query and passage retrieval modes, and availability through both NVIDIA deployment tools and downloadable weights. It is a strong candidate for semantic search, RAG, code retrieval, and agent memory when those requirements align with the project's language coverage, licensing, hosting, and evaluation needs.


Answers to Frequently Asked Questions

How can developers deploy NVIDIA Nemotron 3 Embed 1B?
Developers can download the Hugging Face checkpoint, nvidia/Nemotron-3-Embed-1B-BF16, under the OpenMDW-1.1 license, or use NVIDIA NIM and NVIDIA serving interfaces with the identifier nvidia/nemotron-3-embed-1b. NVIDIA also documents full-weight and merged LoRA fine-tuning options.
Can Nemotron 3 Embed 1B generate answers or summaries?
No. Nemotron 3 Embed 1B only produces text embeddings. A separate generative language model is required to write answers, summaries, code, or other natural-language responses from the retrieved content.
Does NVIDIA Nemotron 3 Embed 1B support multilingual retrieval and RAG?
Yes. NVIDIA documents evaluation across 34 languages, making the model suitable for multilingual search collections and RAG pipelines. However, teams should test retrieval quality for their specific languages, terminology, and document types.
What are the key specifications of Nemotron-3-Embed-1B-BF16?
The model has approximately 1.14 billion parameters, uses BF16 precision, produces 2,048-dimensional embeddings, supports inputs up to 32,768 tokens, and has been evaluated across 34 languages.
What is NVIDIA Nemotron 3 Embed 1B used for?
NVIDIA Nemotron 3 Embed 1B converts text into dense numerical vectors for semantic search, retrieval-augmented generation (RAG), code search, agentic retrieval, and document matching. It is an embedding model, not a chatbot or text-generation model.


Sources 6
Provider

About NVIDIA AI