What is NVIDIA Nemotron 3 Embed 1B?
NVIDIA Nemotron-3-Embed-1B-BF16 is a text-embedding model for converting written content into numerical vectors. These vectors represent the meaning of text, allowing software to compare documents, queries, code, or passages by semantic similarity rather than relying only on exact keyword matches.
The model is provided by NVIDIA and is available as downloadable weights on Hugging Face under the OpenMDW-1.1 license. It can also be accessed through NVIDIA NIM and NVIDIA's model-serving interfaces. The downloadable checkpoint is named nvidia/Nemotron-3-Embed-1B-BF16, while the NVIDIA API and NIM identifier is nvidia/nemotron-3-embed-1b.
With approximately 1.14 billion parameters, this is a relatively compact embedding model rather than a conversational language model. It does not generate answers, paragraphs, images, audio, or video. Its role is to create representations that other applications can use for retrieval, matching, ranking, and recommendation workflows.
Core specifications
| Specification | Verified detail |
|---|---|
| Provider | NVIDIA |
| Model family | Nemotron 3 Embed |
| Model type | Text embedding |
| Parameters | Approximately 1.14 billion |
| Precision | BF16 |
| Embedding size | 2,048 dimensions |
| Maximum input length | 32,768 tokens |
| Evaluated language coverage | 34 languages |
| Output | Dense numerical vectors |
| Release date | July 16, 2026 |
| License | OpenMDW-1.1 |
The 2,048-dimensional output means that every supported text input is represented by a vector containing 2,048 numerical values. Applications typically store those vectors in a vector database and compare them with vectors generated from user queries or other documents.
What the model is used for
Nemotron 3 Embed 1B is intended for retrieval-oriented systems. Its main uses include semantic search, dense retrieval, retrieval-augmented generation (RAG), code search, agentic retrieval, and vector-based document matching.
- Semantic search: Find documents that express a similar idea even when they do not contain the same words as the query.
- RAG: Retrieve relevant passages from a private knowledge base before passing them to a separate generative model for answer writing.
- Code search: Match natural-language descriptions with relevant functions, files, or code examples.
- Agentic retrieval: Help an AI agent locate relevant information from tools, documents, or memory stores.
- Document matching: Compare support tickets, policies, product descriptions, or other text collections.
The model supports query and passage modes for retrieval. In a typical search system, a passage mode is used to encode stored documents, while query mode is used for the user's search request. The resulting vectors can then be compared using the application's chosen similarity or distance method.
Multilingual coverage and technical design
NVIDIA documents evaluation across 34 languages, making the model suitable for multilingual retrieval collections and search interfaces. The supplied research does not establish that every language has identical quality, so teams should validate performance on their own languages, terminology, and document types before deployment.
The model is based on a pruned and distilled Ministral-3-3B-Instruct-2512 architecture. In practical terms, pruning and distillation are techniques used to produce a smaller, more efficient model from a larger or more capable starting point. That design helps position Nemotron 3 Embed 1B as a retrieval encoder that can offer a smaller deployment footprint than a much larger embedding model, while retaining a long input limit and multilingual functionality.
The 32,768-token input limit is useful for long documents or large passages, but it does not mean that sending an entire long document is always the best retrieval strategy. Search systems often achieve better precision by splitting documents into meaningful sections and attaching metadata before indexing them.
Capabilities and limitations
This model produces embeddings rather than natural-language responses. Its verified output type is a 2,048-dimensional embedding vector; it has no text, image, video, audio, music, speech, or action output. It accepts text input only, with no verified image, audio, or video input support.
Because it is not a generative assistant, the model should not be selected for conversation, long-form writing, direct question answering, summarization, code generation, or general-purpose reasoning. A separate language model is required when an application must turn retrieved passages into a natural-language response.
The supplied model data records no tool or function-calling support, no streaming support, and no JSON mode. Those features are generally associated with generative model responses and are not central to an embedding encoder. The model also has no documented knowledge cutoff in the supplied sources; knowledge-cutoff terminology is less directly applicable because the model's principal task is encoding text for retrieval rather than generating factual answers.
Fine-tuning and deployment options
NVIDIA's NeMo documentation supports full-weight fine-tuning and merged LoRA fine-tuning for the model. Fine-tuning can be relevant when the default embedding behavior does not align well with a specialized domain, such as internal legal terminology, scientific documents, technical support content, or a proprietary codebase. Any fine-tuning decision should be evaluated against the quality of the base model, the size and cleanliness of the training data, and the operational cost of maintaining a custom checkpoint.
There are two main supplied deployment routes. Developers can use the downloadable Hugging Face checkpoint, which provides more control over hosting and integration. Alternatively, NVIDIA NIM provides a serving path intended to simplify model deployment and inference through NVIDIA's infrastructure and APIs. The exact hardware, hosting, throughput, and operational requirements depend on the selected deployment configuration and are not specified in the supplied research.
Pricing and availability
No public per-token or per-request price was verified for the exact model. NVIDIA Build lists a free API endpoint subject to service conditions, but that should not be treated as an unlimited or permanently free production guarantee. Hosted usage, infrastructure, rate limits, and commercial support may have separate terms.
The downloadable Hugging Face weights are available under the OpenMDW-1.1 license. Before commercial deployment, users should review the license and the current terms for the chosen NVIDIA service. Self-hosting may avoid hosted inference charges, but it introduces hardware, scaling, monitoring, and maintenance costs.
When to choose Nemotron 3 Embed 1B
Choose Nemotron 3 Embed 1B when the main task is converting multilingual text into dense vectors for retrieval. It is especially relevant when a project needs a long 32,768-token input limit, 2,048-dimensional embeddings, query-and-passage retrieval modes, downloadable weights, or NVIDIA-oriented deployment options.
- Choose it for multilingual semantic search across a large document collection.
- Choose it as the retrieval component of a RAG pipeline, paired with a separate text-generation model.
- Choose it for code or technical-document search when vector similarity is more useful than exact keyword matching.
- Choose it when self-hosting or fine-tuning is important and the OpenMDW-1.1 licensing terms fit the project.
- Choose it when a compact embedding model with a long context limit is preferable to a larger, potentially more expensive encoder.
Another embedding model may be more appropriate when an application requires a different vector dimension, a different language or domain specialization, a hosted pricing model with clearly published production rates, or independently verified benchmarks that better match the target workload. A generative language model is more appropriate when the system must answer users directly, reason over retrieved material, call tools, or produce structured text.
Performance and cost trade-offs
The supplied research does not provide verified speed, latency, or cost scores, so performance should not be described with numerical claims. The model's approximately 1.14-billion-parameter size and retrieval-focused design suggest a positioning between very small embedding encoders and substantially larger models, but actual throughput depends on hardware, batching, sequence length, quantization or precision choices, serving configuration, and workload shape.
Longer inputs can improve coverage for some documents but usually require more computation than short passages. A practical evaluation should therefore measure both retrieval quality and total indexing and query cost. Useful tests include multilingual queries, long and short documents, code-related searches, difficult near-duplicate queries, and RAG answer quality after retrieval.
Bottom line
NVIDIA Nemotron 3 Embed 1B is a specialized multilingual embedding model, not a chatbot. Its main distinguishing characteristics are a 2,048-dimensional dense-vector output, a 32,768-token input limit, support for 34 evaluated languages, query and passage retrieval modes, and availability through both NVIDIA deployment tools and downloadable weights. It is a strong candidate for semantic search, RAG, code retrieval, and agent memory when those requirements align with the project's language coverage, licensing, hosting, and evaluation needs.

