⋈
Models by output type

Embeddings Explained: Vector Outputs for Search, Retrieval, and Similarity

An embedding is a machine-readable representation of content. Instead of returning a paragraph, image, or audio file, an embedding model usually returns an ordered list of numbers—a vector—that captures useful relationships between the input and other items. Similar content is intended to produce vectors that are close together in a mathematical space. This makes embeddings useful for semantic search, retrieval-augmented generation, recommendations, clustering, classification, and similarity detection. The term “project-global embeddings” is not a universal model category; in most contexts, “project” or “global” describes where vectors can be accessed, not a different kind of output.
What this means

Embeddings represent meaning as numbers rather than generating a normal human-readable answer. They are commonly used for semantic search, similarity matching and retrieval systems.

Embeddings models

60 models currently match this capability.

View all models →
◎
Amazon

Amazon Nova Multimodal Embeddings

Amazon Nova Multimodal Embeddings

Cross-modal semantic search, multimodal RAG, digital asset discovery, recommendations, classification, and clustering

Embeddings Other 8,172 ctx Image input Audio input Video input
View model →
◎
Amazon

Amazon Titan Embeddings G1 - Text

Titan Text Embeddings

Semantic search, vector indexing, retrieval-augmented generation, personalization, clustering, classification, and recommendation pipelines.

Embeddings Other 8,192 ctx
View model →
◎
Amazon

Amazon Titan Multimodal Embeddings G1

Amazon Titan

Multimodal search, text-to-image retrieval, image similarity, visual recommendations, personalization, and image-text matching.

Embeddings Multimodal 256 ctx Image input
View model →
◎
Cerebras

Cerebras Dragon-DocChat

Dragon-DocChat

Multi-turn document retrieval and retrieval-augmented generation pipelines

Embeddings Other
View model →
◎
OpenAI

CLIP

CLIP

Zero-shot image classification, image-text similarity, semantic image retrieval, multimodal indexing, and computer-vision research

Embeddings Multimodal 77 ctx Image input
View model →
◎
NAVER

clir-emb-dolphin

Dolphin

General-purpose semantic search, retrieval, document similarity, clustering, and text classification

Embeddings Embedding 500 ctx
View model →
◎
Naver Cloud

clir-sts-dolphin

Dolphin

Sentence similarity, semantic search, document relatedness, clustering, and text classification features

Embeddings Embedding 500 ctx
View model →
◎
Mistral AI

Codestral Embed

Codestral

Semantic code search, repository retrieval, coding-agent RAG, code similarity, duplicate detection, clustering, and code analytics

Embeddings Embedding 8,192 ctx
View model →
◎
Meta AI

DINOv3

DINOv3

Image embeddings, dense feature extraction, image retrieval, classification, segmentation, depth estimation, object discovery, video tracking pipelines, and geospatial computer vision

Embeddings Multimodal Image input
View model →
◎
Cohere

Embed English v3.0

Embed v3.0

English semantic search, retrieval-augmented generation, vector indexing, classification, clustering, and similarity matching

Embeddings Other 512 ctx Image input
View model →
◎
Cohere

Embed v4.0

Embed

Multilingual semantic search, multimodal RAG, PDF and document retrieval, image-to-text retrieval, classification, clustering, and enterprise vector indexing.

Embeddings Multimodal 128,000 ctx Image input
View model →
◎
Cohere

embed-english-light-v3.0

Embed v3.0

Fast, storage-efficient English semantic search, retrieval, classification, clustering, and large-scale embedding workloads

Embeddings Lightweight 512 ctx Image input
View model →
◎
Cohere

embed-multilingual-light-v3.0

Embed v3.0

Fast multilingual semantic search, cross-lingual retrieval, RAG, clustering, classification features, and compact vector indexes

Embeddings Lightweight 512 ctx Image input
View model →
◎
Cohere

embed-multilingual-v3.0

Embed v3.0

Multilingual semantic search, cross-lingual retrieval, RAG indexing, classification, clustering, and image-text similarity

Embeddings Other 512 ctx Image input
View model →
◎
Baidu

Embedding-V1

Embedding-V1

Semantic search, vector retrieval, recommendation, semantic matching, knowledge bases, and retrieval-augmented generation

Embeddings Embedding 384 ctx
View model →
◎
LG AI Research

EXAONE Path 2.0

EXAONE Path

Computational pathology research, whole-slide image representation learning, cancer biomarker prediction, and biomedical image analysis

Embeddings Other Image input
View model →
◎
Google DeepMind

Gemini Embedding

Gemini Embedding

Text semantic search, RAG retrieval, document matching, classification, clustering, and recommendation systems

Embeddings Other 2,048 ctx
View model →
◎
Google DeepMind

Gemini Embedding 2

Gemini Embedding

Cross-modal semantic search, multimodal RAG, vector retrieval, recommendations, classification, clustering, and indexing mixed text and media collections.

Embeddings Embedding 8,192 ctx Image input Audio input Video input
View model →
◎
IBM

Granite-Embedding-107M-Multilingual

Granite Embedding

Low-cost multilingual semantic search, cross-lingual retrieval, vector databases, similarity matching, and RAG pipelines

Embeddings Embedding 512 ctx
View model →
◎
IBM

Granite-Embedding-125M-English

Granite Embedding

English semantic search, vector retrieval, RAG, similarity matching, enterprise knowledge-base search, and local embedding deployment

Embeddings Embedding 512 ctx
View model →
◎
IBM

Granite-Embedding-278M-Multilingual

Granite Embeddings

Multilingual semantic search, vector retrieval, RAG, clustering, similarity matching, and text classification features

Embeddings Embedding 512 ctx
View model →
◎
IBM

granite-embedding-30m-english

Granite Embedding

Fast, low-footprint English semantic search, RAG retrieval, similarity matching, and vector indexing

Embeddings Embedding 512 ctx
View model →
◎
IBM

Granite-Embedding-311M-Multilingual-R2

Granite Embedding

Multilingual semantic search, retrieval-augmented generation, cross-lingual retrieval, long-document search, similarity, and code retrieval

Embeddings Other 32,768 ctx
View model →
◎
IBM

Granite-Embedding-97M-Multilingual-R2

Granite Embedding

Low-latency multilingual semantic search, retrieval-augmented generation, vector search, document similarity, long-document retrieval, and cross-lingual code retrieval

Embeddings Lightweight 32,768 ctx
View model →
◎
IBM

granite-embedding-english-r2

Granite Embedding

English semantic search, vector retrieval, retrieval-augmented generation, document similarity, clustering, and enterprise information retrieval

Embeddings Embedding 8,192 ctx
View model →
◎
IBM

granite-embedding-small-english-r2

Granite Embedding R2

Compact English semantic search, retrieval-augmented generation, document similarity, and private vector-search deployments

Embeddings Embedding 8,192 ctx
View model →
◎
Tencent

Kinfra-Text-Embedding-0.6b

Kinfra

High-volume semantic retrieval, vector search, FAQ matching, text clustering, classification, and cost- or latency-sensitive knowledge-base applications

Embeddings Embedding 32,768 ctx
View model →
◎
Tencent

Kinfra-Text-Embedding-4b

Kinfra

High-quality multilingual semantic search, retrieval-augmented generation, enterprise knowledge bases, similarity matching, and text classification.

Embeddings Embedding 32,000 ctx
View model →
◎
Tencent

Kinfra-VL-Embedding-2b

Kinfra

Fast cross-modal image-text retrieval, multimodal semantic matching, and video search

Embeddings Multimodal 32,768 ctx Image input Video input
View model →
◎
Tencent

Kinfra-VL-Embedding-8b

Kinfra-VL-Embedding

High-precision multimodal retrieval, cross-modal image-text search, video search, and semantic matching across text, image, and video collections

Embeddings Multimodal Embedding 32,768 ctx Image input Video input
View model →
◎
NVIDIA

Llama Nemotron Embed VL 1B v2

Llama Nemotron Embed VL

Multimodal semantic search, visual document retrieval, question-answer retrieval, vector databases, and retrieval-augmented generation

Embeddings Embedding 10,240 ctx Image input
View model →
◎
Aleph Alpha

Luminous-Explore

Luminous

Semantic search, information retrieval, query-document matching, clustering, classification, similarity scoring, and text feature extraction

Embeddings Embedding
View model →
◎
Microsoft

MedImageInsight Premium

MedImageInsight

Medical image and text embeddings, similarity search, multimodal retrieval, downstream classification, outlier detection, dataset curation, and healthcare AI development workflows

Embeddings Multimodal Image input
View model →
◎
Mistral AI

Mistral Embed

Mistral Embed

Semantic search, retrieval-augmented generation, vector databases, document classification, clustering, duplicate detection and general text retrieval

Embeddings Other 8,192 ctx
View model →
◎
NVIDIA

MolMIM

MolMIM

Small-molecule generation, molecular embeddings, chemical-space exploration, lead optimization, and oracle-guided drug-design workflows

Embeddings Other 128 ctx
View model →
◎
Moonshot AI

MoonViT-SO-400M

MoonViT

Native-resolution image feature extraction, vision-language model backbones, high-resolution document and image understanding, and multimodal research

Embeddings Other Image input
View model →
◎
Alibaba Cloud

multimodal-embedding-v1

Multimodal Embedding

Cross-modal retrieval, text-to-image search, image similarity, video search, semantic classification, clustering, and multimodal vector indexing.

Embeddings Multimodal Embedding 512 ctx Image input Video input
View model →
◎
NVIDIA

Nemotron-3-Embed-1B-BF16

Nemotron 3 Embed

Multilingual semantic search, dense retrieval, RAG, agentic retrieval, code search, and vector-based document matching

Embeddings Embedding 32,768 ctx
View model →
◎
Allen Institute for AI

OlmoEarth-v1_2-Base

OlmoEarth v1.2

Satellite-image and Earth-observation embeddings, remote-sensing representation learning, geospatial classification, segmentation, and downstream fine-tuning

Embeddings Multimodal Image input
View model →
◎
Allen Institute for AI

OlmoEarth-v1_2-Nano

OlmoEarth v1.2

Efficient Earth observation embeddings, satellite image and time-series representation learning, geospatial classification, segmentation, and large-scale remote-sensing inference

Embeddings Multimodal Image input
View model →
◎
Allen Institute for AI

OlmoEarth-v1_2-Small

OlmoEarth v1.2

Satellite-image embeddings, remote-sensing representation learning, geospatial segmentation, land-cover analysis, and Earth observation research

Embeddings Multimodal Image input
View model →
◎
Allen Institute for AI

OlmoEarth-v1_2-Tiny

OlmoEarth

Efficient satellite-image and Earth-observation embeddings, remote-sensing research, and downstream classification or segmentation

Embeddings Multimodal Image input
View model →
◎
Meta

Omnilingual wav2vec 2.0

Omnilingual wav2vec 2.0

Multilingual speech representation learning, audio embeddings, low-resource language research, and custom downstream speech systems

Embeddings Other Audio input
View model →
◎
Meta

Perception Encoder Audiovisual

Perception Encoder

Cross-modal audio-video-text retrieval, audiovisual embeddings, sound-event understanding, media indexing, and multimodal perception systems.

Embeddings Multimodal Audio input Video input
View model →
◎
Aleph Alpha

Pharia-1-Embedding-4608-control

Pharia-1 Embedding

Multilingual information retrieval, semantic search, reranking, clustering, and instruction-guided text embeddings

Embeddings Embedding 2,048 ctx
View model →
◎
Aleph Alpha

Pharia-1-Embedding-4608-control-256

Pharia-1 Embedding

Compact vector representations for semantic search, information retrieval, reranking, clustering, and similarity-based classification

Embeddings Other 2,048 ctx
View model →
◎
Alibaba Cloud

qwen3.7-text-embedding

Qwen3.7

Multilingual semantic search, retrieval-augmented generation, code retrieval, recommendation, clustering, classification, and large-scale text vectorization

Embeddings Embedding 131,072 ctx
View model →
◎
ByteDance

Seed-1.6-Embedding

Seed1.6

Cross-modal semantic search, text-image retrieval, video retrieval, multimodal knowledge bases, classification, clustering, and recommendation

Embeddings Other 128,000 ctx Image input Video input
View model →
◎
IBM

slate-125m-english-rtrvr-v2

Slate

English semantic search, vector database indexing, retrieval-augmented generation, document matching, and query-passage retrieval

Embeddings Embedding 512 ctx
View model →
◎
IBM

slate-30m-english-rtrvr-v2

Slate

English semantic search, dense retrieval, vector indexing, duplicate-question matching, and lightweight retrieval-augmented generation

Embeddings Other 512 ctx
View model →
◎
OpenAI

text-embedding-3-large

text-embedding-3

High-quality semantic search, multilingual retrieval, RAG, recommendations, clustering, classification and similarity matching

Embeddings Embedding 8,192 ctx
View model →
◎
OpenAI

text-embedding-3-small

text-embedding-3

Cost-efficient semantic search, retrieval-augmented generation, clustering, recommendations, anomaly detection, and text or code similarity

Embeddings Embedding 8,192 ctx
View model →
◎
OpenAI

text-embedding-ada-002

text-embedding-ada-002

Legacy semantic search, retrieval, clustering, recommendations, anomaly detection, and classification systems already built around ada-002 vectors

Embeddings Embedding 8,192 ctx
View model →
◎
Amazon

Titan Text Embeddings V2

Titan Text Embeddings

Cost-efficient text embeddings for semantic search, RAG, document retrieval, classification, clustering, reranking, and recommendations

Embeddings Embedding 8,192 ctx
View model →
◎
Alibaba Cloud

tongyi-embedding-vision-flash

Embedding-Vision

Cost-sensitive cross-modal retrieval, image and video search, multimedia catalog indexing, and vector search

Embeddings Multimodal 1,024 ctx Image input Video input
View model →
◎
Alibaba Cloud

qwen3-vl-embedding

Qwen3-VL-Embedding

Multimodal vector search, cross-modal retrieval, image and video search, semantic clustering, tagging, and retrieval pipelines combining text with visual content.

Embeddings Multimodal 32,000 ctx Image input Video input
View model →
◎
Alibaba Cloud

qwen3.7-text-embedding-flash

Qwen3.7-Text-Embedding

Cost-sensitive, high-throughput multilingual text vectorization, semantic search, RAG, recommendation, clustering, and classification.

Embeddings Lightweight 131,072 ctx
View model →
◎
Alibaba Cloud Model Studio

text-embedding-v3

Qwen3-Embedding

Multilingual text embeddings, semantic search, RAG, recommendation, clustering, classification, and migration of existing v3 vector indexes

Embeddings Embedding 8,192 ctx
View model →
◎
Alibaba Cloud

text-embedding-v4

Qwen3-Embedding

Multilingual semantic search, RAG pipelines, vector databases, document retrieval, clustering, classification, recommendation, and code retrieval

Embeddings Other 8,192 ctx
View model →
◎
Alibaba Cloud

tongyi-embedding-vision-plus

Tongyi Embedding Vision

Cross-modal retrieval, image and video similarity search, multimodal semantic indexing, recommendation, and content classification

Embeddings Multimodal Embedding Image input Video input
View model →
Learn more

About embedding models

What an embedding model produces

An embedding model converts an input such as text, code, an image, audio, video, or a document into a numerical vector. A vector is simply an ordered list of values, commonly floating-point numbers. An API may return it in a field named embedding or values.

The vector is not normally intended for people to read. It is an intermediate representation that software can compare mathematically. For example, a search system may convert a collection of documents and a user query into vectors, then look for documents whose vectors are closest to the query vector.

A useful analogy is a map. The map does not provide a written explanation of every location, but it places related locations near one another. Similarly, an embedding space attempts to place semantically related items near one another, although the individual coordinates usually do not have simple human-readable meanings.

What “project-global” usually means

Project-global embeddings is not a broadly standardized name for a special embedding format or model family. The phrase most plausibly describes the access scope of vectors within an application.

  • Project-scoped: vectors and their metadata are available only within one project, workspace, repository, tenant, or application.
  • Organization-wide: vectors can be shared across approved teams or projects within an organization.
  • Global: vectors may be searchable across several projects or applications, subject to permissions and governance.

This scope affects indexing, permissions, privacy, deletion, and tenancy. It does not inherently change the vector’s mathematical structure or prove that a different embedding model was used. If a platform uses “global” as an internal product term, its documentation should define the exact boundary.

Input and output are different capabilities

An embedding model may accept one or more modalities, but that does not mean it generates those modalities. A text embedding model takes text and returns a vector. A multimodal embedding model may accept text, images, video, audio, or documents and map them into a shared vector space, while still returning vectors rather than media.

For example, a cross-modal search system might embed a written query and product images into the same space. The system can then find images related to the query. The embedding model has enabled image retrieval, but it has not generated a new image.

This distinction separates embeddings from generative and transformation capabilities:

  • Text generation returns readable language.
  • Image, video, or audio generation returns media or a media stream.
  • Speech recognition converts audio into text.
  • Classification returns labels or probabilities.
  • Reranking scores an existing set of search results.
  • Embeddings return vectors for comparison and downstream computation.

How embeddings support search and retrieval

In a typical semantic-search workflow, documents are divided into useful sections, sometimes called chunks. An embedding model converts each chunk into a vector, and the application stores those vectors in a vector database or approximate-nearest-neighbor index.

When a user submits a query, the application embeds the query with the same or a compatible model. It then compares the query vector with stored vectors using a measure such as cosine similarity, dot product, or Euclidean distance. The closest matches are returned as candidate results.

In a retrieval-augmented generation system, the retrieved documents are usually passed to a separate language model as context. The language model may write the final answer, but it is not necessarily the component that created the embeddings or performed the vector search.

This architecture matters when evaluating claims about native capability. A chat application may offer “memory,” “semantic search,” or “document Q&A” through a combination of an embedding service, a database, permission controls, retrieval code, and a generative model. Those product features should not automatically be treated as native embedding output from the conversational model itself.

Practical uses

Semantic search

Embeddings can match meaning rather than only exact words. A query such as “ways to reduce cloud spending” may retrieve a document titled “controlling infrastructure costs,” even though the wording differs. Exact identifiers, numbers, and rare names may still require keyword or hybrid search.

Retrieval-augmented generation

An organization can embed internal documents, retrieve relevant passages for a question, and provide those passages to a language model. The vectors help locate potentially relevant material; they do not guarantee that the final answer is correct or that the retrieved source is authoritative.

Cross-modal search

Multimodal embeddings can connect different types of content. A retailer might embed product descriptions and product photos so a text query can find visually associated items. A media archive might search images, video segments, transcripts, or audio using a shared representation when the selected model supports it.

Recommendations and similarity

Products, articles, songs, users, or support tickets can be represented as vectors. An application can then find items with similar representations and use those relationships to suggest related content or detect near-duplicates.

Clustering and classification

Vectors can help organize a large collection into groups or serve as features for a separate classifier. For example, support requests may be grouped by topic before an operations team defines categories and routing rules.

What matters when comparing embedding models

Embedding models should be compared on the retrieval or similarity task they need to support, not simply on the size of the underlying language model. Important criteria include:

  • Task quality: performance on representative queries and labeled relevant results.
  • Modality support: whether the model handles text, code, images, audio, video, documents, or cross-modal comparisons.
  • Language and domain coverage: multilingual performance and suitability for fields such as law, medicine, science, or software.
  • Vector dimensions: larger vectors may preserve useful information but increase storage, indexing, and comparison costs.
  • Input limits: maximum text or document length, truncation behavior, and whether chunking is required.
  • Task instructions: some models distinguish query and document embeddings or expect task-specific instructions and prefixes.
  • Similarity compatibility: the metric, normalization method, and index configuration recommended by the provider.
  • Latency and throughput: especially important when indexing millions of records or serving interactive search.
  • Pricing: costs may depend on tokens, characters, requests, batch processing, or hosting infrastructure.
  • Output controls: configurable dimensions, batching, normalization, and precision options.
  • Version stability: changing models can make existing vectors incompatible or reduce retrieval quality, often requiring a full re-embedding process.
  • Privacy and scope: data isolation, access controls, retention, deletion, and regional processing.

The most useful evaluation uses real queries and measures results with metrics such as recall, precision, mean reciprocal rank, or nDCG. Test synonyms, paraphrases, multilingual input, long documents, exact names, numbers, permission boundaries, and difficult or ambiguous queries.

Limitations and trade-offs

Embeddings represent relationships learned by a model; they do not preserve every fact or detail in a source. Similarity is also task-dependent. Two items may be close because they share a topic or writing style without being factually equivalent.

Long documents commonly need to be split into chunks. Poor chunk boundaries can separate important context, while excessively large chunks can reduce retrieval precision. Metadata filters, hybrid keyword search, and reranking may be needed to handle exact terms, dates, identifiers, or access restrictions that vector similarity alone misses.

Vectors are difficult to interpret directly. If a result seems wrong, the cause may be the embedding model, chunking strategy, query formulation, similarity metric, index, metadata filter, or reranker. A vector database also does not automatically solve these integration problems; it generally stores and searches vectors produced elsewhere, although some platforms bundle the steps together.

Global sharing introduces additional governance risks. A vector may reveal useful information about a document even when the original document is not returned, and weak tenant or permission controls can expose results across projects. Applications should apply authorization checks during retrieval, define deletion behavior, and decide whether shared indexes are appropriate for sensitive data.

When do you need an embedding model?

You probably need embeddings when an application must find, compare, group, or recommend items based on meaning or learned similarity. Typical signals include a document search system that must understand paraphrases, a RAG pipeline over a private knowledge base, a recommendation workflow, or a multimodal archive that needs text-to-image or text-to-video retrieval.

You may not need embeddings for a simple exact-match lookup, a fixed set of categories, or a task where a conventional database query is more precise and easier to audit. In many production systems, the best approach combines embeddings with keyword search, structured filters, reranking, and explicit permission checks.

The practical definition

For this model-output category, the safest definition is: an embedding is a numerical vector used by software to measure relationships between inputs. “Project-global” should generally be treated as a scope qualifier describing where those vectors can be used or shared, unless a particular platform documents a more specific meaning.

The vector is rarely the final user-facing result. Its value comes from the system built around it: the embedding model, indexing strategy, similarity metric, retrieval logic, access controls, and any later reranking or answer-generation step.