Dolphin

clir-emb-dolphin

by NAVER AI · Current and accessible through CLOVA Studio's Embedding API

clir-emb-dolphin is NAVER Cloud Platform’s general-purpose text embedding model for CLOVA Studio. It accepts up to 500 tokens and returns 1,024-dimensional vectors, with inner-product similarity recommended for semantic search, retrieval, document similarity, clustering, and classification.

Embeddings Reasoning Coding
clir-emb-dolphin is a text embedding model from NAVER Cloud Platform’s CLOVA Studio. Rather than generating prose or other media, it converts text into numerical representations that applications can compare by meaning. It accepts between 1 and 500 tokens per request and returns a vector containing 1,024 floating-point values. NAVER positions it as a broadly generalizable model and the default model for the classic Embedding API.
Outputs

What clir-emb-dolphin can produce

Embeddings
Inputs

What it can understand

Text
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
Specifications

Technical details

Model family Dolphin
Model type Embedding
Context window 500 tokens
Status Current and accessible through CLOVA Studio's Embedding API
Knowledge cutoff notes

Knowledge-cutoff information is not applicable to this embedding endpoint in the referenced official documentation, and NAVER does not publish a cutoff date for the exact model.

Model notes

clir-emb-dolphin is a text embedding model rather than a generative language model. The CLOVA Studio Embedding API accepts 1 to 500 tokens per request and returns a list of 1,024 floating-point values. NAVER recommends inner-product similarity for this model. The documentation describes it as highly generalizable across domains and identifies it as the default model for the classic Embedding API. NAVER's current public documentation does not provide a model-specific release date, knowledge cutoff, or separate public input/output price for this exact model. The 500-token value is an API input limit, not a conventional generative context window.

Model guide

clir-emb-dolphin: NAVER’s General-Purpose Embedding Model for Semantic Search

clir-emb-dolphin is NAVER Cloud Platform’s general-purpose text embedding model. Available through the CLOVA Studio Embedding API, it converts text of up to 500 tokens into 1,024-dimensional vectors for semantic search, retrieval, document similarity, clustering, and classification.

What clir-emb-dolphin does

clir-emb-dolphin is designed to turn text into vectors, also called embeddings. A vector is a numerical representation of the meaning and characteristics of a piece of text. When two texts have related meanings, their vectors can be close to one another according to a selected similarity calculation.

This makes the model useful as a component inside a larger application rather than as a conversational or generative assistant. For example, a search system can embed a collection of documents, store the resulting vectors, embed a user's query, and retrieve documents whose vectors are most similar to the query. The model produces the representations; the surrounding application performs storage, similarity search, ranking, and any final response generation.

NAVER Cloud Platform provides clir-emb-dolphin through the CLOVA Studio Embedding API. NAVER describes it as a general-purpose model that can be applied across multiple domains. In the classic Embedding API, it is identified as the default embedding model.

Verified technical specifications

SpecificationVerified detail
Model typeText embedding model
ProviderNAVER Cloud Platform
InputText containing 1 to 500 tokens per request
OutputA vector with 1,024 floating-point values
Recommended similarity calculationInner product, also called dot product or scalar product
AccessCLOVA Studio Embedding API
EnvironmentsCLOVA Studio Classic and VPC environments

The 500-token figure is an input limit for this embedding endpoint. It should not be interpreted as the context window of a generative language model. The supplied documentation does not specify a maximum generated-token limit because clir-emb-dolphin does not generate text.

How the embedding workflow works

A typical implementation has two stages. First, documents are divided into suitable passages and each passage is sent to clir-emb-dolphin. The application stores each returned 1,024-dimensional vector together with the original passage and metadata such as its title, source, or access permissions.

When a user submits a search query, the application sends the query to the same model. It then compares the query vector with the stored document vectors using inner-product similarity, the metric recommended by NAVER for this model. The highest-scoring passages can be returned as search results or passed to another system for retrieval-augmented generation.

Using the same embedding model and similarity method for both indexing and querying is important. Mixing models can place documents and queries in incompatible vector spaces, while changing similarity methods can alter ranking behavior. These integration choices belong to the application rather than to the model itself.

The 500-token limit and document chunking

Each request can contain up to 500 tokens. Long documents therefore need to be divided before they are embedded. A useful chunk should normally preserve a coherent idea, such as a paragraph, a short group of paragraphs, a section, or a complete passage that can stand on its own in search results.

Arbitrarily cutting text can weaken retrieval because a definition may be separated from its subject or a qualification may be detached from the statement it qualifies. At the same time, very large chunks can make search results less precise by combining several unrelated topics. The supplied research does not prescribe a particular chunk size below the 500-token maximum, so the best choice should be evaluated against the structure of the source documents and the application's retrieval requirements.

Metadata and source text are not part of the vector output itself. The application must retain them separately if users need citations, document titles, filtering, or access-control checks.

Primary use cases

clir-emb-dolphin can support search based on meaning instead of exact word overlap. A query about cancelling a reservation, for example, may retrieve a passage that uses different wording but discusses the same policy. This is particularly useful when users use informal language or when documents use terminology that differs from the search query.

Retrieval-augmented generation

In a retrieval-augmented generation workflow, the model can identify relevant passages from a private knowledge base. Those passages may then be supplied to a separate text-generation model. clir-emb-dolphin does not write the final answer, summarize the retrieved material, or perform tool calls; it supplies the vector-based retrieval layer.

Document and passage similarity

Applications can compare vectors to identify related documents, duplicate or near-duplicate content, and thematically similar passages. The model can therefore be used for content discovery, recommendation logic, and document organization when the application defines the comparison and business rules.

Clustering and classification

Embedding vectors can be grouped to discover themes or organized into categories using downstream clustering and classification methods. clir-emb-dolphin provides the input representation, but it does not itself return a category label or execute a complete classification pipeline. The surrounding system must supply the clustering algorithm, classifier, thresholds, and evaluation process.

Strengths and trade-offs

The clearest strength of clir-emb-dolphin is its broad positioning. NAVER describes it as highly generalizable across domains, making it a reasonable starting point when an application needs one embedding model for ordinary search, retrieval, similarity, and organization tasks.

Its 1,024-dimensional output provides a fixed representation that can be indexed by a vector database or another similarity-search system. The explicit recommendation to use inner-product similarity also gives implementers a documented default for comparing vectors.

There are practical trade-offs. The model is limited to 500 input tokens per request, so long-document systems need chunking and additional storage or processing logic. It also returns representations rather than answers. It cannot replace a generative model when the application needs explanations, summaries, dialogue, code, or other generated content.

The supplied research does not publish a model-specific benchmark, release date, knowledge cutoff, or separate public input and output price for clir-emb-dolphin. CLOVA Studio usage is subject to the applicable service pricing and usage terms. Cost should therefore be confirmed in NAVER Cloud Platform's current commercial documentation rather than inferred from the model's vector size.

Modalities and unsupported generative capabilities

clir-emb-dolphin accepts text input and produces numerical embedding output. It does not produce text, images, audio, video, or music. It is not documented as supporting image, audio, or video input, and the supplied research does not identify tool use, function calling, streaming generation, structured JSON output, caching, batching, or fine-tuning features for this endpoint.

Reasoning and coding are not meaningful primary capabilities for this model. It can represent text about reasoning or code for similarity and retrieval purposes, but it does not execute reasoning workflows, run programs, or generate code. Those tasks require a separate model or application component.

NAVER describes clir-emb-dolphin as the broadly applicable option in the classic Embedding API. The related clir-sts-dolphin model is described as being focused more specifically on precise semantic textual similarity. That distinction matters when the main task is comparing sentence meaning rather than building a general retrieval or document-organization system.

NAVER also documents Embedding v2, which uses the open-source bge-m3 model and supports inputs of up to 8,192 tokens. A longer input limit may be more suitable for applications that need to process larger passages with less chunking. However, the supplied research does not provide a direct benchmark or pricing comparison between these models. Indexes should not mix their vectors without a deliberate migration or compatibility strategy.

When to choose clir-emb-dolphin

Choose clir-emb-dolphin when you need a general-purpose text representation for semantic search, retrieval, document similarity, clustering, or downstream classification and your content can be divided into passages of no more than 500 tokens. It is especially appropriate when you want the default model for the classic CLOVA Studio Embedding API and are prepared to use inner-product similarity in your vector-search workflow.

Consider clir-sts-dolphin when precise sentence-level semantic similarity is the central requirement. Consider Embedding v2 when the documented 8,192-token input capacity is more valuable than using clir-emb-dolphin's general-purpose positioning. A generative language model is more appropriate when the system must produce answers, summaries, explanations, or code after retrieval.

In practical terms, clir-emb-dolphin is best viewed as a focused retrieval and text-representation component. Its value comes from how consistently an application chunks content, indexes vectors, applies similarity search, and uses the retrieved material—not from standalone conversation or generation.


Answers to Frequently Asked Questions

When should an application choose clir-emb-dolphin instead of another NAVER embedding model?
Choose clir-emb-dolphin for general-purpose semantic search, retrieval, similarity, clustering, or classification when text can be divided into passages of up to 500 tokens. clir-sts-dolphin may be preferable for precise sentence-level similarity, while Embedding v2 may be more suitable when inputs up to 8,192 tokens are important.
Does clir-emb-dolphin generate answers or text?
No. clir-emb-dolphin only produces text embeddings. A surrounding application must handle vector storage, similarity search, ranking, and any answer generation, summarization, or dialogue using a separate model if needed.
Which similarity method should be used with clir-emb-dolphin?
NAVER recommends inner-product similarity, also known as dot product or scalar product. Applications should use the same model and similarity method when embedding documents for indexing and queries for retrieval.
What is clir-emb-dolphin used for?
clir-emb-dolphin is a general-purpose text embedding model from NAVER Cloud Platform. It converts text into numerical vectors for semantic search, document and passage similarity, retrieval-augmented generation, clustering, and downstream classification.
What are the input and output limits of clir-emb-dolphin?
Each request can contain between 1 and 500 tokens. The model returns a vector containing 1,024 floating-point values. The 500-token limit applies to each embedding request and is not a generative model context window.


Sources 4
Provider

About NAVER AI