What clir-emb-dolphin does
clir-emb-dolphin is designed to turn text into vectors, also called embeddings. A vector is a numerical representation of the meaning and characteristics of a piece of text. When two texts have related meanings, their vectors can be close to one another according to a selected similarity calculation.
This makes the model useful as a component inside a larger application rather than as a conversational or generative assistant. For example, a search system can embed a collection of documents, store the resulting vectors, embed a user's query, and retrieve documents whose vectors are most similar to the query. The model produces the representations; the surrounding application performs storage, similarity search, ranking, and any final response generation.
NAVER Cloud Platform provides clir-emb-dolphin through the CLOVA Studio Embedding API. NAVER describes it as a general-purpose model that can be applied across multiple domains. In the classic Embedding API, it is identified as the default embedding model.
Verified technical specifications
| Specification | Verified detail |
|---|---|
| Model type | Text embedding model |
| Provider | NAVER Cloud Platform |
| Input | Text containing 1 to 500 tokens per request |
| Output | A vector with 1,024 floating-point values |
| Recommended similarity calculation | Inner product, also called dot product or scalar product |
| Access | CLOVA Studio Embedding API |
| Environments | CLOVA Studio Classic and VPC environments |
The 500-token figure is an input limit for this embedding endpoint. It should not be interpreted as the context window of a generative language model. The supplied documentation does not specify a maximum generated-token limit because clir-emb-dolphin does not generate text.
How the embedding workflow works
A typical implementation has two stages. First, documents are divided into suitable passages and each passage is sent to clir-emb-dolphin. The application stores each returned 1,024-dimensional vector together with the original passage and metadata such as its title, source, or access permissions.
When a user submits a search query, the application sends the query to the same model. It then compares the query vector with the stored document vectors using inner-product similarity, the metric recommended by NAVER for this model. The highest-scoring passages can be returned as search results or passed to another system for retrieval-augmented generation.
Using the same embedding model and similarity method for both indexing and querying is important. Mixing models can place documents and queries in incompatible vector spaces, while changing similarity methods can alter ranking behavior. These integration choices belong to the application rather than to the model itself.
The 500-token limit and document chunking
Each request can contain up to 500 tokens. Long documents therefore need to be divided before they are embedded. A useful chunk should normally preserve a coherent idea, such as a paragraph, a short group of paragraphs, a section, or a complete passage that can stand on its own in search results.
Arbitrarily cutting text can weaken retrieval because a definition may be separated from its subject or a qualification may be detached from the statement it qualifies. At the same time, very large chunks can make search results less precise by combining several unrelated topics. The supplied research does not prescribe a particular chunk size below the 500-token maximum, so the best choice should be evaluated against the structure of the source documents and the application's retrieval requirements.
Metadata and source text are not part of the vector output itself. The application must retain them separately if users need citations, document titles, filtering, or access-control checks.
Primary use cases
Semantic search and retrieval
clir-emb-dolphin can support search based on meaning instead of exact word overlap. A query about cancelling a reservation, for example, may retrieve a passage that uses different wording but discusses the same policy. This is particularly useful when users use informal language or when documents use terminology that differs from the search query.
Retrieval-augmented generation
In a retrieval-augmented generation workflow, the model can identify relevant passages from a private knowledge base. Those passages may then be supplied to a separate text-generation model. clir-emb-dolphin does not write the final answer, summarize the retrieved material, or perform tool calls; it supplies the vector-based retrieval layer.
Document and passage similarity
Applications can compare vectors to identify related documents, duplicate or near-duplicate content, and thematically similar passages. The model can therefore be used for content discovery, recommendation logic, and document organization when the application defines the comparison and business rules.
Clustering and classification
Embedding vectors can be grouped to discover themes or organized into categories using downstream clustering and classification methods. clir-emb-dolphin provides the input representation, but it does not itself return a category label or execute a complete classification pipeline. The surrounding system must supply the clustering algorithm, classifier, thresholds, and evaluation process.
Strengths and trade-offs
The clearest strength of clir-emb-dolphin is its broad positioning. NAVER describes it as highly generalizable across domains, making it a reasonable starting point when an application needs one embedding model for ordinary search, retrieval, similarity, and organization tasks.
Its 1,024-dimensional output provides a fixed representation that can be indexed by a vector database or another similarity-search system. The explicit recommendation to use inner-product similarity also gives implementers a documented default for comparing vectors.
There are practical trade-offs. The model is limited to 500 input tokens per request, so long-document systems need chunking and additional storage or processing logic. It also returns representations rather than answers. It cannot replace a generative model when the application needs explanations, summaries, dialogue, code, or other generated content.
The supplied research does not publish a model-specific benchmark, release date, knowledge cutoff, or separate public input and output price for clir-emb-dolphin. CLOVA Studio usage is subject to the applicable service pricing and usage terms. Cost should therefore be confirmed in NAVER Cloud Platform's current commercial documentation rather than inferred from the model's vector size.
Modalities and unsupported generative capabilities
clir-emb-dolphin accepts text input and produces numerical embedding output. It does not produce text, images, audio, video, or music. It is not documented as supporting image, audio, or video input, and the supplied research does not identify tool use, function calling, streaming generation, structured JSON output, caching, batching, or fine-tuning features for this endpoint.
Reasoning and coding are not meaningful primary capabilities for this model. It can represent text about reasoning or code for similarity and retrieval purposes, but it does not execute reasoning workflows, run programs, or generate code. Those tasks require a separate model or application component.
Relationship to related NAVER embedding models
NAVER describes clir-emb-dolphin as the broadly applicable option in the classic Embedding API. The related clir-sts-dolphin model is described as being focused more specifically on precise semantic textual similarity. That distinction matters when the main task is comparing sentence meaning rather than building a general retrieval or document-organization system.
NAVER also documents Embedding v2, which uses the open-source bge-m3 model and supports inputs of up to 8,192 tokens. A longer input limit may be more suitable for applications that need to process larger passages with less chunking. However, the supplied research does not provide a direct benchmark or pricing comparison between these models. Indexes should not mix their vectors without a deliberate migration or compatibility strategy.
When to choose clir-emb-dolphin
Choose clir-emb-dolphin when you need a general-purpose text representation for semantic search, retrieval, document similarity, clustering, or downstream classification and your content can be divided into passages of no more than 500 tokens. It is especially appropriate when you want the default model for the classic CLOVA Studio Embedding API and are prepared to use inner-product similarity in your vector-search workflow.
Consider clir-sts-dolphin when precise sentence-level semantic similarity is the central requirement. Consider Embedding v2 when the documented 8,192-token input capacity is more valuable than using clir-emb-dolphin's general-purpose positioning. A generative language model is more appropriate when the system must produce answers, summaries, explanations, or code after retrieval.
In practical terms, clir-emb-dolphin is best viewed as a focused retrieval and text-representation component. Its value comes from how consistently an application chunks content, indexes vectors, applies similarity search, and uses the retrieved material—not from standalone conversation or generation.

