Dolphin

clir-sts-dolphin

by NAVER AI · Current and available through the CLOVA Studio Embedding API

clir-sts-dolphin is Naver Cloud’s CLOVA Studio model for sentence-level semantic similarity. It returns 1,024-dimensional vectors, accepts up to 500 tokens per request, and is designed for cosine-based comparison in semantic search, clustering, document relatedness, and classification workflows.

Embeddings Reasoning Coding
clir-sts-dolphin is a sentence-embedding model available through Naver Cloud’s CLOVA Studio Embedding API. Rather than generating answers or other media, it converts a sentence or short text into a numerical vector that represents its meaning. Applications can compare those vectors with cosine similarity to estimate how closely two pieces of text are related.
Outputs

What clir-sts-dolphin can produce

Embeddings
Inputs

What it can understand

Text
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
Specifications

Technical details

Model family Dolphin
Model type Embedding
Context window 500 tokens
Status Current and available through the CLOVA Studio Embedding API
Knowledge cutoff notes

No model-specific knowledge cutoff is published in the reviewed Naver Cloud documentation. Embedding models generally produce vector representations rather than conversational knowledge responses, but no cutoff should be inferred.

Model notes

clir-sts-dolphin is a sentence-embedding model specialized for precise semantic comparison. It returns a 1,024-dimensional vector and is intended to be evaluated with cosine similarity. The maximum input length is 500 tokens per request. Naver Cloud distinguishes it from clir-emb-dolphin, which is positioned as a more general-purpose embedding model and uses inner product as its recommended distance metric. The model is available through the CLOVA Studio Embedding API and is listed in the Naver LangChain integration. Public first-party documentation reviewed for this record does not specify a release date, knowledge cutoff, exact pricing, maximum generated-output value, or caching support.

Model guide

clir-sts-dolphin: Naver’s Sentence Similarity Embedding Model

clir-sts-dolphin is a Naver Cloud CLOVA Studio embedding model designed specifically for measuring the semantic similarity of sentences. It converts text into 1,024-dimensional vectors, supports inputs of up to 500 tokens per request, and is intended for sentence comparison, semantic search, clustering, document relatedness, and classification workflows.

What is clir-sts-dolphin?

clir-sts-dolphin is a sentence-embedding model from Naver Cloud’s CLOVA Studio catalog. Its purpose is to represent sentences as numerical vectors so that software can compare their meaning. The model is designed for semantic textual similarity, often abbreviated as STS, rather than for conversation or text generation.

When an application sends text to the CLOVA Studio Embedding API, clir-sts-dolphin returns a vector containing 1,024 floating-point values. The vector is not intended to be read directly by a person. Instead, an application compares it with another vector to determine whether the corresponding sentences are semantically close.

For example, a search system might compare the query how can I reset my password? with a support article titled Steps for recovering account access. Although the wording differs, an embedding model can represent the two texts in a way that helps the system recognize their related meaning.

Where it fits in CLOVA Studio

clir-sts-dolphin is part of Naver Cloud’s CLOVA Studio embedding tooling and is available through the CLOVA Studio Embedding API. The provider positions it as a model specialized for precise sentence-meaning comparison.

Naver Cloud distinguishes it from clir-emb-dolphin, a more general-purpose embedding model. The documented distinction is important when selecting a similarity workflow: clir-sts-dolphin is intended for sentence similarity and uses cosine similarity as its recommended comparison metric, while clir-emb-dolphin is described separately and uses inner product as its recommended distance metric.

The model is also listed in Naver’s CLOVA Studio LangChain integration. That can make it practical for applications that already use LangChain-compatible embedding components, although the model still remains an embedding service rather than a generative language model.

Verified technical specifications

SpecificationVerified detail
Modelclir-sts-dolphin
ProviderNaver Cloud
Model typeSentence embedding
Vector size1,024 dimensions
Maximum input length500 tokens per request
Recommended similarity metricCosine similarity
Access methodCLOVA Studio Embedding API
Text inputSupported
Embedding outputSupported
Natural-language text outputNot supported as a model output
Image, audio, and video input or outputNot supported

The 500-token limit applies to the input handled in a single Embedding API request. The reviewed documentation does not publish a model-specific knowledge cutoff, generated-output limit, or output-token limit. Those omissions are expected for an embedding model because it returns vectors rather than generated passages.

How semantic similarity works

Each input sentence is transformed into a point in a 1,024-dimensional mathematical space. Sentences with related meanings should generally produce vectors that are closer according to the selected distance calculation. With clir-sts-dolphin, Naver Cloud recommends cosine similarity.

Cosine similarity compares the angle between two vectors rather than simply comparing their raw length. In practical terms, an application can use the resulting score to rank candidate sentences, documents, or records from most related to least related. The exact threshold for calling two texts similar depends on the application and should be evaluated with representative data rather than assumed in advance.

The model does not decide by itself that two sentences are duplicates, relevant, or interchangeable. It supplies vector representations; the surrounding application must define ranking rules, thresholds, filtering logic, and any user-facing result.

Best use cases

clir-sts-dolphin is most useful when the main task is comparing the meaning of relatively short pieces of text. Suitable applications include:

  • Sentence similarity: Compare paraphrases, support questions, titles, product descriptions, or other short text units.
  • Semantic search: Match a user query with documents or passages based on meaning rather than exact keyword overlap.
  • Document relatedness: Estimate how closely two articles, records, or content items are connected.
  • Clustering: Group sentences or short documents into themes based on their vector representations.
  • Classification features: Use embeddings as input features for a separate classifier, such as a topic or sentiment classification system.
  • Duplicate and near-duplicate analysis: Identify text that expresses similar information even when the wording is not identical.

For semantic search, a typical workflow is to embed documents in advance, store their vectors in a suitable vector index, embed the incoming query, and rank the stored vectors using cosine similarity. The model itself does not provide the search index, ranking interface, or final answer generation; those functions belong to the surrounding application.

Input length and document chunking

The maximum input length is 500 tokens per request. This is a practical limit for sentence-level and short-text work, but it means that a long document should not be sent as one unprocessed input.

Longer material should be divided into meaningful sections before embedding. Paragraphs, headings, individual support entries, or other semantic units are usually more useful chunk boundaries than arbitrary cuts in the middle of a sentence. After retrieval, an application can use the most relevant chunks for display or pass them to a separate system for further processing.

Chunking also affects search quality. Very small chunks may lose necessary context, while very large chunks may mix several unrelated topics and make similarity results less precise. The 500-token limit therefore needs to be treated as part of the application design, not merely as an API validation rule.

Strengths and trade-offs

The clearest strength of clir-sts-dolphin is specialization. It is not presented as a general-purpose model that happens to support embeddings; it is specifically intended for sentence-level semantic comparison. That focus can make it a suitable choice when the central requirement is comparing the meaning of short texts with a consistent vector representation.

Its 1,024-dimensional output provides a fixed-size representation that can be stored and indexed by downstream retrieval systems. The documented cosine-similarity recommendation also gives developers a clear starting point for evaluating relatedness.

There are corresponding trade-offs. The model does not generate explanations, summaries, classifications, or conversational answers by itself. It also does not process images, audio, or video according to the supplied model record. A system that needs a natural-language response will need a separate generative model or application component after retrieval or classification.

Its 500-token input limit makes it less convenient for embedding long documents without preprocessing. For large files, a model or service designed around longer inputs may be easier to operate, although the supplied documentation does not identify a specific alternative or provide a direct quality comparison.

clir-sts-dolphin compared with clir-emb-dolphin

The most relevant documented comparison is with Naver Cloud’s clir-emb-dolphin. Both belong to the embedding category, but Naver Cloud gives them different positioning.

  • clir-sts-dolphin: Specialized for sentence similarity and precise comparison of sentence meaning; cosine similarity is recommended.
  • clir-emb-dolphin: Positioned as a more general-purpose embedding model; inner product is recommended as its distance metric.

This distinction does not establish that one model is universally better. It indicates that the similarity objective and the distance metric should match the model selected. Teams comparing the two should use their own representative queries, documents, languages, and retrieval targets rather than transferring thresholds or indexing assumptions from one model to the other.

Capabilities it does and does not provide

clir-sts-dolphin accepts text and produces embeddings. It is not a reasoning or coding assistant in the conventional sense: the reviewed record does not identify native reasoning workflows, code generation, tool calling, function calling, streaming generation, structured text output, or conversational response generation.

It also has no documented image, audio, or video modality. The output is a numerical embedding rather than text, an image, audio, or video. As a result, it should be evaluated as an infrastructure component for retrieval and text comparison, not as a standalone chatbot or multimodal assistant.

The model record does not specify fine-tuning, caching support, batch API availability, release date, or public per-token pricing for this exact model. These values should not be inferred from other CLOVA Studio services. Applications should confirm current commercial terms and operational options in their Naver Cloud account or the latest provider documentation before deployment.

Pricing and availability

clir-sts-dolphin is currently documented as available through the CLOVA Studio Embedding API. The supplied first-party references do not publish a verified price for this exact model, so no numeric price is stated here.

Because embedding costs can depend on request volume, account configuration, region, or the applicable CLOVA Studio commercial plan, prospective users should check Naver Cloud’s current pricing and service terms directly. The absence of a price in the reviewed model documentation should not be interpreted as free access.

When to choose clir-sts-dolphin

Choose clir-sts-dolphin when the primary task is sentence-level semantic similarity and you want a Naver Cloud embedding model explicitly positioned for that purpose. It is a sensible candidate for semantic search, related-content discovery, clustering, and classification pipelines in which inputs can be kept within 500 tokens and cosine similarity is appropriate.

Consider another option when the application needs generated answers, summaries, code, tool use, or multimodal processing. A different embedding model may also be more appropriate if the workflow is designed around a different similarity metric, requires substantially longer inputs, or has a documented need for capabilities not provided by this model.

In short, clir-sts-dolphin is best understood as a focused text-representation component. Its value comes from turning short sentences into comparable vectors, while search, ranking, classification, generation, and user interaction remain responsibilities of the larger application.


Answers to Frequently Asked Questions

What should developers consider when using clir-sts-dolphin for long documents?
Because the model supports a maximum of 500 tokens per request, long documents should be divided into meaningful sections such as paragraphs, headings, or support entries before embedding. The resulting chunks can then be indexed and ranked using cosine similarity.
Can clir-sts-dolphin generate text or process images, audio, and video?
No. clir-sts-dolphin produces numerical text embeddings rather than natural-language responses. It has no documented support for image, audio, or video input or output, so applications needing generation or multimodal processing require separate components.
How does clir-sts-dolphin differ from clir-emb-dolphin?
clir-sts-dolphin is specialized for sentence similarity and precise comparison of sentence meaning, with cosine similarity as the recommended metric. clir-emb-dolphin is positioned as a more general-purpose embedding model and uses inner product as its recommended distance metric.
What is clir-sts-dolphin used for?
clir-sts-dolphin is a Naver Cloud sentence-embedding model designed to compare the meaning of short text. It can support semantic search, sentence similarity, document relatedness, clustering, classification features, and duplicate or near-duplicate detection.
What are the main technical specifications of clir-sts-dolphin?
clir-sts-dolphin accepts text through the CLOVA Studio Embedding API and returns 1,024-dimensional vectors. Its maximum input length is 500 tokens per request, and Naver Cloud recommends cosine similarity for comparing the embeddings.


Sources 3
Provider

About NAVER AI