Kinfra

Kinfra-Text-Embedding-0.6b

by Tencent AI · Current and available through Tencent Cloud TokenHub

Tencent Kinfra-Text-Embedding-0.6b is a lightweight multilingual embedding model for high-volume semantic retrieval. It returns fixed 1,024-dimensional vectors, supports a documented 32K-token context, accepts more than 30 languages, and costs $0.07 per million text-input tokens through Tencent Cloud TokenHub. The guide explains its limits, API access, search use cases, and trade-offs against larger or generative models.

Embeddings Reasoning Coding
Kinfra-Text-Embedding-0.6b is Tencent Cloud's approximately 0.6-billion-parameter model for turning text into numerical vectors that represent meaning. Those vectors can be stored in a vector database and compared to find semantically similar documents, questions, or passages. The model is available through Tencent Cloud TokenHub and is aimed at large-scale retrieval workloads where throughput and operating cost matter more than the highest available embedding benchmark score.
Outputs

What Kinfra-Text-Embedding-0.6b can produce

Embeddings
Inputs

What it can understand

Text
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Kinfra
Model type Embedding
Context window 33K tokens
Status Current and available through Tencent Cloud TokenHub
Knowledge cutoff notes

Tencent's public documentation describes the model's embedding specifications and supported use cases but does not provide a model knowledge-cutoff date. Knowledge cutoff is not generally applicable to this embedding endpoint in the same way it is to generative language models.

Model notes

This is an embedding model rather than a generative language model. It returns fixed 1,024-dimensional vectors and does not support configurable output dimensions. Tencent documents support for more than 30 mainstream languages and positions the model for large-scale text retrieval, latency-sensitive workloads, and cost-sensitive workloads. TokenHub uses an OpenAI-compatible embeddings endpoint. Tencent's documentation recommends no more than 128 text inputs per request and states that a single text input should not exceed 2,000 characters. Tencent's published CMTEB results include Mean(Task) 66.64, retrieval 71.01, clustering 68.60, classification 71.46, paired classification 76.84, reranking 64.16, and semantic textual similarity 54.88. Editorial reasoning and coding scores are minimal because the model is not designed for generative reasoning or code generation.

Cost

Model pricing

Input $0.07 per million text-input tokens
Model guide

Tencent Kinfra-Text-Embedding-0.6b for Fast, Cost-Sensitive Vector Search

Tencent Kinfra-Text-Embedding-0.6b is a lightweight multilingual text embedding model that converts text into fixed 1,024-dimensional vectors. With a 32K-token context specification, more than 30 supported languages, OpenAI-compatible TokenHub access, and pricing of $0.07 per million text-input tokens, it is designed for high-volume retrieval, semantic search, clustering, classification, and other latency- or cost-sensitive workloads.

What is Kinfra-Text-Embedding-0.6b?

Kinfra-Text-Embedding-0.6b is a text embedding model provided by Tencent Cloud. Instead of producing a conversational answer, it converts an input text into a fixed-length numerical representation called an embedding. Each output contains 1,024 dimensions. Applications can compare these vectors to estimate semantic similarity, retrieve related content, organize documents into groups, or classify text.

This makes the model a component for search and machine-learning systems rather than a general-purpose chatbot. For example, a knowledge-base application could embed both a user's question and its stored documents, then retrieve documents whose vectors are closest to the question. A retrieval-augmented generation system could use those results as context for a separate language model.

The model identifier in Tencent Cloud TokenHub is kinfra-text-embedding-0.6b. Tencent positions it as a lightweight option for large-scale text recall, high-concurrency services, and applications where lower latency or lower cost is more important than maximum semantic accuracy.

Technical specifications and input limits

The following are the principal specifications documented for the model:

SpecificationDocumented value
ProviderTencent Cloud
Model sizeApproximately 0.6 billion parameters
Model typeText embedding
OutputFixed 1,024-dimensional vector
Context length32,768 tokens
Supported languagesMore than 30 mainstream languages
Custom output dimensionsNot supported
Input modalityText
Output modalityNumerical embeddings

Supported languages include Chinese, English, Japanese, Korean, French, German, Russian, Portuguese, and Spanish, among others. The multilingual coverage makes the model suitable for search collections that contain more than one language, although the supplied documentation does not establish that its quality is identical across every supported language.

Tencent documents a 32,768-token context length, but the TokenHub usage guidance also states that one text input should contain no more than 2,000 characters. These are different kinds of limits: the context specification describes the model's token capacity, while the endpoint guidance is a practical request constraint for submitted text. Implementations should follow the stricter endpoint instruction and split or preprocess long documents before embedding them.

Tencent also recommends no more than 128 text inputs in a single request. The returned vector size is always 1,024 dimensions; applications cannot request a smaller or larger output through a configurable dimensions parameter according to the supplied documentation.

What the model is designed to do

Kinfra-Text-Embedding-0.6b is primarily a retrieval and text representation model. Its vector output can support:

  • Semantic search: Find passages that match the meaning of a query even when they do not use exactly the same words.
  • Large-scale document retrieval: Index substantial collections for knowledge bases, enterprise search, or content discovery.
  • FAQ matching: Compare a new question with previously prepared questions and answers.
  • Clustering: Group documents, support tickets, or user queries by semantic similarity.
  • Classification: Use embeddings as features for categorizing text.
  • Similarity calculation: Measure relationships between documents, queries, or other text records.
  • Retrieval-augmented generation: Supply relevant retrieved text to a separate generative model.

The model does not itself generate a natural-language response. A typical retrieval workflow therefore includes text preprocessing, embedding generation, vector indexing, similarity search, and optionally a separate answer-generation step. This separation lets an application use Kinfra-Text-Embedding-0.6b for inexpensive retrieval while reserving a larger language model for final responses.

TokenHub access and pricing

Tencent exposes the model through TokenHub's OpenAI-compatible embeddings endpoint at https://tokenhub.tencentmaas.com/v1/embeddings. A request supplies the model identifier and either one text string or an array of text strings. The compatibility layer can simplify integration for applications that already use the common embeddings request pattern, but it does not change the model's output type or its documented input limits.

Tencent's published pricing lists Kinfra-Text-Embedding-0.6b at $0.07 per million text-input tokens. The supplied pricing information specifies billing for text input and does not list a separate output-token charge for the returned vectors. The practical cost of a complete search system can still include storage, vector indexing, network traffic, and any separate model used to generate answers.

At this price, the model is positioned for frequent embedding calls and large collections. Teams should nevertheless estimate token volume from their own documents and query traffic. Re-embedding unchanged content unnecessarily, sending oversized text chunks, or using a more expensive model for every routine query can increase costs without improving the search experience proportionally.

Strengths and trade-offs

The central strength of Kinfra-Text-Embedding-0.6b is its balance between a relatively small model size, broad language coverage, fixed vector output, and low published input pricing. A 1,024-dimensional output is large enough for many semantic-search systems while remaining predictable for index design and storage planning. The fixed size also makes deployment straightforward because every record has the same vector shape.

Its approximately 0.6-billion-parameter scale is relevant to throughput and latency-sensitive deployments. Tencent specifically presents the model for high-concurrency and cost-sensitive workloads. These are provider positioning claims rather than a guarantee of a particular response time: actual performance depends on service load, request size, batching, network conditions, and the surrounding vector database.

The main trade-off is embedding quality. In Tencent's published CMTEB comparison, Kinfra-Text-Embedding-0.6b records a Mean(Task) score of 66.64 and a retrieval score of 71.01. The larger Kinfra-Text-Embedding-4b scores higher on those measures, according to the supplied research. This supports a straightforward positioning distinction: the 0.6b model favors efficiency, while the larger sibling may be more appropriate when retrieval quality is more important than cost or latency.

The model also lacks configurable output dimensions. Systems designed around a different vector width must either adapt their index or choose another embedding model. Its documented output is an embedding rather than text, an image, audio, or video, so it cannot replace a generative, vision, speech, or multimodal model.

Reasoning, coding, and tool capabilities

Kinfra-Text-Embedding-0.6b is not intended for reasoning dialogue, code generation, or autonomous tool use. It accepts text and returns vectors; it does not return a natural-language completion, support image or audio understanding, or provide a documented function-calling workflow. It therefore has no useful standalone maximum output-token setting for a generated answer.

Editorial capability assessments rate its reasoning and coding suitability as minimal because those tasks require generative output that this model does not provide. Those are editorial evaluations, not provider-published benchmark scores. The model can still be used around reasoning or coding systems—for example, to retrieve relevant programming documentation or prior support cases—but another model must perform the actual explanation, code generation, or decision-making.

When to choose this model

Choose Kinfra-Text-Embedding-0.6b when the application needs a large number of text embeddings and the following characteristics are useful:

  • Low published input cost is important.
  • High request volume or latency-sensitive retrieval favors a lightweight model.
  • A fixed 1,024-dimensional vector fits the planned vector index.
  • The data includes several of the more than 30 documented mainstream languages.
  • The workload involves semantic search, FAQ matching, clustering, classification, or knowledge-base retrieval rather than answer generation.
  • The application can follow the 2,000-character single-input guidance and the recommendation of no more than 128 inputs per request.

It may be a particularly sensible starting point for a high-volume search system that can measure retrieval quality with its own queries. Testing on representative data is important because a model that is inexpensive and fast may still be unsuitable if missed results have a high business cost.

When another option may be more appropriate

Consider Tencent's larger Kinfra-Text-Embedding-4b when the supplied CMTEB comparison and the application's own evaluation indicate that higher retrieval quality justifies additional cost, latency, or resource use. The 0.6b model is not automatically the best choice for every search task simply because its input price is low.

A generative language model is more appropriate when the requirement is to write answers, summarize retrieved passages, explain code, or conduct a conversation. A vision, audio, or multimodal model is needed when the searchable content includes those modalities rather than text alone. Another embedding service may be preferable if the system requires a different vector dimension, a capability not documented for this model, or an input and output format beyond TokenHub's text-embedding endpoint.

Bottom line

Kinfra-Text-Embedding-0.6b is a focused Tencent Cloud embedding model for turning multilingual text into fixed 1,024-dimensional vectors. Its strongest case is high-volume semantic retrieval where predictable vector size, low published input pricing, and efficiency matter. Its limitations are equally clear: it does not generate answers, its output dimensions cannot be customized, endpoint guidance limits individual inputs to 2,000 characters and batches to 128 text inputs, and a larger model may deliver better retrieval quality. For teams that understand those boundaries, it is a practical lightweight component for search, recommendation, classification, and retrieval-augmented applications.


Answers to Frequently Asked Questions

When should a team choose Kinfra-Text-Embedding-0.6b instead of a larger embedding model?
Choose Kinfra-Text-Embedding-0.6b when low input cost, high throughput, lower latency, multilingual text support, and a fixed 1,024-dimensional vector are priorities. A larger model such as Kinfra-Text-Embedding-4b may be more suitable when higher retrieval quality is worth additional cost, latency, or resource use.
How much does Kinfra-Text-Embedding-0.6b cost through TokenHub?
Tencent's published pricing lists Kinfra-Text-Embedding-0.6b at $0.07 per million text-input tokens. The supplied pricing information does not list a separate output-token charge, although complete deployments may also incur costs for storage, vector indexing, network traffic, and answer-generation models.
What is Tencent Kinfra-Text-Embedding-0.6b used for?
Kinfra-Text-Embedding-0.6b converts text into fixed 1,024-dimensional numerical vectors for semantic search, document retrieval, FAQ matching, clustering, classification, similarity calculation, and retrieval-augmented generation. It does not generate natural-language answers itself.
What are the input limits and output dimensions of Kinfra-Text-Embedding-0.6b?
The model produces a fixed 1,024-dimensional vector, and custom output dimensions are not supported. Tencent's TokenHub guidance recommends limiting each text input to 2,000 characters and sending no more than 128 text inputs in one request. The documented model context length is 32,768 tokens.


Sources 4
Provider

About Tencent AI