Kinfra

Kinfra-Text-Embedding-4b

by Tencent AI · Current and available through Tencent Cloud TokenHub and the Tencent Cloud embeddings API.

Tencent Cloud's Kinfra-Text-Embedding-4b is a text-only embedding model for multilingual semantic search, retrieval-augmented generation, knowledge bases, clustering, and similarity matching. It supports more than 30 languages, has a documented 32,000-token context length, returns fixed 2,560-dimensional vectors, and costs USD 0.084 per million input tokens internationally.

Embeddings Reasoning Coding
Kinfra-Text-Embedding-4b is Tencent Cloud's larger Kinfra text embedding model for applications that need to compare meaning rather than match exact words. It converts text into numerical vectors that can be indexed in a vector database and searched for semantic similarity. The model supports more than 30 languages, returns fixed 2,560-dimensional vectors, and is priced per million input tokens. Its main trade-off is specialization: it can provide useful representations for retrieval and matching, but it does not generate prose, call tools, or produce images, audio, or video.
Outputs

What Kinfra-Text-Embedding-4b can produce

Embeddings
Inputs

What it can understand

Text
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Kinfra
Model type Embedding
Context window 32K tokens
Status Current and available through Tencent Cloud TokenHub and the Tencent Cloud embeddings API.
Knowledge cutoff notes

A knowledge cutoff is not documented for this embedding model. Embedding models are generally used for vector representation rather than direct factual text generation.

Model notes

The canonical service identifier is kinfra-text-embedding-4b. The model returns fixed 2,560-dimensional floating-point vectors; output dimensions cannot be customized. Tencent Cloud documents support for more than 30 mainstream languages. The embeddings API accepts a single text string or an array, with a documented maximum of 2,000 characters per string and a recommendation of no more than 128 strings per request. The model is available through an OpenAI-compatible embeddings endpoint. Editorial scores are comparative estimates for an embedding model and do not represent vendor benchmarks.

Cost

Model pricing

Input USD 0.084 per million input tokens internationally; RMB 0.6 per million input tokens on the China TokenHub pricing page.
Output Not applicable; the model is billed for text input tokens and returns embedding vectors.
Model guide

Kinfra-Text-Embedding-4b for Multilingual Semantic Search and Retrieval

Tencent's Kinfra-Text-Embedding-4b is a text-only embedding model for converting documents and queries into fixed 2,560-dimensional vectors. With support for more than 30 languages, a documented 32,000-token context length, and access through Tencent Cloud's OpenAI-compatible embeddings API, it is aimed at high-quality semantic search, retrieval-augmented generation, knowledge bases, similarity matching, and classification rather than conversational text generation.

What Kinfra-Text-Embedding-4b does

Kinfra-Text-Embedding-4b is a text embedding model provided through Tencent Cloud. Instead of replying to a prompt with an answer, it transforms text into a numerical representation called an embedding. Texts with related meanings can produce vectors that are close together in a vector database or similarity calculation, even when they use different words.

For example, a knowledge-base application could embed both a customer's question and a collection of support documents. A semantic search system would then retrieve documents whose meaning is relevant to the question, rather than relying only on exact keyword matches. The retrieved passages could subsequently be supplied to a separate generative model, but Kinfra-Text-Embedding-4b itself is not the component that writes the final response.

The model is a text-only member of Tencent's Kinfra embedding offering. Its documented input is text, not images, audio, or video, and its output is a floating-point embedding vector rather than natural-language text.

Where it fits in Tencent Cloud

Tencent Cloud lists Kinfra-Text-Embedding-4b as a current model available through TokenHub and the Tencent Cloud embeddings API. The canonical service identifier is kinfra-text-embedding-4b. The service uses an OpenAI-compatible embeddings request and response structure, which may reduce integration work for applications already organized around that general interface. Compatibility with an API format should not be confused with support for chat completion, tool calling, or text generation; the model remains an embeddings service.

Within the Kinfra family, the 4b model is positioned as a larger text embedding option intended for high semantic quality and deeper text understanding. The supplied documentation does not provide a benchmark table against other Kinfra models, so claims about its relative accuracy should be treated as Tencent's product positioning rather than as a verified independent ranking.

Verified technical specifications

SpecificationDocumented detail
Model typeText embedding model
Input modalityText strings or arrays of text strings
OutputFloating-point embedding vectors
Output dimension2,560 dimensions, fixed
Documented context length32,000 tokens
LanguagesMore than 30 mainstream languages
Per-string request limitA single input string must not exceed 2,000 characters according to the API documentation
Batch guidanceNo more than 128 text strings are recommended in one request

The 32,000-token context figure describes the model's documented context capability, while the API documentation separately specifies a 2,000-character limit for an individual input string. In practice, an application should observe the stricter request-level rule and split long documents into chunks before embedding them. The exact number of tokens represented by 2,000 characters varies with language and text content, so character limits should not be treated as a direct token conversion.

The output dimension cannot be customized according to the supplied model notes. A vector database index, similarity-search implementation, and any downstream machine-learning layer must therefore be configured for exactly 2,560 values per vector.

Languages and supported modalities

Tencent Cloud documents support for more than 30 mainstream languages. Examples listed in the research include Chinese, English, Japanese, Korean, French, German, Russian, Portuguese, and Spanish. This makes the model relevant to multilingual search systems, cross-language knowledge bases, and organizations whose documents and user queries span several major languages. The documentation does not provide a language-by-language quality ranking, so teams should validate performance on their own terminology, domains, and language pairs.

Kinfra-Text-Embedding-4b accepts text and returns embeddings. It does not natively accept image, audio, or video inputs, and it does not produce text, images, audio, video, or music as its primary output. It is consequently not a multimodal understanding model or a generative model. If an application needs to search across images and text, it would need a separate multimodal embedding system or an additional processing pipeline; the supplied research does not identify a particular Tencent alternative for that purpose.

Pricing and API access

Tencent Cloud's international pricing page lists Kinfra-Text-Embedding-4b at USD 0.084 per million input tokens. A China TokenHub pricing page lists it at RMB 0.6 per million input tokens. These are region-specific prices, and the applicable amount can depend on the Tencent Cloud region, account, and service configuration. The model is billed for input tokens; there is no separately priced generated-text output because the response is an embedding vector.

The API accepts one text string or an array of text strings. A typical ingestion workflow divides documents into suitable chunks, sends those chunks in batches while observing the documented request guidance, stores the returned 2,560-dimensional vectors, and later embeds user queries with the same model before running a similarity search.

The OpenAI-compatible request and response format can be useful for portability at the integration layer. It does not guarantee that every SDK feature from another provider is available, however, and the Tencent Cloud documentation remains the authority for request limits, authentication, regional availability, and operational behavior.

Main strengths and trade-offs

  • High-dimensional fixed vectors: The 2,560-dimensional output provides a consistent representation size for indexing and comparison. The trade-off is that storage and index requirements may be greater than for a smaller embedding model.
  • Multilingual coverage: Support for more than 30 languages is useful for international search and document collections. Actual quality can vary by language and subject area, so evaluation with representative data remains important.
  • Long documented context: The 32,000-token context length gives the model a substantial documented context capability, although the 2,000-character per-string API limit still governs individual requests.
  • Usage-based pricing: The listed international rate of USD 0.084 per million input tokens can make large-scale embedding economically attractive, especially when compared with using a general-purpose generative model for retrieval preparation. Cost comparisons with other embedding services require matching their regions, quotas, dimensions, and billing rules.
  • Specialized behavior: Because it is designed for embeddings, it avoids the unnecessary cost and complexity of using a conversational model to produce vectors. The same specialization means it cannot answer questions, write summaries, or serve as the final response generator.

The editorial data rates the model favorably for speed and cost, with comparative scores of 8 for speed and 9 for cost. These are editorial estimates, not Tencent-published benchmarks or guarantees. The supplied research does not include independent latency, recall, throughput, or quality benchmark results.

Best use cases

  • Semantic search: Retrieve relevant documents when the query and document use different wording but express related ideas.
  • Retrieval-augmented generation: Build the retrieval stage that finds passages for a separate language model to use as context.
  • Enterprise knowledge bases: Index policies, manuals, product documentation, and internal records for meaning-based discovery.
  • Multilingual retrieval: Search across Chinese, English, Japanese, Korean, and other supported languages, subject to validation on the organization's data.
  • Similarity and duplicate detection: Compare documents, tickets, questions, or product descriptions using vector distance.
  • Text classification support: Use embeddings as features for a downstream classifier or grouping system.
  • Semantic clustering: Organize large collections by meaning rather than by exact terms.

For a retrieval system, consistent preprocessing matters. The application should use an appropriate chunking strategy, preserve useful metadata such as document identifiers and access permissions, and apply the same embedding model to indexed content and incoming queries. These implementation choices are not model specifications, but they directly affect the usefulness of the resulting search system.

When to choose Kinfra-Text-Embedding-4b

Choose this model when the central requirement is high-quality text representation for multilingual search, retrieval, or similarity workflows and when a fixed 2,560-dimensional index is acceptable. It is particularly relevant if the application already uses Tencent Cloud TokenHub, needs the listed regional pricing, or benefits from an OpenAI-compatible embeddings interface.

A smaller or lower-cost embedding option may be more appropriate when storage, memory, indexing speed, or query latency is more important than the semantic capacity suggested by Tencent's positioning. The supplied research does not identify a named smaller sibling or provide comparative benchmark results, so the choice should be made through testing rather than assumed scores.

A generative language model is more appropriate when the application must answer questions, produce summaries, write code, reason through a task, or generate customer-facing text. Kinfra-Text-Embedding-4b can support such a system's retrieval layer, but it is not a replacement for the generation model. A multimodal model is more appropriate when source data includes images, audio, or video, because this model's documented input is text only.

Limitations to check before deployment

  • The output dimension is fixed at 2,560, so existing indexes designed for another vector size cannot be reused without a compatible migration or separate index.
  • Each input string is documented as limited to 2,000 characters, and the API recommends no more than 128 strings per request. Long-document ingestion therefore requires deliberate chunking and batching.
  • The model does not generate prose and has no documented conversational, reasoning, coding, web-search, tool-calling, or function-calling capability.
  • No knowledge cutoff is documented. This is generally less central for an embedding model than for a factual generation model, but it means the provider has not supplied a cutoff date to use as a data-freshness guarantee.
  • Fine-tuning, batch API availability, caching, and a customizable output dimension are not verified in the supplied research. They should not be assumed when designing the system.
  • Pricing and availability may differ by Tencent Cloud region and account configuration. Confirm the applicable commercial terms before estimating production costs.

Overall, Kinfra-Text-Embedding-4b is best understood as a specialized retrieval component: a multilingual text-to-vector model with a large fixed output and a documented long context, rather than a general-purpose AI assistant. Its value depends on whether the resulting semantic quality and language coverage justify the storage, indexing, and integration requirements of 2,560-dimensional vectors.


Answers to Frequently Asked Questions

What is the Tencent Cloud service identifier for Kinfra-Text-Embedding-4b?
The canonical Tencent Cloud service identifier is kinfra-text-embedding-4b. The model is available through TokenHub and the Tencent Cloud embeddings API, which uses an OpenAI-compatible embeddings request and response structure.
How much does Kinfra-Text-Embedding-4b cost?
Tencent Cloud's international pricing page lists the model at USD 0.084 per million input tokens, while a China TokenHub pricing page lists RMB 0.6 per million input tokens. Actual pricing depends on the Tencent Cloud region, account, and service configuration, so production costs should be confirmed with Tencent Cloud.
Which languages and input types does Kinfra-Text-Embedding-4b support?
Kinfra-Text-Embedding-4b supports more than 30 mainstream languages, including Chinese, English, Japanese, Korean, French, German, Russian, Portuguese, and Spanish. It accepts text only and does not natively process images, audio, or video.
What are the main technical specifications of Kinfra-Text-Embedding-4b?
The model produces fixed 2,560-dimensional floating-point vectors, supports a documented context length of 32,000 tokens, accepts text strings or arrays of text strings, and supports more than 30 mainstream languages. The API documentation limits each input string to 2,000 characters and recommends no more than 128 text strings per request.
What is Kinfra-Text-Embedding-4b used for?
Kinfra-Text-Embedding-4b converts text into numerical embedding vectors for semantic search, retrieval-augmented generation, multilingual knowledge bases, similarity matching, duplicate detection, classification, and clustering. It retrieves semantically related content but does not generate final answers or other natural-language text.


Sources 3
Provider

About Tencent AI