Qwen3.7

qwen3.7-text-embedding

by Qwen · Current and available through Alibaba Cloud Model Studio internationally

A multilingual embedding model for semantic search, RAG, recommendation, clustering, classification, and code retrieval, with 128K-token input support, 201-language coverage, configurable dimensions up to 2,560, and optional dense or sparse vector output.

Embeddings
Qwen3.7-Text-Embedding is a text-only embedding model from Alibaba Cloud Model Studio. Instead of generating an answer, it converts text into numerical vectors that applications can compare for semantic similarity. Its main differentiators are broad multilingual coverage, a 128K-token input limit, selectable vector sizes, and support for dense, sparse, or combined retrieval representations.
Outputs

What qwen3.7-text-embedding can produce

Embeddings
Inputs

What it can understand

Text
Capabilities

Supported features

Batch API
Model profile

Performance characteristics

8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Qwen3.7
Model type Embedding
Context window 131K tokens
Release date 2026-07-31
Status Current and available through Alibaba Cloud Model Studio internationally
Knowledge cutoff notes

Alibaba Cloud does not publish a separate training-data or knowledge-cutoff date for this embedding model. The model's current API documentation specifies input limits and embedding behavior but does not define a knowledge cutoff.

Model notes

A multilingual text embedding model based on Qwen3.7. Alibaba Cloud reports approximately 20% better performance than text-embedding-v4 on MTEB multilingual, Chinese-English, and code retrieval tasks. It supports 201 languages and dialects, configurable dimensions of 256, 512, 768, 1,024, 1,536, 2,048, or 2,560, with 1,024 as the default. The API supports task-specific instructions, query/document text types, and dense, sparse, or combined dense-and-sparse output. The maximum input length is 128,000 tokens, and synchronous requests support up to 20 input rows. Batch inference is supported. The official model page does not publish a knowledge cutoff or a separate output-token limit. Output is an embedding vector rather than generated natural-language text, so reasoning and coding scores are not applicable.

Cost

Model pricing

Input $0.07 per 1 million input tokens in Singapore international deployment; China pricing may differ by region and billing mode
Model guide

Qwen3.7-Text-Embedding: Multilingual Search Vectors with Configurable Dimensions

Qwen3.7-Text-Embedding is Alibaba Cloud Model Studio's multilingual text embedding model for semantic search, recommendation, clustering, classification, and retrieval-augmented generation. It supports 201 languages and dialects, inputs up to 128,000 tokens, configurable embedding dimensions from 256 to 2,560, task-specific instructions, and dense or sparse vector output.

What is Qwen3.7-Text-Embedding?

Qwen3.7-Text-Embedding is a multilingual text embedding model provided through Alibaba Cloud Model Studio. An embedding model transforms text into a list of numbers, called a vector, so software can compare the meaning of documents, search queries, product descriptions, support tickets, or code without relying only on exact keyword matches.

For example, a search for “how to reset a forgotten password” can be matched with a document titled “account credential recovery” even though the wording differs. This makes embeddings useful as a foundation for semantic search, retrieval-augmented generation (RAG), recommendations, clustering, and classification.

Qwen3.7-Text-Embedding is based on Qwen3.7 and is positioned in Alibaba Cloud's current embedding catalog for demanding multilingual text and code retrieval workloads. Alibaba Cloud describes it as a successor-level option to text-embedding-v4. That positioning is a provider claim; the practical choice still depends on retrieval quality, latency, vector-storage cost, and the languages in an application's data.

Key specifications

SpecificationVerified detail
ProviderAlibaba Cloud Model Studio
Model IDqwen3.7-text-embedding
Model typeText embedding
Release dateJuly 31, 2026
Input languages201 languages and dialects
Maximum input length128,000 tokens
Context window listed by the catalog131,072 tokens
Embedding dimensions256, 512, 768, 1,024, 1,536, 2,048, or 2,560
Default dimension1,024
Output typesDense, sparse, or combined dense-and-sparse embeddings
Synchronous request limitUp to 20 input rows
Batch inferenceSupported
International price$0.07 per 1 million input tokens in Singapore international deployment

The documentation does not publish a separate generated-output limit because this model returns vectors rather than natural-language text. It also does not publish a knowledge-cutoff date.

How the model works in practice

An application sends text to the embedding endpoint and receives a numerical representation for each input. The application can then store those vectors in a vector database or another retrieval system. At query time, it embeds the user's search text and compares that vector with stored document vectors.

Qwen3.7-Text-Embedding supports task-specific instructions and distinguishes between query and document text types. This distinction is useful in retrieval systems because a search query and a document play different roles. Applying the intended instruction and text type can help an application produce vectors that are better aligned with its retrieval task, although the supplied documentation does not provide a universal accuracy guarantee.

The model can return dense vectors, sparse vectors, or both. Dense vectors represent meaning through a fixed-length numerical array and are commonly used for semantic similarity. Sparse vectors preserve a more selective representation in which many values are zero or absent. A combined dense-and-sparse setup can support hybrid retrieval, pairing semantic matching with stronger lexical or term-level matching.

Multilingual coverage and input capacity

Alibaba Cloud documents support for 201 languages and dialects, making the model suitable for search systems whose content or users span multiple language communities. This can be useful for multilingual knowledge bases, cross-language discovery, international product catalogs, and support content.

The maximum input length is 128,000 tokens, while the catalog lists a 131,072-token context window. That is substantially more room than a typical short-text embedding workflow requires. It allows applications to process long documents or larger text segments, but a larger limit does not automatically mean that embedding an entire long document is the best retrieval design. Splitting content into meaningful sections can make search results more precise and can reduce unnecessary token usage.

For synchronous requests, the documented limit is up to 20 input rows. Larger workloads can use batch inference, which is more appropriate for indexing a large document collection, product database, or code repository.

Choosing an embedding dimension

The model supports seven selectable dimensions: 256, 512, 768, 1,024, 1,536, 2,048, and 2,560. The default is 1,024. Dimension selection affects the size of every stored vector and therefore influences memory, storage, transfer, and similarity-search costs.

  • Smaller dimensions: useful when storage efficiency, lower transfer overhead, or high search throughput matters most.
  • Medium dimensions: a practical starting point for many general semantic-search and RAG systems.
  • Larger dimensions: potentially useful when an evaluation shows that preserving more representation capacity improves the target retrieval task.

These are engineering trade-offs rather than guarantees that the largest vector is always best. A team should evaluate candidate dimensions on its own documents, languages, queries, and relevance judgments. Changing dimensions after indexing generally requires regenerating the stored vectors so that queries and documents use the same representation.

Main strengths and limitations

Strengths

  • Broad multilingual scope: support for 201 languages and dialects is relevant to international search and cross-language retrieval.
  • Flexible vector output: dense, sparse, and combined output supports different retrieval architectures, including hybrid search.
  • Configurable vector size: seven dimensions let teams balance retrieval experiments against storage and serving costs.
  • Long input support: the 128,000-token maximum can accommodate long documents and large text segments when that design is appropriate.
  • Retrieval-oriented controls: task instructions and query/document text types are designed for embedding workflows rather than conversational generation.
  • Large-scale processing: batch inference is available for indexing and other offline workloads.

Limitations

  • It does not generate answers: the output is an embedding vector, not a natural-language response. A separate generation model is needed for an RAG system that explains retrieved results.
  • No multimodal input: the documented model accepts text and does not provide image, audio, or video understanding.
  • No tools or function calling: tool use, web search, and action execution are not supported model capabilities.
  • No published reasoning or coding score: reasoning and coding scores are not applicable to this embedding model. It can support code retrieval, but it is not a code-generation model.
  • Pricing varies by deployment: the quoted $0.07 per 1 million input tokens applies to Singapore international deployment; China pricing can differ by region and billing mode.
  • Quality still requires testing: broad language coverage and a large context limit do not establish equal retrieval quality for every language, domain, or query type.

Pricing, speed, and cost trade-offs

The supplied international price is $0.07 per 1 million input tokens for Singapore international deployment. There is no output-token charge listed because the model produces embeddings rather than generated text. China-region prices and billing modes may differ, so production estimates should use the price for the selected deployment.

The catalog gives the model a speed score of 8 and a cost score of 8 in the supplied evaluation data. These are editorial or catalog scores, not provider-published benchmark results, and should not be treated as a guaranteed latency or quality measurement. Actual performance depends on request size, concurrency, batch design, region, network conditions, vector-database configuration, and the selected dimension.

For online search, a smaller dimension may reduce vector storage and similarity-search work. For offline indexing, batch inference can improve operational efficiency. The sensible approach is to measure end-to-end retrieval quality and latency with representative workloads rather than selecting a dimension or deployment solely from a general score.

Best use cases

  • Multilingual semantic search: finding conceptually related content across languages and dialects.
  • Retrieval-augmented generation: retrieving relevant passages before passing them to a separate text-generation model.
  • Code retrieval: locating related functions, documentation, examples, or issue discussions in a code repository.
  • Recommendation: matching products, articles, listings, or support resources by semantic similarity.
  • Clustering: grouping documents, tickets, reviews, or messages by meaning.
  • Classification: using vector representations as features for assigning content to categories.
  • Large-scale indexing: embedding document collections with batch inference and selecting a dimension appropriate for storage constraints.

When to choose this model

Choose Qwen3.7-Text-Embedding when multilingual coverage, long inputs, configurable vector sizes, or dense-and-sparse retrieval are important requirements. It is especially suitable when the application needs a dedicated embedding model rather than a chat model and when the team can evaluate retrieval quality on its own data.

A smaller or simpler embedding option may be more appropriate when the dataset is small, language requirements are narrow, or minimal operational complexity matters more than flexible output modes. A model designed for multimodal embeddings is a better fit for applications that must search images, audio, or video alongside text. A chat or reasoning model is necessary when the application must write answers, analyze content interactively, call tools, or generate code rather than only retrieve related items.

For RAG, Qwen3.7-Text-Embedding should be viewed as the retrieval component, not the complete application. The system still needs document preparation, chunking, vector storage, ranking or filtering, and a separate model if it must produce a final answer.

Bottom line

Qwen3.7-Text-Embedding is a retrieval-focused model with a strong combination of multilingual coverage, long-input support, configurable dimensions, and dense or sparse output. Its main value is flexibility: teams can adapt the vector representation and retrieval style to their storage, latency, and search requirements. Its boundaries are equally clear: it is text-only, does not generate natural-language responses, and requires task-specific evaluation before production deployment.


Answers to Frequently Asked Questions

How much does Qwen3.7-Text-Embedding cost?
The listed international price is $0.07 per 1 million input tokens for the Singapore international deployment. Pricing for China-region deployments and different billing modes may vary, so production estimates should use the price of the selected deployment.
Can Qwen3.7-Text-Embedding produce dense and sparse vectors?
Yes. Qwen3.7-Text-Embedding can return dense vectors, sparse vectors, or combined dense-and-sparse embeddings. Combined output can support hybrid retrieval that uses both semantic similarity and lexical or term-level matching.
What embedding dimensions are available in Qwen3.7-Text-Embedding?
The model supports dimensions of 256, 512, 768, 1,024, 1,536, 2,048, and 2,560. The default is 1,024. Smaller dimensions reduce storage and search costs, while larger dimensions may preserve more representation capacity, so teams should evaluate the options on their own data.
Which languages and input lengths does Qwen3.7-Text-Embedding support?
The model supports 201 languages and dialects, with a maximum input length of 128,000 tokens. Although it can process long inputs, splitting documents into meaningful sections may improve retrieval precision and reduce token usage.
What is Qwen3.7-Text-Embedding used for?
Qwen3.7-Text-Embedding converts text into numerical vectors for semantic search, retrieval-augmented generation, recommendations, clustering, classification, multilingual discovery, and code retrieval. It serves as the retrieval component and does not generate natural-language answers.


Sources 5
Provider

About Qwen