Qwen-Turbo

Qwen-Turbo

by Qwen · Currently accessible; no longer being updated; Alibaba Cloud recommends switching to Qwen-Flash

Qwen-Turbo is Alibaba Cloud’s fast, cost-optimized text-generation model for simple to moderate workloads. It supports thinking and non-thinking modes, structured JSON output, a 131,072-token context window, streaming, context caching, and batch inference. It remains accessible but will no longer be updated, so Alibaba Cloud recommends Qwen-Flash for new deployments.

Text Reasoning Coding
Qwen-Turbo is a verified Alibaba Cloud Model Studio model built for high-throughput, cost-sensitive text applications. It is suitable for customer support, summarization, rewriting, classification, structured extraction, and straightforward question answering. Its low listed price and fast response profile are balanced by limited advanced capabilities and an announced end to future model updates.
Outputs

What Qwen-Turbo can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Web search Streaming JSON mode Structured output Prompt caching Batch API
Model profile

Performance characteristics

6/10 Reasoning
5/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Qwen-Turbo
Model type Lightweight
Context window 131K tokens
Maximum output 16K tokens
Release date 2025-04-28
Status Currently accessible; no longer being updated; Alibaba Cloud recommends switching to Qwen-Flash
Knowledge cutoff notes

No exact knowledge-cutoff date was identified in the reviewed first-party model documentation.

Model notes

The canonical model ID is qwen-turbo. Alibaba Cloud states that the model is functionally equivalent to qwen-turbo-2025-04-28 and that Qwen-Turbo will no longer be updated. The provider recommends Qwen-Flash as the successor for new deployments. The model supports hybrid thinking and non-thinking modes, but structured output is documented for the non-thinking mode. International batch inference is available at a 50% discount. Pricing varies by deployment region and thinking mode. The official model page reports function calling and web search as unsupported, while separate Alibaba Cloud documentation lists qwen-turbo as compatible with the first-party web-search feature; this may reflect differences between the base model capability table and platform-level grounding support.

Cost

Model pricing

Input $0.05 per 1M tokens international non-thinking; $0.05 per 1M tokens international thinking; $0.044 per 1M tokens China (Beijing)
Output $0.20 per 1M tokens international non-thinking; $0.50 per 1M tokens international thinking; $0.087 per 1M tokens China non-thinking; $0.431 per 1M tokens China thinking
Model guide

Qwen-Turbo: Alibaba Cloud’s Fast, Low-Cost Model for Routine Text Workloads

Qwen-Turbo is Alibaba Cloud Model Studio’s fast, cost-optimized text-generation model for simple to moderate workloads. It supports thinking and non-thinking modes, structured JSON output, context caching, streaming, batch inference, and a 131,072-token context window. The model remains accessible through the qwen-turbo identifier, but Alibaba Cloud says it will no longer receive updates and recommends Qwen-Flash for new deployments.

What is Qwen-Turbo?

Qwen-Turbo is Alibaba Cloud Model Studio’s lightweight Qwen text-generation model for simple to moderate workloads. Its canonical model identifier is qwen-turbo, and Alibaba Cloud describes it as functionally equivalent to the qwen-turbo-2025-04-28 snapshot.

The model accepts text and returns text. It is not a multimodal model in the documented configuration: the reviewed specifications do not list image, audio, or video input, and it does not produce non-text media. Its role is therefore narrower than that of models designed to interpret images or operate as broad multimodal assistants.

Qwen-Turbo is currently accessible through Model Studio, but it is not a strong choice for a new long-lived deployment without considering its lifecycle status. Alibaba Cloud’s current documentation says the model will no longer be updated and recommends migrating to Qwen-Flash for new deployments. No exact shutdown date is published in the reviewed first-party documentation.

Capabilities and model limits

Qwen-Turbo has a documented context window of 131,072 tokens. A context window is the total amount of text the model can consider across the request and conversation. Within that overall window, the documented maximum input length is 98,304 tokens and the maximum output length is 16,384 tokens.

These limits make the model suitable for long documents, large support histories, and sizeable batches of text, although applications still need to manage prompts carefully. A large context window does not by itself mean that every long prompt will receive equally detailed or reliable reasoning.

The model supports two operating styles:

  • Non-thinking mode: the lower-cost, faster option for routine generation, extraction, summarization, and question answering.
  • Thinking mode: an option that allocates additional reasoning effort and has a higher output price.

Alibaba Cloud documents structured output in non-thinking mode, including JSON Object output. This is useful when an application needs fields such as a category, sentiment label, entity list, or extracted values rather than free-form prose. Structured output should not be treated as evidence that every mode or deployment supports identical JSON behavior; the documentation identifies the feature as mode- and region-sensitive.

Qwen-Turbo also supports streaming, which delivers generated text incrementally instead of waiting for the complete response. It supports context caching and batch inference in the international Model Studio deployment. Batch inference can reduce cost for offline workloads, while streaming is more useful when an application needs to show partial output quickly.

Pricing and cost trade-offs

For the international deployment, Alibaba Cloud lists non-thinking input at $0.05 per 1 million tokens and non-thinking output at $0.20 per 1 million tokens. Thinking-mode output is listed at $0.50 per 1 million tokens. The supplied pricing information also lists international thinking-mode input at $0.05 per 1 million tokens.

China pricing differs. The reviewed information lists input at $0.044 per 1 million tokens, non-thinking output at $0.087 per 1 million tokens, and thinking output at $0.431 per 1 million tokens for the China deployment. Regional availability, pricing, and capabilities should therefore be checked before estimating production costs.

International batch inference receives a 50% discount where supported. Cached input is priced separately, so applications that repeatedly send the same context should consult the applicable Model Studio pricing rules rather than assume that all input tokens are charged at the standard rate.

These prices position Qwen-Turbo toward high-volume text processing rather than premium, deep-reasoning workloads. Non-thinking mode is the cost-focused default. Thinking mode can be reserved for requests that benefit from additional reasoning, but its output price is substantially higher than non-thinking output.

Reasoning, coding, and tool support

Qwen-Turbo’s thinking mode provides a documented way to request additional reasoning effort, but the supplied documentation does not establish that it matches the capabilities of a leading advanced reasoning model. It is best understood as a choice between faster routine generation and more deliberate processing within the same model family.

For coding, Qwen-Turbo can be used for ordinary text-based programming assistance, such as generating small snippets, explaining code, or rewriting simple technical text. However, its documented positioning is focused on simple to moderate tasks, and the supplied evaluation rates its coding suitability below its speed and cost characteristics. It should not be selected primarily for complex software engineering, extensive repository work, or difficult debugging without task-specific testing.

Function calling is not listed as supported for this exact model. That limits its suitability for agentic applications that need the model to select and invoke application-defined tools. The official model capability information also lists web search as unsupported, although separate Alibaba Cloud documentation lists qwen-turbo as compatible with the platform’s first-party web-search feature. This apparent difference may reflect a distinction between the base model capability table and platform-level grounding support, so web-search behavior should be verified for the specific deployment before relying on it.

Supported inputs and outputs

AreaDocumented status
Text inputSupported
Text outputSupported
Image, audio, and video inputNot listed as supported for this model
Image, audio, and video outputNot supported
Structured JSON outputSupported in non-thinking mode, subject to deployment and mode limits
StreamingSupported
Context cachingSupported
Batch inferenceSupported in the international deployment
Fine-tuningNot listed as supported
Function callingNot listed as supported

Where Qwen-Turbo fits best

Qwen-Turbo is most useful when an application processes a large volume of conventional text and values low cost and fast responses more than frontier-level reasoning. Suitable workloads include:

  • Customer-support replies and conversation summarization
  • Simple question answering over supplied content
  • Summarization of reports, tickets, and documents
  • Rewriting, classification, tagging, and moderation workflows
  • Structured extraction into JSON fields
  • Streaming text experiences where users should see output quickly
  • Offline or asynchronous processing that can use discounted batch inference

For example, a support system could use non-thinking mode to classify incoming tickets, extract order details, and draft a response. A document pipeline could summarize large text collections in batches. These are relatively constrained tasks where predictable throughput and token cost may matter more than advanced planning.

When to choose Qwen-Turbo

Choose Qwen-Turbo when the workload is primarily text-based, the task is simple or moderately difficult, and operating cost and response speed are important. Its 131,072-token context window is useful for applications that need to provide substantial source material, while streaming, caching, and batch inference offer different ways to manage user experience and processing cost.

Non-thinking mode is the natural starting point for routine production work. Thinking mode may be appropriate for selected requests that need more deliberate reasoning, but it should be tested against the higher output price and the actual quality requirement.

Another option may be more appropriate when the application requires image or audio understanding, reliable tool invocation, fine-tuning, complex coding, advanced reasoning, or an actively maintained model for a new long-term system. Alibaba Cloud specifically recommends Qwen-Flash for new deployments because Qwen-Turbo will no longer receive updates.

Practical limitations and lifecycle status

The most important limitation is not only technical capability but product status. Qwen-Turbo remains available, yet Alibaba Cloud has stated that it will no longer be updated. This creates a distinction between using it for an existing compatible workload and selecting it as the foundation for a new system.

Applications should also account for regional differences. Pricing, structured-output behavior, batch availability, and other platform capabilities may vary between international and China deployments. The maximum output is 16,384 tokens, while the maximum input is 98,304 tokens, so requests that approach the full context window must leave room for the generated response.

The reviewed official documentation does not specify a knowledge-cutoff date. Developers should avoid assuming that Qwen-Turbo has current knowledge of events or rapidly changing information, and should provide authoritative source material when the task depends on information that may have changed.

Bottom line

Qwen-Turbo is a fast, inexpensive text model for routine, high-volume workloads. Its strongest practical advantages are low listed token prices, a large context window, streaming, caching, structured output in non-thinking mode, and batch inference. Its weaker fit is equally clear: it is not documented as a multimodal, function-calling, fine-tunable, or advanced coding model, and it is no longer on an active update path. It can remain a sensible option for compatible existing workloads, but teams starting a new deployment should evaluate Alibaba Cloud’s recommended Qwen-Flash successor first.


Answers to Frequently Asked Questions

What is Qwen-Turbo best used for?
Qwen-Turbo is best suited to high-volume, text-based workloads such as summarization, classification, rewriting, moderation, simple question answering, customer-support drafting, and structured data extraction. Its low cost and fast non-thinking mode make it more appropriate for routine tasks than advanced reasoning or complex software engineering.
What are Qwen-Turbo’s context, input, and output limits?
Qwen-Turbo has a documented context window of 131,072 tokens, with a maximum input length of 98,304 tokens and a maximum output length of 16,384 tokens. Applications using large prompts must leave enough room within the total context window for the model’s response.
How much does Qwen-Turbo cost?
For the international deployment, non-thinking input costs $0.05 per 1 million tokens and non-thinking output costs $0.20 per 1 million tokens. Thinking-mode input is listed at $0.05 per 1 million tokens and thinking-mode output at $0.50 per 1 million tokens. China pricing differs, and international batch inference may receive a 50% discount where supported.
Does Qwen-Turbo support multimodal input, function calling, or structured output?
Qwen-Turbo is documented primarily as a text-in, text-out model and does not list image, audio, or video input support. Function calling and fine-tuning are also not listed as supported. Structured JSON output is supported in non-thinking mode, subject to deployment and regional limitations.
Should I use Qwen-Turbo for a new deployment?
Qwen-Turbo can remain suitable for existing compatible workloads that prioritize low cost and speed, but Alibaba Cloud states that it will no longer be updated and recommends Qwen-Flash for new deployments. Teams building a long-term system should evaluate Qwen-Flash or another actively maintained model, especially if they need advanced reasoning, complex coding, multimodal capabilities, or tool invocation.


Sources 8
Provider

About Qwen