Qwen-Plus

Qwen-Plus

by Qwen · Current; qwen-plus currently resolves to the qwen-plus-2025-12-01 snapshot

Qwen-Plus is Alibaba Cloud Model Studio’s general-purpose text model, currently mapped to qwen-plus-2025-12-01. It offers up to 1 million tokens of context, 32,768-token outputs, non-thinking and thinking modes, structured outputs, function calling, and web search in selected regions. The model is best suited to long-context business workflows, multilingual text tasks, and tool-using applications, but it does not support native media input or output, fine-tuning, or batch inference according to the supplied specifications.

Text Reasoning Coding
Qwen-Plus is Alibaba Cloud’s middle-ground general-purpose language model for production applications that need more than basic text generation but do not necessarily require the highest-cost reasoning configuration. Its current canonical identifier, qwen-plus, maps to the qwen-plus-2025-12-01 snapshot. The model accepts and produces text, supports long contexts of up to 1 million tokens, and can use tools such as function calls and web search in supported Model Studio regions. Its main trade-off is that pricing and some capabilities vary by deployment region and input length.
Outputs

What Qwen-Plus can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Web search Streaming Structured output
Model profile

Performance characteristics

7/10 Reasoning
7/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Qwen-Plus
Model type General Purpose
Context window 1M tokens
Maximum output 33K tokens
Release date 2025-12-01
Status Current; qwen-plus currently resolves to the qwen-plus-2025-12-01 snapshot
Knowledge cutoff notes

Alibaba Cloud's current Qwen-Plus model documentation does not provide a verified knowledge-cutoff date for the exact qwen-plus-2025-12-01 snapshot.

Model notes

Alibaba Cloud documents qwen-plus as functionally equivalent to qwen-plus-2025-12-01. The December 1, 2025 snapshot improved reasoning, agent behavior, multi-turn tool invocation, and subjective creative-task performance. The model supports separate non-thinking and thinking modes, but pricing and capability availability vary by deployment region. Structured outputs are documented, while a separate legacy JSON-mode capability is not independently confirmed. Web search and function calling are supported in some Model Studio regions, including China (Beijing), Singapore, and Germany; US and Hong Kong capability tables list more restrictions. The model is text-in/text-out and does not natively generate images, audio, or video.

Cost

Model pricing

Input $0.40 per 1M input tokens for 0–256K input; $1.20 per 1M input tokens above 256K in the International Singapore pricing tier. International Global pricing is $0.115 per 1M tokens for up to 128K, $0.345 for 128K–256K, and $0.689 for 256K–1M.
Output $1.20 per 1M non-thinking output tokens and $4.00 per 1M thinking output tokens for 0–256K input; $3.60 non-thinking and $12.00 thinking per 1M output tokens above 256K in the International Singapore pricing tier. International Global pricing differs by r
Model guide

Qwen-Plus: Alibaba Cloud’s Balanced Long-Context Model for Text and Tool Use

Qwen-Plus is Alibaba Cloud Model Studio’s general-purpose text model for applications that need a practical balance of reasoning quality, speed, cost, multilingual capability, long-context processing, structured output, and tool use. The current qwen-plus identifier maps to the qwen-plus-2025-12-01 snapshot, which supports up to 1 million input tokens, up to 32,768 output tokens, function calling, web search, structured outputs, and separate non-thinking and thinking modes.

What is Qwen-Plus?

Qwen-Plus is a general-purpose large language model provided through Alibaba Cloud Model Studio. It is intended for applications that need a balance between output quality, reasoning, responsiveness, tool use, and operating cost. Rather than being limited to a narrow task, it can handle common business and developer workflows such as drafting, summarization, multilingual text processing, long-document analysis, structured data extraction, and agent-style interactions.

Alibaba Cloud currently maps the general qwen-plus identifier to the qwen-plus-2025-12-01 snapshot. That distinction matters for reproducibility: the short identifier is convenient for ongoing use, while the dated snapshot identifies the current model version documented in the supplied research. Alibaba Cloud describes the December 2025 snapshot as improving reasoning, agent behavior, multi-turn tool invocation, and subjective performance on creative tasks.

Qwen-Plus is text-in and text-out. It is not a native image, audio, or video model, and it does not generate those media types directly.

Where Qwen-Plus fits in Alibaba Cloud’s lineup

Qwen-Plus occupies a broad middle position in Alibaba Cloud’s current model catalog. It is more feature-oriented than a text-only completion endpoint because it supports structured outputs, function calling, web search, long context, and an optional thinking mode. At the same time, it is positioned for cost-conscious general-purpose use rather than as a specialist image, audio, video, or fine-tuning model.

This positioning makes it suitable when one application must cover several types of work. A customer-service workflow might use ordinary generation for a quick response, structured output to return a predictable case record, a function call to retrieve account information, and web search when current external information is required. The same model can also process a large collection of documents when the application stays within the documented context limit.

Key specifications

SpecificationQwen-Plus
Current snapshotqwen-plus-2025-12-01
ProviderAlibaba Cloud
Maximum context length1,000,000 tokens
Maximum output32,768 tokens
Input and outputText input and text output
Reasoning modesNon-thinking and thinking modes
Structured outputsSupported
Function callingSupported
Web searchSupported in some Model Studio regions
Fine-tuningNot supported according to the supplied model data

The 1-million-token context limit is an especially important specification. Context is the amount of text the model can consider in one request, including instructions, conversation history, documents, tool material, and the requested response. A large context can reduce the need to split long reports or code repositories into many separate requests. It does not, however, guarantee that every long document will receive equal attention or that a long prompt will be economical.

The maximum output is 32,768 tokens. This is a ceiling rather than a promise that every request will produce that much text. Actual output depends on the prompt, generation settings, selected mode, and application limits.

Reasoning, coding, speed, and cost

Qwen-Plus provides separate non-thinking and thinking modes. Non-thinking mode is the more direct choice when an application prioritizes lower latency, shorter responses, and lower output cost. Thinking mode is intended for tasks where additional reasoning effort is useful, such as multi-step analysis, difficult planning, or agent workflows that require more deliberate decisions. The supplied research confirms the two modes and their separate pricing, but does not provide a standardized public benchmark result for this page.

Editorial evaluations in the supplied data rate Qwen-Plus at 7 out of 10 for reasoning, coding, and speed, and 8 out of 10 for cost. These are comparative editorial assessments, not scores published by Alibaba Cloud and not substitutes for testing the model on a particular workload. In practice, coding quality can depend heavily on the programming language, repository size, requested change, and whether tools are available to inspect or test code.

The practical trade-off is straightforward: use the less expensive non-thinking path for routine generation, extraction, classification, and straightforward coding assistance; reserve thinking mode for requests where deeper analysis is worth additional latency and output cost. Because the model supports a very large context, the total input bill can also become significant when repeatedly sending large documents or conversation histories.

Qwen-Plus pricing

Pricing depends on the deployment region, input length, and whether the model uses thinking mode. The following figures are from the International Singapore pricing tier supplied for this model:

  • For 0–256K input tokens, input costs $0.40 per 1 million tokens.
  • For more than 256K input tokens, input costs $1.20 per 1 million tokens.
  • For 0–256K input tokens, non-thinking output costs $1.20 per 1 million tokens.
  • For 0–256K input tokens, thinking output costs $4.00 per 1 million tokens.
  • Above 256K input tokens, non-thinking output costs $3.60 per 1 million tokens.
  • Above 256K input tokens, thinking output costs $12.00 per 1 million tokens.

Alibaba Cloud also lists different International Global prices: $0.115 per 1 million input tokens for up to 128K input, $0.345 for 128K–256K, and $0.689 for 256K–1M. The supplied research indicates that output pricing also differs by region and tier, so users should check the pricing table for the exact deployment location rather than applying Singapore rates universally.

These are token-based prices, not a flat subscription price. A request with a large prompt and a long generated answer consumes both input and output tokens. Thinking mode can be substantially more expensive on the Singapore tier, particularly above 256K input tokens. For predictable operating costs, applications should limit unnecessary history, avoid resending unchanged documents where possible, and select thinking mode only for tasks that benefit from it.

Tools, function calling, and structured workflows

Qwen-Plus supports function calling, a mechanism that lets the model request an operation defined by the application. For example, the model can decide that it needs a customer record, inventory lookup, calculation, or internal search, then return the arguments for the application to execute. The application remains responsible for performing the operation and deciding whether the result is safe to use.

Structured outputs are also documented. They are useful when the application needs a response in a predictable schema rather than free-form prose, such as a list of extracted fields, a support-ticket object, or a set of classification labels. Structured output support should not automatically be treated as proof of a separate legacy JSON-mode feature; the supplied research specifically confirms structured outputs but does not independently confirm that distinct capability.

Web search is supported in some Model Studio regions, including China (Beijing), Singapore, and Germany. The supplied capability information also notes greater restrictions in some other regions, including the United States and Hong Kong. Availability therefore depends on the deployment region and the relevant Model Studio capability table. An application that relies on web search should verify availability before committing to a regional architecture.

Modalities and important limitations

Qwen-Plus is a text model. Its supported modality combination is text input and text output. It does not natively accept images, audio, or video according to the supplied specifications, and it does not produce image, audio, or video output. Applications built around visual documents, speech, or video should use a suitable multimodal or media-specific model before passing any resulting text to Qwen-Plus.

The model is also listed as not supporting fine-tuning, batch inference, or caching in the supplied model data. That limits some optimization and customization strategies. If an application needs a privately adapted model, a dedicated fine-tuning-capable option may be more appropriate. If it needs large-scale offline processing through a batch interface, a model and service with verified batch support would be a better fit.

Pricing is another limitation. There is no single universal rate that applies across every Model Studio region, and the price changes with input length and reasoning mode. Teams should therefore evaluate the full request pattern rather than comparing only the lowest advertised per-token number.

Best use cases for Qwen-Plus

  • Long-document analysis: The 1-million-token context can help applications work with very large collections of text in one request, subject to practical cost and quality testing.
  • Business automation: Structured outputs and function calling fit workflows such as ticket routing, record extraction, approval preparation, and internal knowledge tasks.
  • Multilingual text applications: Its broad general-purpose positioning makes it suitable for translation-related workflows, multilingual drafting, and cross-language information handling, although the supplied research does not provide language-specific benchmark scores.
  • Agent workflows: Function calling, web search in supported regions, and the documented improvements to multi-turn tool invocation support applications that must combine model decisions with external actions.
  • Mixed-difficulty workloads: Applications can use non-thinking mode for routine requests and thinking mode for harder cases, provided the deployment and pricing configuration supports the required behavior.

When to choose Qwen-Plus

Choose Qwen-Plus when one text model needs to cover ordinary generation, long-context analysis, structured responses, and tool-assisted workflows without requiring native media understanding or generation. It is particularly compelling when a large context window reduces document chunking and when the application can control costs by using non-thinking mode for simpler requests.

Another option may be more appropriate when the workload is primarily image, audio, or video based; when fine-tuning is a core requirement; when batch inference is essential; or when the application needs a clearly uniform price across regions. A smaller or simpler text model may also be preferable for high-volume, low-complexity requests where Qwen-Plus’s long context and tool features are unnecessary. Conversely, a more specialized reasoning model may be worth evaluating for tasks that consistently require deeper reasoning than the general-purpose balance offered here.

For production selection, test representative prompts in both modes, measure latency and token use, verify structured-output behavior against the exact schema, and confirm function-calling and web-search availability in the intended region. Those checks are more informative than relying on a general model score because Qwen-Plus’s value depends on how much an application benefits from its long context and integrated tool capabilities.


Answers to Frequently Asked Questions

Does Qwen-Plus support function calling, structured outputs, and web search?
Yes. Qwen-Plus supports function calling and structured outputs. Web search is available in some Alibaba Cloud Model Studio regions, including China (Beijing), Singapore, and Germany, but availability should be verified for the intended deployment region.
What is the context window of Qwen-Plus?
Qwen-Plus supports a maximum context length of 1,000,000 tokens and a maximum output of 32,768 tokens. The context includes instructions, conversation history, documents, tool results, and the requested response.
What is the difference between Qwen-Plus thinking and non-thinking modes?
Non-thinking mode is designed for lower latency, shorter responses, and lower cost on routine tasks. Thinking mode provides additional reasoning effort for difficult analysis, planning, and agent workflows, but it generally has higher latency and output pricing.
What is Qwen-Plus?
Qwen-Plus is a general-purpose text-in and text-out large language model provided through Alibaba Cloud Model Studio. It supports drafting, summarization, multilingual processing, long-document analysis, structured data extraction, function calling, web search in some regions, and agent-style workflows.


Sources 6
Provider

About Qwen