QwQ

qwq-plus

by Qwen · Legacy; currently accessible as of October 7, 2026; scheduled to cease service on October 10, 2026 at 00:00 UTC+08, subject to the actual change time

QwQ-Plus is Alibaba Cloud’s Qwen2.5-based reasoning model for mathematics, coding, and complex analytical tasks. It offers a 131,072-token context window, regional function calling and web search, and low China deployment pricing, but it is legacy software scheduled to cease service on October 10, 2026.

Text Reasoning Coding
QwQ-Plus is an Alibaba Cloud Model Studio reasoning model built for difficult mathematical, programming, and analytical problems. It accepts text and returns text, supports a 131,072-token context window, and can use tools such as web search and function calling in the China (Beijing) deployment. Its low listed token prices make it attractive for supported workloads, but its legacy status and announced October 10, 2026 service shutdown make it unsuitable for new long-term production dependencies.
Outputs

What qwq-plus can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Web search Streaming Batch API
Model profile

Performance characteristics

9/10 Reasoning
8/10 Coding
6/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family QwQ
Model type Reasoning
Context window 131K tokens
Maximum output 8K tokens
Status Legacy; currently accessible as of October 7, 2026; scheduled to cease service on October 10, 2026 at 00:00 UTC+08, subject to the actual change time
Shutdown date 2026-10-10
Knowledge cutoff notes

Alibaba Cloud's current model documentation does not publish a knowledge-cutoff date for qwq-plus.

Model notes

Canonical Alibaba Cloud Model Studio model ID is qwq-plus. The model is described as an enhanced QwQ reasoning model trained from Qwen2.5 with reinforcement learning. Alibaba reports strong AIME, LiveCodeBench, IFEval, and LiveBench performance. Function calling, web search, and batch inference are supported in the China (Beijing) deployment but unsupported in the Singapore international deployment. Structured outputs, prefix completion, context caching, and fine-tuning are listed as unsupported. Batch file pricing in China is $0.115 per 1M input tokens and $0.287 per 1M output tokens. Alibaba Cloud announced service cessation for October 10, 2026.

Cost

Model pricing

Input $0.230 per 1M tokens in China (Beijing); $0.80 per 1M tokens in Singapore international deployment
Output $0.574 per 1M tokens in China (Beijing); $2.40 per 1M tokens in Singapore international deployment
Model guide

QwQ-Plus: Alibaba’s Reasoning Model for Mathematics and Complex Analysis

QwQ-Plus is Alibaba Cloud Model Studio’s Qwen2.5-based reasoning model for mathematics, coding, and other multi-step analytical tasks. It offers a 131,072-token context window, region-dependent tools such as web search and function calling, and relatively low token pricing, but it is classified as a legacy model scheduled to cease service on October 10, 2026.

What is QwQ-Plus?

QwQ-Plus is Alibaba Cloud Model Studio’s enhanced QwQ reasoning model. Alibaba describes it as a model trained from Qwen2.5 and improved with reinforcement learning, a training approach intended to strengthen performance on problems that require multiple reasoning steps rather than a short factual response.

The model’s primary focus is complex problem solving. Typical tasks include working through mathematics, generating or debugging code, comparing evidence, and answering research-style questions that benefit from a longer chain of analysis. Its canonical Model Studio model ID is qwq-plus.

QwQ-Plus is a text-in, text-out model. It does not natively produce images, audio, or video, so it should be evaluated as a reasoning and language model rather than as a general multimodal generation system.

Where QwQ-Plus fits in Alibaba’s catalog

QwQ-Plus belongs to Alibaba’s QwQ reasoning family and is offered through Alibaba Cloud Model Studio. Its role is more specialized than that of a general-purpose chat model: the model is intended to spend more effort on multi-step reasoning, especially for mathematics and programming.

However, Alibaba Cloud classifies QwQ-Plus as a legacy model. The provider has announced that service is scheduled to cease at 00:00 UTC+08 on October 10, 2026, subject to the actual change time. It remained accessible according to the supplied documentation as of October 7, 2026, but this planned shutdown is an important part of its practical positioning. Users selecting a model for a new application should generally prefer a currently supported reasoning successor where one is available instead of building a long-term dependency on QwQ-Plus.

Reasoning and benchmark positioning

QwQ-Plus is designed for tasks where the answer depends on several linked deductions. In practice, that makes it a better fit for solving a mathematical exercise, explaining a debugging path, or analyzing a complicated question than for a simple high-volume classification task where the main priorities are minimal latency and low cost.

Alibaba reports that QwQ-Plus reached results comparable to the full version of DeepSeek-R1 on selected mathematics, coding, and general benchmarks. The cited evaluations include AIME 2024 and AIME 2025, LiveCodeBench, IFEval, and LiveBench. These are provider-reported benchmark claims rather than an independent assessment, and benchmark performance should not be treated as a guarantee for every production workload.

The model’s reasoning orientation can increase the usefulness of answers to difficult prompts, but it also contributes to an important trade-off: a reasoning-focused model may be slower or more expensive in practical use than a smaller, speed-oriented text model. The supplied specifications do not publish a universal response-latency figure, so actual speed will depend on deployment, prompt size, output length, and service conditions.

Context window and output limits

QwQ-Plus supports a maximum context of 131,072 tokens. A token is a unit of text used by the model; the context includes the material supplied to the model and the conversation or other content retained for the request. A large context is useful when a task involves long source documents, extensive code, or a multi-step discussion.

The documented maximum input length is 98,304 tokens, while the maximum output length is 8,192 tokens. These limits are distinct: the model can receive a very large request, but it cannot return an unlimited-length answer in one generation. Applications should also account for system instructions, conversation history, tool-related content, and the requested response when planning within the available context.

The model supports text input and text output only. It does not support native image, audio, or video input or output according to the supplied specifications. Structured outputs are listed as unsupported, so applications that require responses conforming reliably to a predefined JSON schema should consider a model and deployment that explicitly supports that feature.

Tools and regional support

QwQ-Plus has different capabilities depending on where it is deployed. In the China (Beijing) deployment, the documented feature set includes function calling, web search, and batch inference. Function calling allows an application to connect the model to external operations, while web search can provide access to retrieved information through the supported Model Studio workflow.

These tools should not be assumed to be available everywhere. The Singapore international deployment supports text generation and model experience, but the supplied documentation does not list function calling, web search, or batch inference there. A deployment decision therefore affects more than price: it can change which application features are possible.

Batch inference is useful when many requests can be processed without interactive, one-at-a-time behavior. In China, Alibaba also lists separate batch file-processing prices. Tool and batch support should be verified for the exact region and interface before an application is designed around them.

Pricing by deployment

Alibaba Cloud lists token-based pricing, with separate rates for input and output. The rates below are stated per one million tokens and differ substantially between the China (Beijing) and Singapore international deployments.

DeploymentInputOutputNotable support
China (Beijing)$0.230 per 1M tokens$0.574 per 1M tokensFunction calling, web search, and batch inference listed as supported
China batch file processing$0.115 per 1M tokens$0.287 per 1M tokensBatch pricing; interactive capabilities may differ
Singapore international$0.80 per 1M tokens$2.40 per 1M tokensText generation and model experience listed; regional tools are not listed

These are usage prices rather than a monthly subscription price. Input and output are billed separately, so a workload that generates long reasoning responses can cost more than one with short answers even when the prompts are similar. Alibaba also lists a limited free quota for the Singapore deployment, subject to the provider’s applicable conditions.

On the supplied prices, the China deployment is considerably less expensive than the Singapore international deployment. That difference may make China attractive for eligible workloads, but region, data-handling requirements, tool availability, and operational constraints should be evaluated alongside the nominal token rate.

Strengths and limitations

Main strengths

  • Reasoning focus: The model is specifically positioned for mathematics, coding, and multi-step analysis rather than only short conversational responses.
  • Large context: The 131,072-token context window can accommodate long documents, sizeable codebases, or extended analytical exchanges.
  • Useful China deployment tools: Function calling, web search, and batch inference are listed for the China (Beijing) deployment.
  • Low China token rates: The listed China prices are substantially lower than the Singapore international rates.
  • Strong provider-reported evaluations: Alibaba reports results comparable to the full version of DeepSeek-R1 on selected mathematics, coding, instruction-following, and general benchmarks.

Important limitations

  • Planned retirement: The scheduled October 10, 2026 shutdown creates continuity risk for new applications.
  • Regional differences: Web search, function calling, and batch inference are listed for China but not for the Singapore international deployment.
  • Text-only modality: The model does not natively handle or generate images, audio, or video.
  • No listed structured outputs: Workflows that need schema-constrained responses may require another model or additional application-side validation.
  • No fine-tuning or context caching: The supplied specifications list both capabilities as unsupported.
  • Reasoning versus speed: Its focus on difficult reasoning may be unnecessary for simple, latency-sensitive tasks, and the documentation does not provide a universal speed guarantee.

Best use cases for QwQ-Plus

QwQ-Plus is most appropriate for workloads where answer quality on difficult reasoning problems matters more than immediate response speed. Suitable examples include:

  • Solving or explaining advanced mathematical problems.
  • Generating, reviewing, debugging, and analyzing software code.
  • Working through long research-style questions or analytical documents.
  • Producing explanations that require several connected steps.
  • Building China-based agent workflows that use supported web search or function calling.
  • Processing large batches of text in the China deployment when the batch pricing and regional availability meet the project’s requirements.

For example, a developer could use QwQ-Plus to inspect a lengthy code sample, identify interacting bugs, and propose a step-by-step repair plan. A research workflow could provide a large source collection and ask for a structured comparison, although the application should not assume native structured-output enforcement.

When to choose QwQ-Plus—and when not to

Choose QwQ-Plus when you need a text reasoning model for mathematics, coding, or complex analysis, have access to a suitable deployment region, and can accept its announced end-of-service date. The China deployment is particularly relevant when low token cost, web search, function calling, or batch inference are important and the deployment’s operational requirements are acceptable.

A speed-oriented or smaller language model may be more appropriate for straightforward extraction, short answers, simple classification, or high-volume interactive requests where every increment of latency matters. A multimodal model is a better choice when the workflow requires image, audio, or video input or generation. A model with explicit structured-output support is preferable when downstream software depends on guaranteed schema-conforming responses.

Most importantly, QwQ-Plus is a poor choice for a new production system that must remain available beyond October 10, 2026. Its capabilities and pricing may still be useful for an existing deployment, a short-lived project, evaluation, or migration period, but teams should identify and test a currently supported replacement before committing to a long-term architecture.

Bottom line

QwQ-Plus is a specialized Alibaba Cloud reasoning model with a large context window, strong provider-reported mathematics and coding results, and useful regional tools in China. Its main practical advantages are its focus on multi-step problem solving and its low China deployment pricing. Its decisive disadvantage is lifecycle: Alibaba Cloud has classified it as legacy and scheduled its service to end on October 10, 2026. Treat it as a capable but transitional option, not as the foundation for a new long-lived application.


Answers to Frequently Asked Questions

What is QwQ-Plus and what is it designed for?
QwQ-Plus is Alibaba Cloud Model Studio’s reasoning-focused text model, trained from Qwen2.5 and enhanced with reinforcement learning. It is designed for multi-step tasks such as advanced mathematics, coding, debugging, evidence comparison, and research-style analysis. Its canonical model ID is "qwq-plus".
What are QwQ-Plus’s context window and output limits?
QwQ-Plus supports a maximum context of 131,072 tokens, with a documented maximum input length of 98,304 tokens and a maximum output length of 8,192 tokens. It supports text input and text output only and does not natively support images, audio, or video.
How much does QwQ-Plus cost in different deployments?
Alibaba Cloud lists separate input and output token rates by deployment. China (Beijing) costs $0.230 per 1 million input tokens and $0.574 per 1 million output tokens. China batch file processing costs $0.115 per 1 million input tokens and $0.287 per 1 million output tokens. Singapore international costs $0.80 per 1 million input tokens and $2.40 per 1 million output tokens.
Which tools and features are available with QwQ-Plus?
The China (Beijing) deployment lists function calling, web search, and batch inference. These features are not listed for the Singapore international deployment, which lists text generation and model experience. QwQ-Plus does not list support for structured outputs, fine-tuning, or context caching.
Should developers use QwQ-Plus for a new long-term production application?
Generally, no. Alibaba Cloud classifies QwQ-Plus as a legacy model and has scheduled service cessation for 00:00 UTC+08 on October 10, 2026, subject to the actual change time. It may still suit existing deployments, short-lived projects, evaluations, or migration periods, but new long-term systems should use and test a currently supported reasoning successor.


Sources 5
Provider

About Qwen