Qwen3-Coder

Qwen3-Coder-Next

by Qwen · Current; open-weight model and available through Alibaba Cloud Model Studio

Qwen3-Coder-Next is an Apache 2.0 open-weight coding model with approximately 80B total and 3B active parameters, a 262,144-token context window, fill-in-the-middle completion, and training for repository-scale agentic software development. It supports text input and output, can run through Alibaba Cloud Model Studio or self-hosted vLLM and SGLang deployments, and has regional token pricing. Its main limitations are the lack of native multimodal output and several unsupported hosted features, including structured outputs, web search, caching, batch inference, and function calling.

Text Reasoning Coding
Qwen3-Coder-Next is a specialized Qwen model for code generation, repository analysis, code completion, debugging, refactoring, and software-development agents. It accepts text and produces text, supports long programming contexts, and can be deployed through Alibaba Cloud Model Studio or self-hosted using compatible inference tools. Its main trade-off is focus: it is designed for coding workflows rather than multimodal assistance, and several managed endpoint features remain unavailable.
Outputs

What Qwen3-Coder-Next can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Streaming
Model profile

Performance characteristics

8/10 Reasoning
9/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Qwen3-Coder
Model type Coding
Context window 262K tokens
Maximum output 66K tokens
Release date 2026-02-03
Status Current; open-weight model and available through Alibaba Cloud Model Studio
Knowledge cutoff notes

No direct authoritative knowledge-cutoff date was identified in the official Alibaba Cloud model documentation, Qwen repository, or model card.

Model notes

Qwen3-Coder-Next is an 80B-parameter mixture-of-experts model with approximately 3B active parameters. It is based on Qwen3-Next-80B-A3B and uses a hybrid architecture with gated attention and Gated DeltaNet layers. The open-weight model supports a specialized tool-call format when deployed with compatible vLLM or SGLang parsers, but Alibaba Cloud Model Studio lists function calling as unsupported for its hosted endpoint. Alibaba's hosted documentation lists a 204,800-token maximum input length and 65,536-token maximum output length within a 262,144-token context window. The model supports non-thinking mode only and does not emit think blocks. It supports fill-in-the-middle completion. Pricing varies by region and input-length band; listed prices are original API prices before temporary promotions. The open-weight release is distributed under the Apache 2.0 license.

Cost

Model pricing

Input USD 0.144 per 1M tokens for China Beijing input up to 32K; USD 0.216 for 32K–128K; USD 0.359 for 128K–256K. International Singapore and Frankfurt pricing is USD 0.30, USD 0.50, and USD 0.80 per 1M input tokens across the same bands.
Output USD 0.574 per 1M tokens for China Beijing output up to 32K; USD 0.861 for 32K–128K; USD 1.434 for 128K–256K. International Singapore and Frankfurt pricing is USD 1.50, USD 2.50, and USD 4.00 per 1M output tokens across the same bands.
Model guide

Qwen3-Coder-Next: Open-Weight Coding Model for Long-Context Agent Work

Qwen3-Coder-Next is Alibaba's open-weight coding model for repository-scale software development and agentic coding. Its mixture-of-experts design has 80 billion total parameters but activates approximately 3 billion per token, combining a 262,144-token context window with lower active-compute requirements than a comparably sized dense model.

What is Qwen3-Coder-Next?

Qwen3-Coder-Next is an open-weight causal language model from Alibaba's Qwen team. It is positioned as a coding specialist within the Qwen family, rather than as a general-purpose multimodal assistant. The model is intended for software-development tasks that require understanding multiple files, maintaining context over long interactions, using tools, and recovering from execution errors.

The model was released on February 3, 2026, according to the supplied model research. It is available as downloadable open weights under the Apache 2.0 license and is also offered through Alibaba Cloud Model Studio. That gives users two substantially different deployment paths: a managed regional API and self-hosted inference using an appropriate serving stack.

Qwen3-Coder-Next is based on Qwen3-Next-80B-A3B. The name reflects its mixture-of-experts structure: the model contains approximately 80 billion parameters in total, but about 3 billion parameters are active for each token processed. In practical terms, this means the model has a large parameter capacity while attempting to keep per-token computation and serving cost closer to a much smaller active model. Actual speed and memory requirements still depend on the inference engine, quantization, hardware, batching, and context length.

Architecture and context limits

Qwen3-Coder-Next uses a hybrid architecture combining gated attention, Gated DeltaNet layers, and mixture-of-experts routing. It has 48 layers and a native context window of 262,144 tokens. A context window is the amount of input and generated conversation history the model can process within one request, although the usable amount can be lower in a particular application because prompts, tool results, system instructions, and output all consume context.

Alibaba Cloud's hosted documentation lists a maximum input length of 204,800 tokens and a maximum output length of 65,536 tokens. These hosted limits should not be confused with the model's native 262,144-token context capacity. The distinction matters for large repositories: a self-hosted deployment may expose different limits, while the managed endpoint applies its own input and output restrictions.

The long context is particularly relevant to repository-level work. A coding agent can potentially keep more source files, documentation, test output, and task history in view than a short-context model. However, a large context does not guarantee that every detail will receive equal attention. Practical results still depend on how the agent selects files, summarizes history, and structures tool results.

Coding and agentic capabilities

The model's primary purpose is software development. Supported use cases include generating new code, completing partially written code, explaining repository structure, debugging failures, refactoring existing implementations, and working through terminal-oriented development tasks. It also supports fill-in-the-middle completion, where the model generates code that belongs between an existing prefix and suffix.

Qwen3-Coder-Next was trained for longer software tasks rather than only isolated code snippets. The supplied technical material describes repository-level understanding, executable-task synthesis, environment interaction, recovery from execution failures, and multi-turn tool interaction as important parts of its design. These capabilities are useful when a coding assistant must inspect a project, make a change, run tests, interpret an error, and revise its implementation.

For self-hosted use, Qwen documents a specialized tool-call format and compatibility with serving configurations for tools such as vLLM and SGLang. Qwen's documentation also identifies integrations with coding environments and agent frameworks including Qwen Code, Cline, Claude Code, Qoder, Kilo, and Trae. These integrations are not the same as the model independently executing actions: the surrounding agent or serving system must parse requests, invoke tools, return results, and enforce permissions.

Hosted tools and API limitations

There is an important difference between the open-weight release and the Alibaba Cloud Model Studio endpoint. The open-weight model can participate in compatible tool workflows when the serving layer provides the relevant parser and orchestration. In contrast, Alibaba's hosted capability table lists function calling as unsupported for the Qwen3-Coder-Next endpoint.

The hosted endpoint also lists web search, structured outputs, context caching, batch inference, and fine-tuning as unsupported for this exact model. This means users should not assume that a coding model with a tool-oriented training profile automatically provides every managed API feature. A self-hosted tool-call format can help with agent integration, but it does not turn the Model Studio endpoint into a verified structured-output or native function-calling service.

Qwen3-Coder-Next accepts text and returns text. It does not natively generate or understand images, audio, video, spoken speech, or embeddings according to the supplied specifications. Users who need a single assistant for visual files, voice conversations, image creation, or video generation should consider a multimodal model or a provider-level assistant instead of selecting this model solely for its Qwen branding.

Pricing and availability

Alibaba Cloud Model Studio prices Qwen3-Coder-Next by tokens rather than through a consumer subscription. The following figures are the supplied original API prices before temporary promotions and vary by region and input-length band.

RegionInput up to 32KInput 32K–128KInput 128K–256KOutput up to 32KOutput 32K–128KOutput 128K–256K
China Beijing$0.144 per 1M tokens$0.216 per 1M tokens$0.359 per 1M tokens$0.574 per 1M tokens$0.861 per 1M tokens$1.434 per 1M tokens
Singapore and Frankfurt$0.30 per 1M tokens$0.50 per 1M tokens$0.80 per 1M tokens$1.50 per 1M tokens$2.50 per 1M tokens$4.00 per 1M tokens

These are token rates, not a guaranteed price for a complete coding task. Large repositories, repeated tool results, long conversation histories, and generated patches can increase consumption. The higher rates for longer input bands also make prompt and context management important. Self-hosting removes per-token Model Studio charges but shifts costs to hardware, storage, electricity, operations, and model-serving expertise.

Strengths and trade-offs

The main strength of Qwen3-Coder-Next is specialization. Its long context, repository-oriented training, fill-in-the-middle support, and agentic coding focus make it a logical candidate for tasks that go beyond generating a short function. The approximately 3-billion active-parameter design may also offer a useful cost and speed profile compared with dense models containing a similar total number of parameters, although real-world performance must be measured on the target hardware and workload.

Its open-weight Apache 2.0 release is another practical advantage. Organizations can evaluate local deployment, choose compatible serving infrastructure, and retain more control over where code and repository data are processed. Quantized variants can reduce hardware requirements, but the supplied research does not establish a universal minimum memory requirement; deployment needs depend on precision, context size, concurrency, and the serving framework.

The limitations are equally important. Qwen3-Coder-Next is text-only, so it is not appropriate when the coding workflow depends on screenshots, diagrams, audio, or video as native inputs. The hosted endpoint does not provide several capabilities that some managed coding platforms offer, including structured outputs, native function calling, web search, batch inference, caching, and fine-tuning. The open-weight release can support tool interaction through an external serving layer, but that introduces integration and maintenance work.

The supplied model notes also describe the model as supporting non-thinking mode only and not emitting think blocks. Users seeking a model with an explicitly exposed extended-reasoning mode should verify that requirement separately rather than infer it from the model's ability to solve difficult coding tasks.

When to choose Qwen3-Coder-Next

Choose Qwen3-Coder-Next when the central problem is software engineering across a substantial codebase. It is a strong candidate for repository-scale coding assistants, autonomous or semi-autonomous development agents, code completion, debugging, refactoring, terminal workflows, and teams evaluating an open-weight model for local or private deployment.

  • Choose it for long coding contexts: the 262,144-token native context and 204,800-token hosted input limit are useful when a task requires many files, documentation pages, or test logs.
  • Choose it for open deployment: the Apache 2.0 weights allow self-hosting with compatible infrastructure rather than requiring every request to use the managed endpoint.
  • Choose it for cost-sensitive coding workloads: its mixture-of-experts design activates approximately 3B parameters per token, though actual economics depend on hardware and serving configuration.
  • Choose another type of model for multimodal work: image, audio, video, speech, or visual debugging tasks require capabilities this model does not natively provide.
  • Choose another managed endpoint when API features are essential: projects that require verified structured output, native hosted function calling, web search, batch processing, caching, or fine-tuning should confirm that the selected service supports those features.

In short, Qwen3-Coder-Next is best understood as a focused coding engine rather than a complete general-purpose assistant. Its value comes from combining open weights, a very large coding context, mixture-of-experts efficiency, and training aimed at sustained repository work. The correct choice depends on whether those advantages outweigh the additional deployment complexity and the absence of native multimodal and several managed API features.


Answers to Frequently Asked Questions

How much does Qwen3-Coder-Next cost on Alibaba Cloud Model Studio?
Model Studio charges for Qwen3-Coder-Next by input and output tokens, with prices varying by region and context-length band. The supplied rates range from $0.144 to $0.359 per 1 million input tokens and from $0.574 to $1.434 per 1 million output tokens in China Beijing, while Singapore and Frankfurt rates range from $0.30 to $0.80 for input and $1.50 to $4.00 for output per 1 million tokens.
Does Qwen3-Coder-Next support function calling and multimodal inputs?
The open-weight model can participate in tool workflows when an external serving and orchestration layer supports its tool-call format. However, Alibaba Cloud's hosted endpoint lists function calling as unsupported. The model is also text-only and does not natively support images, audio, video, speech, or embeddings.
Can Qwen3-Coder-Next be self-hosted?
Yes. Qwen3-Coder-Next is available as downloadable open weights under the Apache 2.0 license, allowing organizations to self-host it with compatible inference systems such as vLLM or SGLang. Hardware and memory requirements depend on precision, quantization, context length, concurrency, and the serving framework.
What is Qwen3-Coder-Next designed for?
Qwen3-Coder-Next is an open-weight coding model designed for repository-level software development, including code generation, completion, debugging, refactoring, terminal workflows, multi-file reasoning, and agentic tasks involving tools and execution feedback.
How large is Qwen3-Coder-Next's context window?
Qwen3-Coder-Next has a native context window of 262,144 tokens. Alibaba Cloud Model Studio lists a maximum input length of 204,800 tokens and a maximum output length of 65,536 tokens for its hosted endpoint.


Sources 5
Provider

About Qwen