Qwen3.6

Qwen3.6-27B

by Qwen · Current and available; open-weight release and Alibaba Cloud Model Studio API model

Qwen3.6-27B is Alibaba's 27B open-weight vision-language model for coding agents, repository reasoning, visual document analysis, video understanding, and STEM tasks. It accepts text, images, and video, returns text, supports thinking modes and tool calling, and has a 262,144-token native context window. Alibaba Cloud Model Studio offers hosted access with international pricing of $0.60 per million input tokens and $3.60 per million output tokens.

Text Reasoning Coding
Qwen3.6-27B is an open-weight multimodal model from Alibaba's Qwen team, designed especially for coding agents and visual reasoning rather than image or audio generation. It accepts text, images, and video, returns text, supports tool calling, and provides a 262,144-token native context window. Developers can run it through compatible inference frameworks or use the hosted Alibaba Cloud Model Studio endpoint.
Outputs

What Qwen3.6-27B can produce

Text
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output
Model profile

Performance characteristics

8/10 Reasoning
9/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Qwen3.6
Model type Multimodal
Context window 262K tokens
Maximum output 66K tokens
Release date 2026-04-24
Status Current and available; open-weight release and Alibaba Cloud Model Studio API model
Knowledge cutoff notes

No direct authoritative knowledge-cutoff date was identified in the Alibaba Cloud Model Studio documentation or the official Qwen model card.

Model notes

Qwen3.6-27B is a 27B dense causal language model with a vision encoder. It accepts text, images, and video and returns text. The native context length is 262,144 tokens, with Alibaba Cloud Model Studio listing a maximum input length of 260,096 and maximum output length of 65,536. The open-weight model can be extended to approximately 1.01 million tokens using RoPE or YaRN scaling, subject to framework and hardware constraints. It supports thinking and non-thinking modes. Alibaba Cloud Model Studio lists function calling, structured outputs, web search, and prefix completion as supported, while context caching, batch inference, and fine-tuning are unsupported for this endpoint. Editorial scores are comparative estimates, not vendor ratings.

Cost

Model pricing

Input $0.60 per 1M tokens internationally; $0.412564 per 1M tokens in China Beijing and Singapore
Output $3.60 per 1M tokens internationally; $2.475384 per 1M tokens in China Beijing and Singapore
Model guide

Qwen3.6-27B: Open-Weight Vision Model Built for Coding Agents

Qwen3.6-27B is Alibaba's 27-billion-parameter dense vision-language model for coding agents, repository-level reasoning, visual document analysis, video understanding, STEM tasks, and long-context applications. It accepts text, images, and video, produces text, supports thinking and non-thinking modes, and offers a 262,144-token native context window. The model is available as open weights and through Alibaba Cloud Model Studio.

What Qwen3.6-27B is

Qwen3.6-27B is a dense vision-language model developed by Alibaba's Qwen team and released as part of the Qwen3.6 model family. The “27B” designation refers to approximately 27 billion parameters, the learned values that the model uses to process inputs and generate responses. It is positioned as a practical model for software engineering, visual analysis, reasoning, and agent workflows.

The model is available in two main forms. Developers can use the open-weight release with compatible self-hosting tools, or access a managed version through Alibaba Cloud Model Studio. These options are related but not identical: local deployments depend on the chosen hardware and inference framework, while the hosted service defines its own limits, supported features, and pricing.

Alibaba's published materials describe improvements in repository-level reasoning, frontend development workflows, STEM problem solving, spatial intelligence, object localization, video understanding, document OCR, and visual-agent tasks. Those descriptions are provider claims; actual results can vary with prompts, system instructions, tool implementations, and deployment settings.

Input and output modalities

Qwen3.6-27B accepts text, images, and video as input and produces text as output. This makes it suitable for tasks such as asking questions about a screenshot, extracting information from a document image, locating an object in a scene, reviewing a screen recording, or combining a visual reference with a coding request.

  • Text input: supported.
  • Image input: supported for visual analysis and reasoning.
  • Video input: supported for video understanding.
  • Audio input: not listed as supported for this model.
  • Text output: supported.
  • Image, video, audio, music, and speech output: not supported natively.

The distinction between input and output is important. Qwen3.6-27B can understand an image or video, but it does not create a finished image, video, or spoken audio file. Applications that need those outputs should use a model designed for media generation or speech synthesis alongside Qwen3.6-27B.

Context window and reasoning modes

The native context window is 262,144 tokens. In Alibaba Cloud Model Studio, the documented maximum input length is 260,096 tokens and the maximum output length is 65,536 tokens. A token is a unit of text used by the model; the exact number of words represented by a token varies by language and content.

This large context is useful for repository analysis, lengthy technical documents, extended conversations, and multimodal tasks that require the model to consider substantial background information. It does not guarantee that every long input will receive equally reliable attention. Large requests also require more memory and may increase latency or cost.

The open-weight release can reportedly be extended to approximately 1.01 million tokens using RoPE or YaRN scaling techniques. This is an optional deployment configuration rather than the model's native context limit. It may require framework-specific settings, additional memory, and careful testing. The hosted Model Studio limit should be treated separately from experimental or locally configured extensions.

Qwen3.6-27B supports both thinking and non-thinking modes. Thinking mode is intended for more complex reasoning, coding, and STEM problems. Non-thinking mode can reduce response overhead and latency for routine requests. The practical choice is therefore not simply maximum reasoning versus minimum reasoning: developers can use the more deliberate mode for difficult tasks and the faster mode when a short, direct response is sufficient.

Coding and agent capabilities

Coding is one of Qwen3.6-27B's clearest intended uses. Alibaba's published evaluation information reports scores of 77.2 on SWE-bench Verified, 53.5 on SWE-bench Pro, 59.3 on Terminal-Bench 2.0, and 48.2 on SkillsBench under the listed test settings. These figures are benchmark results reported by the provider and should not be treated as a guarantee for every repository or development environment.

In practical use, the model can help inspect a codebase, explain unfamiliar modules, propose patches, review implementation choices, generate frontend code, and work through multi-step debugging tasks. Its repository-level focus is particularly relevant when a request depends on relationships across multiple files rather than on producing an isolated code snippet.

The model supports function calling and tool use. A surrounding application can expose operations such as searching files, running tests, querying a database, or retrieving documentation. Qwen3.6-27B can generate a structured request for one of those operations, but the model does not independently execute tools or control a computer. The application must validate the request, run the tool, and return the result.

Visual and coding abilities can also be combined. For example, a development assistant could receive a screenshot of a broken interface, inspect the relevant source files through tools, and suggest a correction. The quality of that workflow depends not only on the model but also on the tool definitions, permissions, feedback loop, and safeguards implemented by the developer.

Model Studio pricing and deployment choices

Alibaba Cloud Model Studio lists international pricing of $0.60 per million input tokens and $3.60 per million output tokens for Qwen3.6-27B. The documentation also lists lower regional prices of $0.412564 per million input tokens and $2.475384 per million output tokens for the China Beijing and Singapore scopes. The applicable price depends on the endpoint and region used, so developers should verify the current regional documentation before budgeting a production system.

The hosted endpoint supports function calling, structured outputs, web search, and streaming-compatible usage. Structured outputs can help an application request responses in a defined format, while streaming allows partial output to be delivered as it is generated. These features do not turn the model into an autonomous agent: application code still controls tool execution, validation, and external actions.

For this exact Model Studio endpoint, the supplied documentation lists context caching, batch inference, and fine-tuning as unsupported. That makes the hosted model more suitable for direct inference and interactive agent workflows than for a deployment plan that depends on training customized weights, submitting large offline batches, or reusing cached context to reduce repeated processing costs.

Self-hosting provides more control over data handling, deployment, and runtime configuration, but it transfers infrastructure responsibility to the operator. Compatible frameworks identified for the model include Transformers, vLLM, SGLang, and KTransformers. Hardware requirements, throughput, quantization behavior, and achievable context length depend on the selected implementation and configuration; the supplied research does not establish one universal hardware requirement.

Strengths and limitations that matter in practice

Strengths

  • Strong coding orientation: the model is designed for software engineering, repository reasoning, and coding-agent workflows.
  • Broad visual understanding: it combines text processing with image and video analysis, including document and OCR-oriented tasks.
  • Large native context: the 262,144-token window can accommodate substantial codebases, documents, or conversation history.
  • Open-weight availability: organizations can investigate self-hosted deployment instead of relying exclusively on a managed endpoint.
  • Agent integration: function calling and tool use support applications that connect the model to search, code, data, or business systems.
  • Configurable response behavior: thinking and non-thinking modes provide a way to balance deliberation against latency.

Limitations

  • Text-only generation: it cannot directly produce images, videos, audio, or speech.
  • Resource demands: running a 27B model, especially with a very large context, can require substantial hardware and memory.
  • Hosted feature gaps: the documented Model Studio endpoint does not support fine-tuning, batch inference, or context caching for this model.
  • Long-context extensions need care: the approximately 1.01 million-token figure applies to optional RoPE or YaRN configuration, not the native limit.
  • Deployment differences: local frameworks and the managed API may differ in performance, limits, and available features.
  • Tool execution remains external: function calling describes intended actions; it does not itself provide permission to access files, run commands, or change systems.

When to choose Qwen3.6-27B

Choose Qwen3.6-27B when the application needs a combination of coding, visual understanding, long context, and tool integration. It is a particularly reasonable candidate for repository assistants, code-review systems, visual development tools, document-analysis workflows, video question answering, STEM reasoning, and self-hosted multimodal applications.

Its value is strongest when those capabilities are used together. A text-only coding model may be simpler for ordinary code completion, while a smaller model may be faster and less expensive for classification, extraction, or short routine responses. Conversely, a media-generation model is more appropriate when the required result is an image, video, voice recording, or other non-text artifact.

Use the hosted Model Studio version when managed access, streaming, web search, and structured outputs are more important than fine-tuning or batch processing. Consider the open-weight version when deployment control and local operation justify the infrastructure work. In either case, test the model on representative repositories, documents, visual inputs, and tool workflows rather than relying only on published benchmark scores.

Bottom line

Qwen3.6-27B is a text-generating vision-language model aimed at serious coding and agent workloads. Its defining combination is a 27B open-weight architecture, native image and video understanding, a 262,144-token context window, thinking and non-thinking modes, and support for tools through the hosted API. It is less suitable for direct media generation or managed workflows that require fine-tuning, batch inference, or context caching. For developers who need multimodal understanding and software-engineering ability in one model, it offers a substantial capability set, provided its infrastructure and deployment trade-offs are acceptable.


Answers to Frequently Asked Questions

How can Qwen3.6-27B be deployed and what does it cost?
Developers can self-host the open-weight model with compatible frameworks such as Transformers, vLLM, SGLang, and KTransformers, or use Alibaba Cloud Model Studio. Model Studio lists international pricing of $0.60 per million input tokens and $3.60 per million output tokens, with different regional prices available for China Beijing and Singapore. The hosted endpoint does not list fine-tuning, batch inference, or context caching support.
Is Qwen3.6-27B suitable for coding agents and tool use?
Yes. Qwen3.6-27B is designed for software engineering, repository analysis, debugging, code review, frontend development, and multi-step agent workflows. It supports function calling and tool use, but the surrounding application must execute tools, validate requests, and control permissions.
What is Qwen3.6-27B's context window?
The native context window is 262,144 tokens. Alibaba Cloud Model Studio documents a maximum input length of 260,096 tokens and a maximum output length of 65,536 tokens. Optional RoPE or YaRN configurations may extend local deployments to approximately 1.01 million tokens, but this is not the native limit.
What is Qwen3.6-27B?
Qwen3.6-27B is a 27-billion-parameter, open-weight vision-language model from Alibaba's Qwen team. It accepts text, images, and video as input and generates text, with a focus on coding, repository-level reasoning, visual analysis, and agent workflows.
What modalities does Qwen3.6-27B support?
Qwen3.6-27B supports text, image, and video input and produces text output. It can analyze screenshots, document images, and videos, but it does not natively generate images, video, audio, music, or speech.


Sources 4
Provider

About Qwen