Qwen3.8

Qwen3.8-27B

by Qwen · Active and currently available; open-weight release and Alibaba Cloud Model Studio API model

Qwen3.8-27B is a 27-billion-parameter Apache 2.0 open-weight model that accepts text, images, and video and produces text. It targets coding, visual document analysis, long-context research, reasoning, and tool-using agents, with a 262K native local context and a documented 1M-token hosted context in Alibaba Cloud Model Studio.

Text Reasoning Coding
Qwen3.8-27B is designed for users who need more than a text-only coding model but do not want to rely exclusively on a very large mixture-of-experts system. The model combines text and vision-language understanding with video input, long-context processing, configurable reasoning, and tool use. Its open-weight Apache 2.0 release supports self-hosting, while Alibaba Cloud Model Studio provides a hosted version with a documented context window of up to 1 million tokens.
Outputs

What Qwen3.8-27B can produce

Text
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching
Model profile

Performance characteristics

8/10 Reasoning
9/10 Coding
6/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Qwen3.8
Model type Multimodal
Context window 1M tokens
Maximum output 131K tokens
Release date 2026-08-14
Status Active and currently available; open-weight release and Alibaba Cloud Model Studio API model
Knowledge cutoff notes

No direct authoritative knowledge-cutoff date was found for the exact Qwen3.8-27B model. Web search and retrieval features provide external information at runtime but do not establish or change the underlying model cutoff.

Model notes

Canonical open-weight identifier: Qwen/Qwen3.8-27B. Alibaba Cloud Model Studio API identifier: qwen3.8-27b. The model is a native multimodal dense model with 27 billion parameters and Apache 2.0 licensing. The open-weight release has a native 262K-token context and can be extended to approximately 1M tokens using YaRN; Model Studio documents a 1M-token hosted context window. The hosted model supports hybrid reasoning, function calling, structured outputs, web search, context caching, and streaming. Model Studio marks batch inference and fine-tuning unsupported for this exact hosted model. Reasoning, coding, speed, and cost values are editorial comparative estimates rather than vendor scores.

Cost

Model pricing

Input $0.424 per 1M tokens in China Beijing; $0.50 per 1M tokens in Singapore. Implicit cache input is $0.085 per 1M tokens in Beijing and $0.10 per 1M tokens in Singapore. Explicit cache creation is $0.53 per 1M tokens in Beijing and $0.625 per 1M tokens in Si
Output $1.696 per 1M tokens in China Beijing; $3 per 1M tokens in Singapore.
Model guide

Qwen3.8-27B: Open-Weight Multimodal Model for Coding and Agents

Qwen3.8-27B is Alibaba's 27-billion-parameter dense open-weight model for coding, visual understanding, long-context research, office automation, and tool-using agent workflows. It accepts text, images, and video, produces text, supports configurable reasoning and function calling, and is available for local deployment or through Alibaba Cloud Model Studio.

What is Qwen3.8-27B?

Qwen3.8-27B is a 27-billion-parameter dense model from Alibaba's Qwen team. It is positioned for software engineering, professional work, research, document processing, office automation, and long-horizon agentic tasks. A dense model uses the full model for each request rather than selecting only a subset of expert components, which gives this release a relatively straightforward deployment profile compared with larger mixture-of-experts systems, although it still requires substantial hardware when run locally.

The model is available in two main forms. The open-weight release uses the canonical identifier Qwen/Qwen3.8-27B and is licensed under Apache 2.0. Alibaba Cloud Model Studio offers a hosted API identifier, qwen3.8-27b. The model documentation and repository identify the release date as August 14, 2026.

Within the Qwen3.8 catalog, this model occupies the role of a general-purpose multimodal system focused particularly on coding, visual understanding, reasoning, and tool-connected workflows. It is not an image, video, audio, or speech generation model.

Input and output modalities

Qwen3.8-27B accepts text, images, and video as input and returns text as its native output. This combination supports tasks such as asking questions about screenshots, extracting information from visual documents, reviewing diagrams, analyzing video content, and using images as context for code or technical troubleshooting.

Its text-only output is an important practical distinction. Although Qwen Studio and other products in the wider Qwen ecosystem may offer image or video creation, those provider-level features should not be attributed to this model. Qwen3.8-27B does not natively generate images, video, audio, music, or speech.

The model's visual input can be useful in workflows where a conventional text model would require a separate OCR or media-processing system. For example, a coding assistant could receive a screenshot of an error, a diagram of a system architecture, or a video showing a software defect and then explain the evidence in text.

Context window, reasoning, and output limits

The open-weight model has a native context length of 262,144 tokens. Context is the amount of text and other represented input that the model can consider in one request. The open-weight release can be extended to approximately 1 million tokens with YaRN, but local users must configure the serving framework and extension themselves.

Alibaba Cloud Model Studio documents a 1,000,000-token context window for the hosted model. It also lists a maximum output length of 131,072 tokens and a maximum chain-of-thought length of 262,144 tokens. These hosted limits should not be treated as identical to the native local configuration: the API exposes a million-token context, while the open-weight release natively provides 262K and requires additional configuration for a longer context.

Reasoning is enabled by default in the hosted configuration. Developers can adjust reasoning intensity through provider-specific controls such as reasoning_effort. More reasoning can help with complex coding, planning, research, and multi-step tool use, but it can also increase latency and token consumption. Lower settings may be preferable for routine transformations, short answers, or applications where response time matters more than depth.

No authoritative knowledge-cutoff date is documented for this exact model. Web search and retrieval tools can provide current external information at runtime, but they do not establish or change the underlying training-data cutoff.

Coding and agent workflows

Coding is one of Qwen3.8-27B's clearest intended uses. The model can help inspect repositories, explain implementation choices, diagnose errors from screenshots, draft or modify code, and reason through multi-step software tasks. Its combination of coding ability, long context, visual input, and tool calling is especially relevant to development environments that need to work with source files, documentation, terminals, issue trackers, or test systems.

Tool calling allows an application to connect the model to external functions or services. The model can decide that a tool is needed and produce a structured request, while the surrounding application performs the actual operation and returns the result. This makes it suitable for research assistants, file-processing systems, office automation, and agents that need feedback from an execution environment.

Official materials also list function calling, structured outputs, web search, context caching, and streaming as supported capabilities. Structured outputs can help applications receive responses in a predictable schema, while streaming allows partial text to be delivered before the complete response is finished. These features are useful for interactive software, but they do not turn the model into an autonomous system by themselves: developers still need to implement permissions, tool execution, validation, and error handling.

For local serving, official examples reference Transformers, SGLang, vLLM, and TokenSpeed. The serving examples expose an OpenAI-compatible endpoint and include Qwen-specific reasoning and tool-call parsers. Hardware requirements vary substantially with numerical precision, quantization, context length, batching, and serving configuration, so the 27-billion-parameter size alone is not enough to determine the cost of a deployment.

Pricing and deployment options

Alibaba Cloud Model Studio publishes separate prices for the China Beijing and international Singapore deployments. The standard China Beijing price is $0.424 per million input tokens and $1.696 per million output tokens. The Singapore price is $0.50 per million input tokens and $3 per million output tokens.

DeploymentStandard inputStandard output
China Beijing$0.424 per 1M tokens$1.696 per 1M tokens
Singapore$0.50 per 1M tokens$3 per 1M tokens

Cached-input pricing is lower than standard input pricing. In Beijing, implicit cache input is listed at $0.085 per million tokens, explicit cache creation at $0.53, and explicit cache reads at $0.042. In Singapore, the corresponding prices are $0.10, $0.625, and $0.05 per million tokens. Cache pricing matters most for applications that repeatedly send a large, stable prompt such as a codebase index, policy document, or system instruction.

Model Studio marks batch inference and fine-tuning as unsupported for this exact hosted model. That limitation applies to the documented hosted offering, not necessarily to the open-weight model. The open-weight release can be downloaded, adapted, or self-hosted with compatible tools, although the practical cost and difficulty of training or serving it depend on available hardware and software.

Strengths and limitations

Several strengths distinguish Qwen3.8-27B from a conventional text-only model:

  • Multimodal understanding: it processes text, images, and video in a single model workflow.
  • Long context: the hosted version supports a documented 1-million-token context, while the local release provides a 262K native context.
  • Coding focus: the model is intended for software engineering, repository analysis, and environment-connected development tasks.
  • Configurable reasoning: developers can trade depth and quality against latency and token use.
  • Agent support: function calling, structured outputs, web search, streaming, and caching support tool-connected applications.
  • Deployment flexibility: users can choose a hosted Model Studio API or an Apache 2.0 open-weight release for local serving.

There are also important limitations:

  • No native media generation: the model does not produce images, video, audio, music, or speech.
  • Resource demands: a 27-billion-parameter model with a very long context can require substantial memory and careful serving configuration.
  • Hosted feature restrictions: Model Studio does not list batch inference or fine-tuning for this exact model.
  • Configuration differences: the hosted 1-million-token limit and local 262K native context are different deployment configurations.
  • Reasoning cost: deeper reasoning can increase both response time and token charges.
  • Uncertain current facts without retrieval: no verified knowledge-cutoff date is available, so current information should use web search or another retrieval layer.

Performance and market positioning

The supplied editorial assessment rates Qwen3.8-27B at 9 out of 10 for coding, 8 for reasoning, 8 for cost, and 6 for speed. These are comparative editorial estimates, not provider-published benchmark scores. They indicate the intended trade-off: the model is considered particularly suitable for coding and complex workflows, while its size and reasoning behavior make it less appropriate for the lowest-latency applications.

Compared with a small language model, Qwen3.8-27B offers broader visual input, longer-context processing, and more room for complex reasoning, but it generally requires more compute and may respond more slowly. Compared with a much larger model, its 27-billion-parameter size can make self-hosting more practical, although it may not match the capabilities of the largest systems on every task. The most appropriate choice depends on whether coding depth, multimodal understanding, local control, speed, or operating cost is the primary requirement.

When to choose Qwen3.8-27B

Choose Qwen3.8-27B when you need a model that can combine coding with visual or video understanding and connect to external tools. It is a strong candidate for a private coding assistant, repository analysis, document and screenshot review, long-context research, office automation, and multimodal agents that need to inspect evidence before taking a structured action.

The open-weight version is particularly relevant when deployment control, local processing, Apache 2.0 licensing, or customization is more important than turnkey hosting. The Model Studio version is more convenient when a managed endpoint, documented long context, streaming, caching, and provider-supported API access are preferable.

Another option may be more appropriate when the application needs native image or speech generation, extremely low latency, hosted batch inference, hosted fine-tuning, or minimal hardware requirements. Qwen3.8-27B is best understood as a text-output reasoning and understanding model with multimodal input, not as an all-purpose media-generation system.


Answers to Frequently Asked Questions

What is Qwen3.8-27B designed for?
Qwen3.8-27B is a 27-billion-parameter multimodal model designed for software engineering, coding, visual understanding, document processing, research, office automation, and long-horizon agentic workflows.
What is the context window of Qwen3.8-27B?
The open-weight model has a native context length of 262,144 tokens and can be extended to approximately 1 million tokens with YaRN and suitable serving configuration. Alibaba Cloud Model Studio documents a 1-million-token context window for its hosted version, with a maximum output length of 131,072 tokens.
How can Qwen3.8-27B be deployed and what does it cost?
Users can run the Apache 2.0 open-weight release locally with tools such as Transformers, SGLang, vLLM, or TokenSpeed, or use the hosted Alibaba Cloud Model Studio API. Hosted standard pricing is $0.424 per million input tokens and $1.696 per million output tokens in China Beijing, or $0.50 per million input tokens and $3 per million output tokens in Singapore.
Can Qwen3.8-27B be used for coding agents and tool calling?
Yes. Qwen3.8-27B supports coding workflows, function calling, structured outputs, web search, streaming, context caching, and tool-connected applications. Developers must still implement tool permissions, execution, validation, and error handling.
What input and output modalities does Qwen3.8-27B support?
Qwen3.8-27B accepts text, images, and video as input and produces text as output. It can analyze screenshots, diagrams, visual documents, and video, but it does not natively generate images, video, audio, music, or speech.


Sources 7
Provider

About Qwen