Qwen3.8-Omni

Qwen3.8-Omni-Flash

by Qwen · Current and available

Alibaba Cloud’s Qwen3.8-Omni-Flash analyzes text, images, audio, and video and returns text responses. It is built for long-form multimedia understanding, with a 1-million-token context window, adjustable thinking, spatial-audio input, function calling, web search, streaming, JSON Object mode, and context caching. International pricing is based on tokenized input, cache hits, and output.

Text Reasoning Coding
Qwen3.8-Omni-Flash is designed for applications that need to understand long or mixed-media inputs rather than generate images, audio, or video. It accepts text, images, audio, and video, then produces text responses. The model is available through Alibaba Cloud Model Studio under the API identifier qwen3.8-omni-flash, with a 1-million-token context window and a maximum output of 131,072 tokens.
Outputs

What Qwen3.8-Omni-Flash can produce

Text
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming JSON mode Structured output Prompt caching
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Qwen3.8-Omni
Model type Multimodal
Context window 1M tokens
Maximum output 131K tokens
Release date 2026-09-17
Status Current and available
Knowledge cutoff notes

No authoritative provider-published knowledge cutoff was found for the exact Qwen3.8-Omni-Flash model.

Model notes

Canonical API model ID is qwen3.8-omni-flash. Thinking is enabled by default and supports adjustable reasoning_effort levels from none through max. The model accepts text, images, audio, and video but natively outputs text only. It supports custom function calling, Alibaba Cloud's web_search built-in tool, automatic implicit caching, Responses Session caching, streaming, spatial-audio input, and JSON Object mode. Audio input is tokenized at seven tokens per second. Images are generally tokenized at one token per 32-by-32 pixels, with a minimum of 24 tokens and a default maximum of 1,280 tokens; high-resolution image processing can raise the maximum to 16,384 tokens. Pricing varies by region and modality. The separate qwen3.8-omni-flash-realtime model is intended for real-time audio and video interaction with audio output.

Cost

Model pricing

Input USD 0.15 per 1 million input tokens for International deployment; USD 0.016 per 1 million cache-hit input tokens
Output USD 0.47 per 1 million output tokens for International deployment
Model guide

Qwen3.8-Omni-Flash for Long-Form Audio and Video Understanding

Qwen3.8-Omni-Flash is Alibaba Cloud Model Studio’s non-real-time omni-modal model for analyzing text, images, audio, and video and returning text. Its defining strengths are a 1-million-token context window, adjustable reasoning, spatial-audio support, function calling, web search, streaming, and context caching for multimedia analysis and agent workflows.

What Qwen3.8-Omni-Flash is

Qwen3.8-Omni-Flash is Alibaba Cloud Model Studio’s non-real-time omni-modal model for understanding multimedia content. “Omni-modal” means that one request can contain several types of input, including text, still images, audio, and video. The model then responds with text rather than directly producing a new image, audio track, or video.

This makes it a model for interpretation and reasoning over media. Typical tasks include summarizing a long recording, answering questions about a video, reviewing spoken content, extracting information from images, comparing evidence across different media types, and supporting agents that use external tools after examining an audio-visual input.

The canonical API model identifier is qwen3.8-omni-flash. It is offered through Alibaba Cloud Model Studio using Chat Completions or the Responses API. The model is separate from Qwen3.8-Omni-Flash-Realtime, which is intended for real-time audio and video interaction and is therefore a more appropriate option for low-latency speech-to-speech experiences.

Input and output modalities

Qwen3.8-Omni-Flash accepts four main input types:

  • Text
  • Images
  • Audio
  • Video

Its native output is text. It can describe, summarize, classify, transcribe-oriented content, answer questions about, or reason over the supplied media, but the supplied specifications do not identify native image, video, music, or speech generation for this model.

The model also supports spatial-audio input. Multichannel processing can be enabled for stereo audio or four-channel first-order ambisonics through the relevant API settings. This is useful when the position or direction of sound is relevant to the analysis, although applications still need to account for the additional processing and tokenization of audio input.

Context window and output limits

The advertised context window is 1,000,000 tokens. In practical terms, this gives an application substantially more room for long recordings, extended video, documents, instructions, and tool results than a conventional short-context request.

The exact maximum input length depends on the reasoning mode. The documented limits are 991,808 tokens in non-thinking mode and 983,616 tokens in thinking mode. The maximum output length is 131,072 tokens. These are token limits rather than direct measurements of minutes of audio or hours of video: multimedia is converted into tokens according to the provider’s modality-specific rules, so the usable duration depends on the media and request configuration.

Long context does not make every very large request inexpensive. More audio, video, images, and accompanying text generally mean more billable tokens. A practical implementation should monitor token usage and avoid sending irrelevant portions of a recording or video when a smaller, targeted input would answer the question.

Reasoning and agent features

Thinking is enabled by default and can be adjusted with the reasoning_effort parameter. The documented levels are none, minimal, low, medium, high, xhigh, and max. Lower settings can be useful for straightforward extraction or classification, while higher settings are intended for more involved analysis and planning. The provider does not publish a guarantee that a higher setting will be better for every task, and it can increase response time or token consumption.

Qwen3.8-Omni-Flash supports custom function calling. Function calling allows the model to request an application-defined operation, such as saving extracted findings, querying a database, creating a task, or passing a result to another service. The application, not the model, executes the function and returns the result.

It also supports Alibaba Cloud’s built-in web_search tool. This enables workflows in which the model first examines multimedia, then gathers external information before producing an answer. Web search can make an agent more useful for research or current-information tasks, but search results should still be checked because tool access does not guarantee factual accuracy.

Streaming, JSON, and caching

The model supports streaming responses, allowing an application to receive output progressively instead of waiting for the complete response. It also supports JSON Object mode for responses that need a machine-readable object. JSON Object mode should not automatically be treated as the same thing as unrestricted JSON Schema or fully constrained structured output; developers should verify the exact response-format behavior required by their application.

Qwen3.8-Omni-Flash supports automatic implicit context caching and Responses Session caching. Caching can help when related requests reuse context, such as asking several questions about the same long recording or continuing an agent session. Cache behavior, eligibility, and pricing depend on the API implementation and should be checked against the selected deployment documentation.

Multimedia is tokenized for billing. The supplied Model Studio documentation states that audio input is counted at seven tokens per second. Images are generally counted at one token per 32-by-32 pixels, with a minimum of 24 tokens per image and a default maximum of 1,280 tokens. Higher-resolution image processing can increase the maximum to 16,384 tokens. Video and other content are billed according to the provider’s applicable modality-conversion rules.

Pricing and regional availability

For the International deployment scope, the listed prices are:

Usage typePrice
Input tokensUSD 0.15 per 1 million tokens
Cache-hit input tokensUSD 0.016 per 1 million tokens
Output tokensUSD 0.47 per 1 million tokens

These token prices apply to the text and multimedia content after it has been converted into billable tokens. A long audio or video request can therefore cost more than its visible text content might suggest. Cache-hit pricing can materially reduce the cost of repeated context, but only when the request qualifies for caching under the provider’s rules.

The model is listed for Model Studio deployments in China, Singapore, Hong Kong, Japan, Germany, and the United States. Endpoint, account, and API-key requirements vary by region. Availability and prices may change, so production applications should confirm the current regional pricing page before estimating operating costs.

Main strengths and trade-offs

The clearest strength of Qwen3.8-Omni-Flash is the combination of broad media input and very long context. It is suited to workloads where an application needs to keep a large amount of audio-visual evidence available while reasoning over it. Adjustable thinking, function calling, web search, and caching extend it beyond simple media summarization into tool-using workflows.

Its cost profile is also potentially attractive for large-scale processing because the International input price is lower than the output price and cache-hit input is priced substantially lower than ordinary input. However, the low per-million-token figures should not be confused with low total cost for very long media. Audio and video can generate many billable tokens, and higher reasoning settings may require more computation or output.

There are important limitations. The model is non-real-time and produces text only. It is not the best fit for an interactive voice assistant that must respond with synthesized speech at very low latency; the separate Qwen3.8-Omni-Flash-Realtime model is intended for that role. It is also not a native image, video, music, or speech-generation model. Finally, the large context window does not remove the need for careful input selection, because irrelevant media can increase cost and make results harder to evaluate.

Best use cases

  • Summarizing long meetings, lectures, interviews, broadcasts, or recorded calls
  • Answering detailed questions about video scenes, spoken dialogue, and on-screen information
  • Reviewing multimedia content for production, moderation, quality control, or research
  • Combining audio-visual evidence with documents and instructions in one long-context request
  • Building agents that inspect media and then call business tools or search the web
  • Generating video commentary, production notes, editing suggestions, or structured findings
  • Analyzing spatial audio where stereo or first-order ambisonics information matters

When to choose Qwen3.8-Omni-Flash

Choose Qwen3.8-Omni-Flash when the central problem is understanding substantial amounts of mixed media and returning a text explanation, extraction, or decision. It is especially appropriate when a 1-million-token context, adjustable reasoning, tool use, or repeated access to the same multimedia context justifies using a larger omni-modal model.

Another option may be more appropriate when the job is narrowly focused on text and does not benefit from image, audio, or video input. A smaller or text-specialized model may offer lower total cost or faster responses for simple classification, extraction, or short-answer tasks. Qwen3.8-Omni-Flash-Realtime is the better direction for real-time audio and video interaction with audio output. A dedicated generation model is required when the application must create images, video, music, or speech rather than analyze existing media.

Bottom line

Qwen3.8-Omni-Flash is a text-output model for long-form multimedia understanding. Its verified specification highlights include text, image, audio, and video input; a 1-million-token context window; up to 131,072 output tokens; adjustable thinking; spatial-audio support; function calling; web search; streaming; JSON Object mode; and context caching. Those features make it a strong candidate for audio-visual analysis and agent workflows, provided that the application can tolerate non-real-time text responses and manages the token cost of long media inputs.


Answers to Frequently Asked Questions

What is the difference between Qwen3.8-Omni-Flash and Qwen3.8-Omni-Flash-Realtime?
Qwen3.8-Omni-Flash is a non-real-time model that analyzes multimedia and returns text. Qwen3.8-Omni-Flash-Realtime is intended for low-latency, real-time audio and video interaction, making it more suitable for interactive voice assistants and speech-to-speech applications.
How much does Qwen3.8-Omni-Flash cost?
For the International deployment scope, the listed prices are USD 0.15 per 1 million input tokens, USD 0.016 per 1 million cache-hit input tokens, and USD 0.47 per 1 million output tokens. Multimedia is converted into billable tokens, so long audio and video can cost more than their visible text content suggests.
How large is Qwen3.8-Omni-Flash's context window?
Qwen3.8-Omni-Flash has an advertised context window of 1,000,000 tokens. The documented maximum input is 991,808 tokens in non-thinking mode and 983,616 tokens in thinking mode, with a maximum output of 131,072 tokens. Actual audio and video duration depends on how the media is converted into tokens.
What input and output modalities does Qwen3.8-Omni-Flash support?
The model accepts text, still images, audio, and video in a single request, including stereo and first-order ambisonic audio when spatial processing is enabled. Its native output is text; it does not natively generate images, video, music, or speech.
What is Qwen3.8-Omni-Flash used for?
Qwen3.8-Omni-Flash is designed for understanding and reasoning over text, images, audio, and video. Common uses include summarizing long recordings, answering questions about videos, reviewing multimedia content, extracting information, and building agents that analyze media before using external tools.


Sources 6
Provider

About Qwen