Model catalog

Qwen Models

Browse the AI models associated with Qwen. Compare current and historical models by family, capabilities, context window, availability and intended use.

100 models tracked
100 Total models
44 Model families
11 Model types
99 Current / accessible
All models

Qwen model catalog

Cost-sensitive cross-modal retrieval, image and video search, multimedia catalog indexing, and vector search

Type Multimodal
Context 1K
Reasoning 1/10
Speed 8/10
Multimodal Image input Video input
Status

Current and available through Alibaba Cloud Model Studio International deployment

Input Image/video: $0.03 per 1 million input tokens; text: $0.09 per 1 million input tokens
Output Free; embedding output is not charged
View model →
Qwen logo
Multimodal Embedding

multimodal-embedding-v1

Cross-modal retrieval, text-to-image search, image similarity, video search, semantic classification, clustering, and multimodal vector indexing.

Type Multimodal Embedding
Context 512
Reasoning 0/10
Speed 0/10
Multimodal Image input Video input
Status

Current and accessible through Alibaba Cloud Model Studio in the China (Beijing) region; free trial pricing is listed.

Input Free trial
View model →

Agent-environment simulation, tool-interaction modeling, terminal and software-engineering trajectories, and research on language world models

Type Other
Context 262K
Reasoning 8/10
Speed 7/10
Tool use Streaming
Status

Current open-weight model

View model →

Low-latency duplex voice assistants, real-time customer service, AI companions, and streamed speech-to-speech applications

Type Multimodal
Context 41K
Reasoning 6/10
Speed 8/10
Multimodal Audio input Media output
Status

Current and accessible; standard-edition real-time duplex speech model

Input Singapore: text $0.80/M tokens; audio $6.40/M tokens. China (Beijing): text $0.688/M tokens; audio $5.501/M tokens.
Output Singapore: text $6.40/M tokens; audio $24/M tokens. China (Beijing): text $5.501/M tokens; audio $20.628/M tokens.
View model →

Low-latency voice assistants, customer service, AI companions, full-duplex spoken interaction, and voice applications using tools or cloned voices.

Type Multimodal
Context 262K
Reasoning 4/10
Speed 9/10
Multimodal Audio input Media output
Status

Current and available

Input Singapore: $0.80 per 1M text-input tokens; $6.40 per 1M audio-input tokens. China (Beijing): $0.688 per 1M text-input tokens; $5.501 per 1M audio-input tokens.
Output Singapore: $6.40 per 1M text-output tokens; $24 per 1M audio-output tokens. China (Beijing): $5.501 per 1M text-output tokens; $20.628 per 1M audio-output tokens.
View model →
Qwen logo
Qwen-Audio-3.0-Realtime

qwen-audio-3.0-realtime-flash

Low-latency voice assistants, real-time customer service, duplex speech conversations, interactive voice agents, and applications requiring streaming audio responses.

Type Multimodal
Context 41K
Reasoning 6/10
Speed 9/10
Multimodal Audio input Media output
Status

Available; current API access confirmed, but newer models are recommended for some new projects

Input International/Singapore: text input $0.23 per 1M tokens; audio input $0.93 per 1M tokens. China (Beijing): text input $0.413 per 1M tokens; audio input $4.126 per 1M tokens.
Output International/Singapore: text-only output $0.70 per 1M tokens; combined text and audio output $1.87 per 1M tokens. China (Beijing): text-only output $4.126 per 1M tokens; combined text and audio output $13.752 per 1M tokens.
View model →

Real-time multilingual speech transcription, live captions, meeting transcription, voice interfaces, streaming subtitles, and Chinese-dialect recognition.

Type Speech Recognition
Context 8K
Reasoning 1/10
Speed 9/10
Audio input Streaming
Status

Current and available through Alibaba Cloud Model Studio in the International/Singapore and China (Beijing) regions.

Input International/Singapore: USD 0.93 per 1 million input tokens; China (Beijing): USD 0.848 per 1 million input tokens.
Output International/Singapore: USD 0.70 per 1 million output tokens; China (Beijing): USD 0.636 per 1 million output tokens.
View model →

Expressive text-to-speech, audiobooks, film and video dubbing, content creation, premium voice services, multilingual speech, dialect synthesis, and voice cloning

Type Other
Reasoning 1/10
Speed 8/10
Media output Streaming
Status

Current and available

Input USD 0.20 per 10,000 input characters in Singapore/International; USD 0.19253 per 10,000 input characters in China (Beijing)
Output Not separately billed; pricing is based on input characters
View model →

Long-form offline transcription of meetings, interviews, calls, media files, and multilingual or dialect-rich recordings

Type Other
Context 8K
Reasoning 2/10
Speed 7/10
Audio input
Status

Current and publicly available

Input USD 0.15 per 1 million input tokens in Singapore; USD 0.113 per 1 million input tokens in China (Beijing)
Output USD 0.47 per 1 million output tokens in Singapore; USD 0.382 per 1 million output tokens in China (Beijing)
View model →
Qwen logo
Qwen-Drive

Qwen-Drive-1.0

Autonomous-driving research, driving-scene VQA, 3D BEV perception, trajectory prediction, and embodied-AI experimentation

Type Multimodal
Reasoning 6/10
Speed 5/10
Multimodal Image input Media output
Status

Current open-weight research release

View model →
Qwen logo
Qwen-Image

qwen-image-2.0

Text-to-image generation, image editing, text rendering in images, photorealistic scenes, creative design, and producing multiple image variants.

Type Image Generation
Reasoning 3/10
Speed 8/10
Multimodal Image input Media output
Status

Current; accelerated model; functionally equivalent to qwen-image-2.0-2026-03-03

Input $0 per input image; only generated output images are billed
Output $0.035 per image internationally; $0.028671 per image in China (Beijing) and Singapore
View model →
Qwen logo
Qwen-Image

Qwen-Image-2.1

Local text-to-image generation, multi-reference composition, transparent asset creation, product visuals, and localized image editing

Type Multimodal
Speed 7/10
Multimodal Image input Media output
Status

Current open-weight release

Cost-sensitive text-to-image generation, posters, marketing graphics, illustrations, and images containing Chinese or English text

Type Image Generation
Speed 8/10
Media output
Status

Current and accessible; currently equivalent to qwen-image

Input Not applicable; image generation is billed per output image
Output $0.03/image internationally; $0.028671/image in China (Beijing)
View model →
Qwen logo
Qwen-Image-2.1

Qwen-Image-2.1-PE-I2I

Rewriting vague image-editing instructions, preserving source-image details, coordinating multi-image edits, and preparing prompts for Qwen-Image-2.1.

Type Multimodal
Context 262K
Reasoning 5/10
Speed 5/10
Multimodal Image input Streaming
Status

Current; open-weight research release

Qwen logo
Qwen-Image-2.1

Qwen-Image-2.1-PE-T2I

Expanding short or multilingual image requests into detailed prompts for Qwen-Image-2.1

Type Other
Context 262K
Reasoning 5/10
Speed 5/10
Status

Current open-weight release

Qwen logo
Qwen-Image-Edit

qwen-image-edit

Natural-language single-image editing, bilingual text changes, object insertion or removal, style transfer, pose changes, and image fusion

Type Other
Reasoning 2/10
Speed 6/10
Multimodal Image input Media output
Status

Current and accessible

Input $0.045 per image in Singapore for international deployment
View model →
Qwen logo
Qwen-Image-Edit

qwen-image-edit-max

High-quality image editing, multi-image composition, industrial design concepts, geometric transformations, character-consistent edits, and controlled visual revisions

Type Other
Reasoning 2/10
Speed 6/10
Multimodal Image input Media output
Status

Current and available; canonical model ID is functionally equivalent to qwen-image-edit-max-2026-01-16

Input $0.075 per output image in the International/Singapore region; $0.071677 per output image in China (Beijing)
Output $0.075 per output image in the International/Singapore region; $0.071677 per output image in China (Beijing)
View model →
Qwen logo
Qwen-Image 2.0

qwen-image-2.0-pro

Professional text-to-image generation, image editing, posters, infographics, multilingual in-image text, photorealistic scenes, and reference-based creative production

Type Other
Reasoning 6/10
Speed 5/10
Multimodal Image input Media output
Status

Current rolling model; functionally equivalent to qwen-image-2.0-pro-2026-04-22

Output $0.075 per image internationally; $0.071676 per image in China (Beijing)
View model →
Qwen logo
Qwen-Max

Qwen-Max

Complex multilingual text generation, coding, logical reasoning, creative writing, structured extraction, and enterprise applications

Type General Purpose
Context 33K
Reasoning 7/10
Speed 6/10
Tool use Streaming
Status

Current and accessible rolling-update mainline model

Input US$1.60 per 1 million tokens for International deployment; US$0.345 per 1 million tokens in China mainland
Output US$6.40 per 1 million tokens for International deployment; US$1.377 per 1 million tokens in China mainland
Qwen logo
Qwen-Plus

Qwen-Plus

General-purpose text generation, long-context analysis, multilingual applications, structured business workflows, function-calling agents, and applications that need optional reasoning mode.

Type General Purpose
Context 1M
Reasoning 7/10
Speed 7/10
Tool use Web search Streaming
Status

Current; qwen-plus currently resolves to the qwen-plus-2025-12-01 snapshot

Input $0.40 per 1M input tokens for 0–256K input; $1.20 per 1M input tokens above 256K in the International Singapore pricing tier. International Global pricing is $0.115 per 1M tokens for up to 128K, $0.345 for 128K–256K, and $0.689 for 256K–1M.
Output $1.20 per 1M non-thinking output tokens and $4.00 per 1M thinking output tokens for 0–256K input; $3.60 non-thinking and $12.00 thinking per 1M output tokens above 256K in the International Singapore pricing tier. International Global pricing differs by r
View model →
Qwen logo
Qwen-Turbo

Qwen-Turbo

High-volume customer support, simple to moderate question answering, summarization, rewriting, structured text extraction, and cost-sensitive applications

Type Lightweight
Context 131K
Reasoning 6/10
Speed 9/10
Web search Streaming
Status

Currently accessible; no longer being updated; Alibaba Cloud recommends switching to Qwen-Flash

Input $0.05 per 1M tokens international non-thinking; $0.05 per 1M tokens international thinking; $0.044 per 1M tokens China (Beijing)
Output $0.20 per 1M tokens international non-thinking; $0.50 per 1M tokens international thinking; $0.087 per 1M tokens China non-thinking; $0.431 per 1M tokens China thinking
Qwen logo
Qwen-VL

Qwen-VL-Flash

Fast image and video understanding, visual question answering, document analysis, and multimodal extraction

Type Multimodal
Context 33K
Reasoning 6/10
Speed 9/10
Multimodal Image input Video input
Status

Active

Complex image and video understanding, document analysis, chart interpretation, visual question answering, and structured extraction

Type Multimodal
Context 131K
Reasoning 8/10
Speed 6/10
Multimodal Image input Video input
Status

Currently accessible legacy visual language model; current qwen-vl-max endpoint is functionally equivalent to qwen-vl-max-2025-08-13

Input $0.229 per 1M tokens in China (Beijing) and Singapore; $0.80 per 1M tokens for international deployment
Output $0.573 per 1M tokens in China (Beijing) and Singapore; $3.20 per 1M tokens for international deployment
View model →
Qwen logo
Qwen-VL

Qwen-VL-Plus

High-resolution image and video understanding, OCR-style text recognition, document analysis, visual question answering, and multimodal assistants

Type Multimodal
Context 131K
Reasoning 6/10
Speed 7/10
Multimodal Image input Video input
Status

Current rolling model; functionally equivalent to qwen-vl-plus-2025-08-15

Input $0.115 per 1M tokens in China (Beijing); $0.21 per 1M tokens in Singapore
Output $0.287 per 1M tokens in China (Beijing); $0.63 per 1M tokens in Singapore

Self-hosted chat assistants, multilingual text generation, coding and mathematics assistance, long-context document work, structured text generation, and cost-sensitive private deployments

Type General Purpose
Context 131K
Reasoning 7/10
Speed 7/10
Tool use Fine-tuning Streaming
Status

Current open-weight model; publicly available for download and self-hosted deployment

View model →

Self-hosted multilingual assistants, document processing, RAG, coding support, structured extraction, and cost-conscious production deployments.

Type General Purpose
Context 131K
Reasoning 7/10
Speed 7/10
Tool use Fine-tuning Streaming
Status

Legacy open-weight model; downloadable and self-hostable, while Qwen2.5 API models marked deprecated are no longer callable through current Alibaba Cloud Model Studio pricing documentation.

Input Not applicable for the downloadable checkpoint; no current official per-token hosted price verified for this exact model.
Output Not applicable for the downloadable checkpoint; no current official per-token hosted price verified for this exact model.
View model →

Self-hosted assistants, multilingual text generation, long-document processing, coding support, structured extraction, RAG, and agent applications

Type General Purpose
Context 131K
Reasoning 7/10
Speed 5/10
Tool use Fine-tuning Streaming
Status

Available; open-weight model

View model →

Self-hosted multilingual assistants, coding, mathematics, document analysis, structured text generation, and long-context workloads

Type General Purpose
Context 131K
Reasoning 8/10
Speed 4/10
Fine-tuning Streaming
Status

Available as an open-weight model; legacy relative to newer Qwen generations

View model →
Qwen logo
Qwen2.5-Omni

Qwen2.5-Omni-7B

Multimodal assistants, audio and video understanding, visual question answering, voice interaction, speech instruction following, and local multimodal AI research

Type Multimodal
Context 33K
Reasoning 7/10
Speed 7/10
Multimodal Image input Audio input
Status

Current; open-weight model and available through Alibaba Cloud Model Studio

Input International: $0.10 per 1M text tokens; $6.76 per 1M audio tokens; $0.28 per 1M image/video tokens. China Beijing: $0.087 per 1M text tokens; $5.448 per 1M audio tokens; $0.287 per 1M image/video tokens.
Output International: $0.40 per 1M tokens for text-only input; $0.84 per 1M tokens for text after multimodal input; $13.51 per 1M tokens for text-and-audio output. China Beijing: $0.345 per 1M tokens for text-only input; $0.861 per 1M tokens for text after multi
View model →

Local or self-hosted image and video understanding, OCR, document extraction, chart and diagram analysis, visual question answering, visual grounding, and multimodal research

Type Multimodal
Context 33K
Reasoning 7/10
Speed 7/10
Multimodal Image input Video input
Status

Open-weight and downloadable; supported as a fine-tuning base model in Alibaba Cloud Model Studio; current hosted inference availability and standard pricing for this exact model are not clearly listed in the latest Model Studio inference-pricing catalog.

View model →

High-quality image, document, chart, screenshot, OCR, visual-grounding, and video analysis; multimodal agents and self-hosted experimentation

Type Multimodal
Context 131K
Reasoning 8/10
Speed 3/10
Multimodal Image input Video input
Status

Available open-weight model; older Qwen2.5-VL generation and still listed by Alibaba Cloud Model Studio

View model →

Fast, high-volume text generation; long-context analysis; summarization; extraction; structured outputs; and applications needing optional reasoning.

Type Lightweight
Context 1M
Reasoning 7/10
Speed 9/10
Tool use Web search Streaming
Status

Current and available; canonical qwen-flash identifier is functionally equivalent to qwen-flash-2025-07-28

Input Tiered per 1M input tokens. China Beijing: CNY 0.15 up to 128K input, CNY 0.60 above 128K to 256K, CNY 1.20 above 256K to 1M. International Singapore: CNY 0.367 up to 256K, CNY 1.835 above 256K to 1M. Regional pricing and promotional offers may vary.
Output Tiered per 1M output tokens. China Beijing: CNY 1.50 up to 128K input, CNY 6 above 128K to 256K, CNY 12 above 256K to 1M. International Singapore: CNY 2.936 up to 256K, CNY 14.678 above 256K to 1M. Regional pricing and promotional offers may vary.
View model →

Lightweight local assistants, offline prototypes, embedded experimentation, education, simple text generation, and resource-constrained deployments

Type Lightweight
Context 33K
Reasoning 3/10
Speed 9/10
Tool use Fine-tuning Streaming
Status

Current open-weight model

View model →

Efficient local inference, edge applications, multilingual chat, lightweight reasoning, coding assistance, and tool-enabled agents

Type Lightweight
Context 33K
Reasoning 7/10
Speed 9/10
Tool use Fine-tuning Streaming
Status

Current open-weight model

Input No official hosted API price; open-weight model
Output No official hosted API price; open-weight model
View model →

Local and self-hosted chat, compact reasoning, coding assistance, multilingual applications, retrieval-augmented generation, and lightweight tool-using agents

Type General Purpose
Context 33K
Reasoning 7/10
Speed 8/10
Tool use Fine-tuning Streaming
Status

Current open-weight model

View model →

Cost-efficient reasoning, coding, multilingual assistants, local deployment, structured text generation, tool-enabled agents, and fine-tuned applications

Type General Purpose
Context 131K
Reasoning 8/10
Speed 8/10
Tool use Fine-tuning Streaming
Status

Current and accessible; open-weight model available for self-hosting and Alibaba Cloud Model Studio API deployment

Input $0.072 per 1 million tokens for Global deployment; $0.18 per 1 million tokens for International deployment
Output $0.287 per 1 million tokens for non-thinking mode and $0.717 per 1 million tokens for thinking mode in Global deployment; International pricing is $0.70 non-thinking and $2.10 thinking per 1 million tokens
View model →

Local deployment, multilingual assistants, reasoning, mathematics, coding, structured text generation, research, and cost-sensitive agent workflows.

Type General Purpose
Context 131K
Reasoning 7/10
Speed 8/10
Tool use Fine-tuning Streaming
Status

Current open-weight model; also available as qwen3-14b through Alibaba Cloud Model Studio, with regional capability and pricing differences.

Input Alibaba Cloud Model Studio: China (Beijing) $0.144 per 1 million input tokens; Singapore $0.35 per 1 million input tokens. Regional prices and deployment terms may vary.
Output Alibaba Cloud Model Studio: China (Beijing) $0.574 per 1 million output tokens in non-thinking mode and $1.434 per 1 million output tokens in thinking mode; Singapore $1.4 per 1 million output tokens in non-thinking mode and $4.2 per 1 million output toke
View model →

Self-hosted assistants, coding, mathematical and logical reasoning, multilingual applications, long-context document processing, and agentic tool-use systems

Type Reasoning
Context 131K
Reasoning 8/10
Speed 9/10
Tool use Streaming
Status

Current and accessible; open-weight release with hosted Alibaba Cloud Model Studio availability

Input US$0.108 per 1 million input tokens in several global deployments; US$0.20 per 1 million input tokens for the international Singapore deployment
Output US$0.431 per 1 million output tokens in several global deployments; US$0.80 per 1 million output tokens for the international Singapore deployment. Thinking output is separately priced at higher rates where exposed.
View model →

Self-hosted reasoning assistants, coding agents, mathematics, multilingual applications, structured text generation, and tool-calling workflows

Type Reasoning
Context 256K
Reasoning 9/10
Speed 6/10
Tool use Fine-tuning Streaming
Status

Current; open-weight model and available through Alibaba Cloud Model Studio

Input $0.16 per 1 million tokens in international Singapore, Germany Frankfurt, and US Virginia deployments; $0.287 per 1 million tokens in China Beijing
Output $0.64 per 1 million tokens in international Singapore, Germany Frankfurt, and US Virginia deployments; China Beijing: $1.147 per 1 million non-thinking tokens or $2.868 per 1 million thinking tokens
View model →

Complex reasoning, mathematics, software development, multilingual applications, function calling, agentic workflows, research and self-hosted open-weight deployment

Type Reasoning
Context 131K
Reasoning 9/10
Speed 6/10
Tool use Fine-tuning Streaming
Status

Current and accessible through Alibaba Cloud Model Studio; original Qwen3 open-weight release, with newer 2507 instruct and thinking variants available separately

Input $0.287 per 1M tokens for standard input in the United States, Germany and China; $0.700 per 1M tokens in Singapore. Thinking-mode input is priced the same.
Output $1.147 per 1M tokens for standard output in the United States, Germany and China; $2.868 per 1M tokens for thinking-mode output. Singapore pricing is $2.800 standard output and $8.400 thinking-mode output per 1M tokens.
View model →
Qwen logo
Qwen3 Reranker

qwen3-rerank

Multilingual semantic search, RAG candidate reranking, enterprise document retrieval, knowledge-base search, and improving search-result relevance.

Type Other
Context 4K
Reasoning 2/10
Speed 8/10
Status

Current and available

Input $0.10 per 1 million input tokens for the international Singapore deployment; pricing is region-dependent. Output is free.
Output Free
View model →

Low-cost multilingual speech-to-text, language identification, offline transcription, real-time streaming ASR, and high-throughput deployments

Type Other
Context 66K
Reasoning 2/10
Speed 8/10
Multimodal Audio input Streaming
Status

Current, open-weight, Apache 2.0 licensed

View model →

Multilingual speech transcription, language identification, long-audio processing, and self-hosted or streaming ASR applications

Type Other
Reasoning 2/10
Speed 8/10
Multimodal Audio input Fine-tuning
Status

Current; open-weight and downloadable

Input No official hosted API price identified for this exact open-weight checkpoint
Output No separate output charge for the self-hosted checkpoint
View model →
Qwen logo
Qwen3-ASR

Qwen3-ForcedAligner-0.6B

Word- and character-level speech timestamp alignment, subtitle synchronization, transcript timing, and audio annotation

Type Other
Reasoning 1/10
Speed 8/10
Multimodal Audio input
Status

Current open-weight model

Qwen logo
Qwen3-Coder

Qwen3-Coder-Next

Repository-scale coding agents, code generation, code completion, debugging, refactoring, terminal workflows, and cost-sensitive self-hosted deployments.

Type Coding
Context 262K
Reasoning 8/10
Speed 8/10
Tool use Streaming
Status

Current; open-weight model and available through Alibaba Cloud Model Studio

Input USD 0.144 per 1M tokens for China Beijing input up to 32K; USD 0.216 for 32K–128K; USD 0.359 for 128K–256K. International Singapore and Frankfurt pricing is USD 0.30, USD 0.50, and USD 0.80 per 1M input tokens across the same bands.
Output USD 0.574 per 1M tokens for China Beijing output up to 32K; USD 0.861 for 32K–128K; USD 1.434 for 128K–256K. International Singapore and Frankfurt pricing is USD 1.50, USD 2.50, and USD 4.00 per 1M output tokens across the same bands.
View model →
Qwen logo
Qwen3-Coder

Qwen3-Coder-Plus

Large-codebase analysis, code generation, refactoring, debugging, documentation and long-context coding-agent workflows

Type Coding
Context 1M
Reasoning 7/10
Speed 7/10
Tool use Streaming
Status

Current; canonical qwen3-coder-plus currently equivalent to qwen3-coder-plus-2025-09-23

Input US Virginia: $0.574 per 1M tokens up to 32K input; $0.861 for 32K-128K; $1.434 for 128K-256K; $2.868 for 256K-1M. Regional pricing varies.
Output US Virginia: $2.294 per 1M tokens up to 32K input; $3.441 for 32K-128K; $5.735 for 128K-256K; $28.671 for 256K-1M. Regional pricing varies.
View model →
Qwen logo
Qwen3-Embedding

text-embedding-v3

Multilingual text embeddings, semantic search, RAG, recommendation, clustering, classification, and migration of existing v3 vector indexes

Type Embedding
Context 8K
Speed 8/10
Status

Available; retained primarily for compatibility with existing v3 embedding indexes

Input $0.07 per 1 million input tokens in the Singapore international deployment
Output Free
Qwen logo
Qwen3-Embedding

text-embedding-v4

Multilingual semantic search, RAG pipelines, vector databases, document retrieval, clustering, classification, recommendation, and code retrieval

Type Other
Context 8K
Reasoning 1/10
Speed 8/10
Status

Current and accessible through Alibaba Cloud Model Studio

Input $0.07 per 1M input tokens in Singapore/International and Hong Kong; $0.072 per 1M input tokens in China (Beijing). Beijing batch-file processing is listed at $0.036 per 1M tokens.
Output Free; pricing is based on input tokens
Qwen logo
Qwen3-LiveTranslate

Qwen3-LiveTranslate-Flash

Streaming translation of recorded or uploaded audio and video, multilingual subtitles, translated voice tracks, and applications requiring translated text or synthesized speech.

Type Other
Context 53K
Reasoning 2/10
Speed 8/10
Multimodal Audio input Video input
Status

Current stable model

Input Audio input and output are billed at 12.5 tokens per second, with audio shorter than one second billed as one second. Video usage additionally consumes video tokens based on sampled frames and resolution. Monetary rates depend on the applicable Alibaba Cl
Output Audio output is billed at 12.5 tokens per second. Text output uses the applicable Model Studio token rate when charged separately. No standalone fixed monetary rate was verified for this exact model in the model documentation.
View model →

Real-time multilingual speech interpretation, live voice translation, conference translation, streaming media, and audiovisual translation with text or synthesized speech output

Type Other
Context 53K
Reasoning 1/10
Speed 8/10
Multimodal Image input Audio input
Status

Legacy; still available; no longer recommended for new use

Input China (Beijing): CNY 64 per 1M input audio tokens and CNY 8 per 1M input image tokens. Singapore: CNY 73.392 per 1M input audio tokens and CNY 9.541 per 1M input image tokens.
Output China (Beijing): CNY 64 per 1M output text tokens and CNY 240 per 1M output audio tokens. Singapore: CNY 73.392 per 1M output text tokens and CNY 278.891 per 1M output audio tokens.
View model →
Qwen logo
Qwen3-Max

Qwen3-Max

Complex reasoning, coding assistance, web-grounded agents, function calling, structured extraction, and long-context text analysis

Type Reasoning
Context 262K
Reasoning 9/10
Speed 7/10
Tool use Web search Streaming
Status

Current

Input International: $1.20 per 1M input tokens up to 32K; $2.40 per 1M above 32K to 128K; $3.00 per 1M above 128K to 256K. Regional prices vary.
Output International: $6.00 per 1M output tokens up to 32K; $12.00 per 1M above 32K to 128K; $15.00 per 1M above 128K to 256K. Regional prices vary.
View model →

Local visual assistants, OCR, document and chart analysis, image question answering, lightweight video understanding, and multimodal prototyping

Type Multimodal
Context 256K
Reasoning 6/10
Speed 8/10
Multimodal Image input Video input
Status

Current open-weight model

View model →

Local image and video understanding, OCR, document extraction, visual question answering, visual coding, and lightweight multimodal agents

Type Multimodal
Context 262K
Reasoning 6/10
Speed 8/10
Multimodal Image input Video input
Status

Current open-weight model; available on Hugging Face and supported for supervised fine-tuning in Alibaba Cloud Model Studio

View model →

Local or hosted image and video understanding, OCR, document extraction, visual question answering, spatial reasoning, screenshot analysis, multimodal agents, and structured data extraction

Type Multimodal
Context 262K
Reasoning 7/10
Speed 8/10
Multimodal Image input Video input
Status

Current; open-weight model with hosted inference available through Alibaba Cloud Model Studio

Input $0.072 per 1 million input tokens
Output $0.287 per 1 million output tokens
View model →

Image and video understanding, OCR, document analysis, spatial reasoning, visual coding, long-context multimodal tasks, and visual-agent applications

Type Multimodal
Context 256K
Reasoning 8/10
Speed 8/10
Multimodal Image input Video input
Status

Current and available; open-weight release and Alibaba Cloud Model Studio API model

Input $0.20 per 1 million tokens for the listed international deployment; regional pricing may differ
Output $0.80 per 1 million tokens for the listed international deployment; regional pricing may differ
View model →

Document intelligence, OCR, image and video understanding, spatial reasoning, visual coding, and visual-agent applications

Type Multimodal
Context 131K
Reasoning 8/10
Speed 6/10
Multimodal Image input Video input
Status

Current; open-weight checkpoint and available through Alibaba Cloud Model Studio managed inference

Input $0.16 per 1 million input tokens in Alibaba Cloud Model Studio US (Virginia) global deployment; China Beijing pricing is $0.287 per 1 million input tokens
Output $0.64 per 1 million output tokens in Alibaba Cloud Model Studio US (Virginia) global deployment; China Beijing pricing is $1.147 per 1 million output tokens
View model →

High-quality image and video understanding, OCR, document intelligence, visual coding, spatial reasoning, long-context multimodal analysis and visual-agent applications

Type Multimodal
Context 131K
Reasoning 8/10
Speed 5/10
Multimodal Image input Video input
Status

Current and accessible; open-weight release and Alibaba Cloud Model Studio API availability

Input $0.287 per 1 million tokens in the US Virginia global deployment; $0.400 per 1 million tokens in Singapore
Output $1.147 per 1 million tokens in the US Virginia global deployment; $1.600 per 1 million tokens in Singapore
View model →
Qwen logo
Qwen3-VL

qwen3-vl-rerank

Multimodal reranking, cross-modal search, image retrieval, video retrieval, image clustering, and multimodal RAG

Type Multimodal
Context 120K
Reasoning 2/10
Speed 7/10
Multimodal Image input Video input
Status

Current and available through Alibaba Cloud Model Studio in the China (Beijing) region

Input Text input: $0.10 per 1 million tokens; image input: $0.258 per 1 million tokens; China (Beijing)
Output Free or not separately charged for reranking output
Qwen logo
Qwen3-VL-Embedding

qwen3-vl-embedding

Multimodal vector search, cross-modal retrieval, image and video search, semantic clustering, tagging, and retrieval pipelines combining text with visual content.

Type Multimodal
Context 32K
Reasoning 4/10
Speed 7/10
Multimodal Image input Video input
Status

Current and available through Alibaba Cloud Model Studio; open-weight 2B and 8B variants are also available through Qwen's official model repositories.

Input Text: $0.10 per 1 million tokens; image/video: $0.258 per 1 million tokens in China (Beijing).
Output No separate output-token price documented; the model returns vector embeddings.

Long-context multimodal analysis, document understanding, video and image interpretation, general reasoning, coding, tool-enabled assistants, and self-hosted deployment

Type Multimodal
Context 262K
Reasoning 8/10
Speed 8/10
Multimodal Image input Video input
Status

Accessible; no longer recommended for new projects

Input $0.086 per 1M input tokens for requests up to 128K input tokens in US Virginia; $0.258 per 1M input tokens for 128K-256K requests
Output $0.688 per 1M output tokens for requests up to 128K input tokens in US Virginia; $2.064 per 1M output tokens for 128K-256K requests
View model →

Efficient multimodal assistants, coding, reasoning, long-context analysis, local deployment, and tool-using agents

Type Multimodal
Context 262K
Reasoning 8/10
Speed 9/10
Multimodal Image input Video input
Status

Available; open-weight Apache 2.0 model with hosted API access through Alibaba Cloud Model Studio

Input USD 0.057 per 1 million input tokens for up to 128K input; USD 0.229 per 1 million input tokens for 128K–256K input in the Global deployment scope. International flat pricing is USD 0.25 per 1 million input tokens.
Output USD 0.459 per 1 million output tokens for up to 128K input; USD 1.835 per 1 million output tokens for 128K–256K input in the Global deployment scope. International flat pricing is USD 2 per 1 million output tokens.
View model →

Advanced multimodal reasoning, image and video understanding, coding, long-context analysis, document and chart interpretation, function-calling agents, and web-grounded workflows.

Type Multimodal
Context 262K
Reasoning 9/10
Speed 8/10
Multimodal Image input Video input
Status

Current and available through Alibaba Cloud Model Studio; released globally on February 24, 2026.

Input $0.115 per 1M tokens for input up to 128K; $0.287 per 1M tokens for input from 128K to 256K. Pricing varies by deployment scope and region.
Output $0.917 per 1M tokens for requests up to 128K input; $2.294 per 1M tokens for requests from 128K to 256K input. Pricing varies by deployment scope and region.
View model →

Advanced multimodal reasoning, coding, video and image understanding, long-context analysis, tool-using agents, and self-hosted open-weight deployments

Type Multimodal
Context 262K
Reasoning 9/10
Speed 6/10
Multimodal Image input Video input
Status

Current; open-weight model and available through Alibaba Cloud Model Studio

Input $0.172 per 1M tokens for input up to 128K; $0.43 per 1M tokens for input above 128K and up to 256K in Beijing, Frankfurt, and Virginia. Singapore: $0.60 per 1M tokens.
Output $1.032 per 1M tokens for input up to 128K; $2.58 per 1M tokens for input above 128K and up to 256K in Beijing, Frankfurt, and Virginia. Singapore: $3.60 per 1M tokens.
View model →

Fast long-context text, image and video understanding; structured extraction; tool-enabled agents; web-grounded applications; and high-volume multimodal workloads.

Type Multimodal
Context 1M
Reasoning 8/10
Speed 9/10
Multimodal Image input Video input
Status

Current and available

Input $0.029 per 1M tokens for 0–128K input; $0.115 per 1M tokens for 128K–256K input; $0.172 per 1M tokens for 256K–1M input in US Virginia/Global pricing
Output $0.287 per 1M tokens for 0–128K input; $1.147 per 1M tokens for 128K–256K input; $1.72 per 1M tokens for 256K–1M input in US Virginia/Global pricing
View model →

Long-context reasoning, multimodal document and video analysis, coding, structured enterprise automation, function-calling agents, and web-grounded research.

Type Multimodal
Context 1M
Reasoning 8/10
Speed 8/10
Multimodal Image input Video input
Status

Current; canonical rolling model identifier currently equivalent to qwen3.5-plus-2026-02-15

Input International: $0.40 per 1M input tokens for input up to 256K; $0.50 per 1M input tokens for input above 256K and up to 1M. Regional and global pricing varies.
Output International: $2.40 per 1M output tokens for input up to 256K; $3.00 per 1M output tokens for input above 256K and up to 1M. Regional and global pricing varies.
View model →
Qwen logo
Qwen3.5-OCR

Qwen3.5-OCR

OCR, document parsing, text localization, table extraction, handwritten-text recognition, and key information extraction from images

Type Multimodal
Context 66K
Reasoning 4/10
Speed 7/10
Image input
Status

Current and available through Alibaba Cloud Model Studio

Input 0.069 USD per 1 million tokens in China (Beijing)
Output 0.275 USD per 1 million tokens in China (Beijing)

Fast multimodal analysis, long audio understanding, audiovisual question answering, voice assistants and spoken-response applications

Type Multimodal
Context 262K
Reasoning 7/10
Speed 9/10
Multimodal Image input Audio input
Status

Current; canonical model ID functionally equivalent to qwen3.5-omni-flash-2026-03-15

Input Singapore international: $0.40 per 1M tokens for text/image/video input and $3.00 per 1M tokens for audio input; China Beijing: $0.30 per 1M tokens for text/image/video input and $2.48 per 1M tokens for audio input
Output Singapore international: $2.20 per 1M tokens for text output and $11.90 per 1M tokens for text-and-audio output; China Beijing: $1.83 per 1M tokens for text output and $9.90 per 1M tokens for text-and-audio output
View model →

Low-latency voice assistants, speech-to-speech applications, realtime multimedia analysis, interactive agents, and multimodal conversations.

Type Multimodal
Context 262K
Reasoning 6/10
Speed 9/10
Multimodal Image input Audio input
Status

Current and available; rolling model identity functionally equivalent to qwen3.5-omni-flash-realtime-2026-03-15

Input International: $0.55 per 1M tokens for text/image/video input; $4.50 per 1M tokens for audio input. China mainland: $0.45 per 1M tokens for text/image/video input; $3.71 per 1M tokens for audio input.
Output International: $3.30 per 1M tokens for text output; $17.70 per 1M tokens for audio output. China mainland: $2.75 per 1M tokens for text output; $14.71 per 1M tokens for audio output.
View model →
Qwen logo
Qwen3.5-Omni

Qwen3.5-Omni-Plus

Multilingual voice assistants, speech-enabled multimodal applications, audio-visual analysis, spoken explanations, accessibility tools, and interactive media workflows.

Type Multimodal
Context 262K
Reasoning 7/10
Speed 7/10
Multimodal Image input Audio input
Status

current

Input $1.40 per 1M tokens for text/image/video input; $11.00 per 1M tokens for audio input in international deployments. China (Beijing) snapshot pricing is $0.96 per 1M tokens for text/image/video input and $7.29 per 1M tokens for audio input.
Output $8.30 per 1M tokens for text output; $44.00 per 1M tokens for text-and-audio output in international deployments. China (Beijing) snapshot pricing is $5.50 per 1M tokens for text output and $29.29 per 1M tokens for text-and-audio output.
View model →

Real-time voice assistants, speech-to-speech applications, multimodal customer service, visual conversational agents, live multimedia analysis, and interactive applications requiring controllable speech output.

Type Multimodal
Context 262K
Reasoning 6/10
Speed 9/10
Multimodal Image input Audio input
Status

Current and available; canonical model identifier functionally equivalent to snapshot qwen3.5-omni-plus-realtime-2026-03-15

Input China (Beijing): $1.38 per 1M text/image/video tokens; $11 per 1M audio tokens. Singapore: $2.10 per 1M text/image/video tokens; $16.50 per 1M audio tokens.
Output China (Beijing): $8.25 per 1M text-output tokens; $41.26 per 1M text-and-audio-output tokens. Singapore: $12.40 per 1M text-output tokens; $62 per 1M text-and-audio-output tokens.
View model →

Coding agents, repository-level software engineering, visual document analysis, video understanding, STEM reasoning, long-context assistants, and self-hosted multimodal applications

Type Multimodal
Context 262K
Reasoning 8/10
Speed 7/10
Multimodal Image input Video input
Status

Current and available; open-weight release and Alibaba Cloud Model Studio API model

Input $0.60 per 1M tokens internationally; $0.412564 per 1M tokens in China Beijing and Singapore
Output $3.60 per 1M tokens internationally; $2.475384 per 1M tokens in China Beijing and Singapore
View model →

Agentic coding, long-context software engineering, multimodal analysis, tool-using agents, and self-hosted deployments

Type Multimodal
Context 262K
Reasoning 8/10
Speed 8/10
Multimodal Image input Video input
Status

Current; open-weight and available through Alibaba Cloud Model Studio

Input USD 0.248 per 1M tokens in US Virginia, Germany Frankfurt, and China Beijing; USD 0.375 per 1M tokens in Singapore
Output USD 1.485 per 1M tokens in US Virginia, Germany Frankfurt, and China Beijing; USD 2.25 per 1M tokens in Singapore
View model →

Fast multimodal assistants, coding agents, visual document analysis, video understanding, tool-using workflows, object localization, and large-context applications

Type Multimodal
Context 1M
Reasoning 8/10
Speed 9/10
Multimodal Image input Video input
Status

Current

Input $0.25 per 1M input tokens up to 256K input tokens; $1.00 per 1M input tokens above 256K and up to 1M, Singapore international pricing
Output $1.50 per 1M output tokens up to 256K input tokens; $4.00 per 1M output tokens above 256K and up to 1M, Singapore international pricing
View model →

Advanced coding agents, front-end development, long-context analysis, structured API workflows, and text generation with web search

Type General Purpose
Context 262K
Reasoning 8/10
Speed 7/10
Tool use Web search Streaming
Status

Preview; currently accessible; scheduled for deprecation on October 10, 2026

Input $1.30 per 1M tokens for input up to 128K; $2.00 per 1M tokens for input above 128K and up to 256K, international pricing
Output $7.80 per 1M tokens for requests up to 128K input; $12.00 per 1M tokens for requests above 128K and up to 256K input, international pricing
View model →

Long-context multimodal analysis, agentic coding, OCR, object localization, frontend development, visual reasoning, and tool-enabled enterprise assistants

Type Multimodal
Context 1M
Reasoning 9/10
Speed 7/10
Multimodal Image input Video input
Status

Current; rolling model ID qwen3.6-plus is currently equivalent to qwen3.6-plus-2026-04-02

Input $0.276 per 1M tokens for 0-256K input tokens; $1.101 per 1M tokens for 256K-1M input tokens in Global deployment
Output $1.651 per 1M tokens for 0-256K input tokens; $6.602 per 1M tokens for 256K-1M input tokens in Global deployment
View model →

Fast multimodal agents, visual coding, tool-use workflows, search agents, and long-context document or screen analysis

Type Multimodal
Context 1M
Reasoning 8/10
Speed 9/10
Multimodal Image input Video input
Status

Current and available; rolling qwen3.7-flash identifier is currently equivalent to qwen3.7-flash-2026-07-15

Input Global: $0.028 per 1M input tokens for 0-32K; $0.083 for 32K-256K; $0.165 for 256K-1M. Singapore international: $0.030, $0.100, and $0.200 per 1M input tokens respectively.
Output Global: $0.110 per 1M output tokens for 0-32K; $0.330 for 32K-256K; $0.660 for 256K-1M. Singapore international: $0.130, $0.400, and $0.800 per 1M output tokens respectively.
View model →
Qwen logo
Qwen3.7

Qwen3.7-Max

Complex reasoning, advanced coding, long-context analysis, tool-using agents, productivity automation, and long-horizon task execution

Type Reasoning
Context 1M
Reasoning 9/10
Speed 7/10
Tool use Web search
Status

Legacy; currently accessible in supported Alibaba Cloud Model Studio deployments

Input $1.65 per 1M tokens in the US Virginia Global scope; $2.50 per 1M tokens in the US scope
Output $4.951 per 1M tokens in the US Virginia Global scope; $7.50 per 1M tokens in the US scope

Long-context reasoning, multimodal document and video analysis, coding, tool-using agents, structured extraction, and enterprise productivity workflows

Type Multimodal
Context 1M
Reasoning 9/10
Speed 7/10
Multimodal Image input Video input
Status

Current; rolling identifier currently equivalent to qwen3.7-plus-2026-05-26

Input $0.40 per 1M tokens up to 256K input tokens and $1.20 per 1M tokens from 256K to 1M input tokens in US Virginia; regional and promotional pricing varies
Output $1.60 per 1M tokens up to 256K input tokens and $4.80 per 1M tokens from 256K to 1M input tokens in US Virginia; thinking-mode output uses the same listed rates
View model →

Multilingual semantic search, retrieval-augmented generation, code retrieval, recommendation, clustering, classification, and large-scale text vectorization

Type Embedding
Context 131K
Speed 8/10
Status

Current and available through Alibaba Cloud Model Studio internationally

Input $0.07 per 1 million input tokens in Singapore international deployment; China pricing may differ by region and billing mode
View model →
Qwen logo
Qwen3.7-Text-Embedding

qwen3.7-text-embedding-flash

Cost-sensitive, high-throughput multilingual text vectorization, semantic search, RAG, recommendation, clustering, and classification.

Type Lightweight
Context 131K
Reasoning 1/10
Speed 9/10
Status

Current

Input CNY 0.000125 per 1,000 input tokens in China (Beijing); batch pricing CNY 0.000063 per 1,000 input tokens. Regional pricing may differ.

Advanced reasoning, coding, scientific and professional research, long-context analysis, and long-horizon agent workflows

Type Reasoning
Context 1M
Reasoning 10/10
Speed 4/10
Tool use Web search Streaming
Status

Current and available; open-weight release and hosted API model

Input $2 per 1 million tokens internationally; $1.65 per 1 million tokens in China (Beijing)
Output $6 per 1 million tokens internationally; $4.951 per 1 million tokens in China (Beijing)
View model →

Coding assistants, repository analysis, visual document workflows, long-context research, office automation, multimodal agents, and tool-using applications

Type Multimodal
Context 1M
Reasoning 8/10
Speed 6/10
Multimodal Image input Video input
Status

Active and currently available; open-weight release and Alibaba Cloud Model Studio API model

Input $0.424 per 1M tokens in China Beijing; $0.50 per 1M tokens in Singapore. Implicit cache input is $0.085 per 1M tokens in Beijing and $0.10 per 1M tokens in Singapore. Explicit cache creation is $0.53 per 1M tokens in Beijing and $0.625 per 1M tokens in Si
Output $1.696 per 1M tokens in China Beijing; $3 per 1M tokens in Singapore.
View model →

Fast long-context reasoning, coding assistance, visual document and chart analysis, video understanding, function-calling agents, and high-concurrency applications

Type Multimodal
Context 1M
Reasoning 8/10
Speed 9/10
Multimodal Image input Video input
Status

Current and accessible through Alibaba Cloud Model Studio

Input $0.113 per 1 million input tokens; implicit cache input $0.014 per 1 million tokens; explicit cache creation $0.177 per 1 million tokens; explicit cache read $0.014 per 1 million tokens
Output $0.382 per 1 million output tokens
View model →

High-volume coding, agentic workflows, long-context document and codebase analysis, office automation, and multimodal text-image-video understanding

Type Multimodal
Context 262K
Reasoning 8/10
Speed 9/10
Multimodal Image input Video input
Status

Current open-weight experimental preview

View model →

Complex coding, autonomous software engineering, long-horizon agent workflows, professional document analysis, visual reasoning, long videos, and demanding research tasks.

Type Multimodal
Context 1M
Reasoning 9/10
Speed 7/10
Multimodal Image input Video input
Status

Stable official release; currently available through Alibaba Cloud Model Studio

Input US$1.65 per 1M input tokens and US$2.00 per 1M input tokens for International scope; regional pricing varies. Cached-input pricing starts at US$0.206 per 1M tokens for the listed regional/global scope and US$0.25 per 1M tokens for International scope.
Output US$4.951 per 1M output tokens for the listed regional/global scope and US$6.00 per 1M output tokens for International scope.
View model →

Real-time speech translation, multilingual meetings, live interpretation, translated voice communication, and audiovisual translation with low latency.

Type Multimodal
Context 53K
Reasoning 4/10
Speed 9/10
Multimodal Image input Audio input
Status

Current stable model

Input Singapore: audio input $7.50 per 1 million tokens; image input $0.55 per 1 million tokens. China (Beijing): audio input $5.653 per 1 million tokens; image input $0.466 per 1 million tokens.
Output Singapore: text output $20 per 1 million tokens; audio output $30 per 1 million tokens. China (Beijing): text output $14.133 per 1 million tokens; audio output $22.613 per 1 million tokens.
View model →

Long-form audio and video understanding, multimedia analysis, audio-visual agents, content summarization, and tool-using workflows

Type Multimodal
Context 1M
Reasoning 8/10
Speed 8/10
Multimodal Image input Audio input
Status

Current and available

Input USD 0.15 per 1 million input tokens for International deployment; USD 0.016 per 1 million cache-hit input tokens
Output USD 0.47 per 1 million output tokens for International deployment
View model →

Real-time voice assistants, speech-to-speech applications, interactive video agents, live media analysis, multimodal customer service, meeting and collaboration interfaces, and applications requiring tool or MCP integration.

Type Multimodal
Context 197K
Reasoning 7/10
Speed 9/10
Multimodal Audio input Video input
Status

Current and available; international deployment in Singapore

Input China (Beijing): CNY 1.5 per 1 million tokens for text/images/video input and CNY 6 per 1 million tokens for audio input. Singapore: CNY 1.677 per 1 million tokens for text/images/video input and CNY 6.781 per 1 million tokens for audio input.
Output China (Beijing): CNY 4.5 per 1 million tokens for text output and CNY 12 per 1 million tokens for audio output. Singapore: CNY 5.104 per 1 million tokens for text output and CNY 13.636 per 1 million tokens for audio output.
View model →
Qwen logo
Qwen3Guard

Qwen3Guard-Stream-0.6B

Low-latency local moderation of streaming prompts and language-model responses

Type Other
Context 33K
Reasoning 2/10
Speed 8/10
Streaming
Status

Current open-weight model

Real-time moderation of streamed prompts and language-model responses

Type Other
Context 8K
Reasoning 3/10
Speed 8/10
Streaming
Status

Available open-weight model

View model →
Qwen logo
Qwen3Guard

Qwen3Guard-Stream-8B

Real-time multilingual moderation of prompts and streamed language-model responses

Type Other
Context 8K
Reasoning 1/10
Speed 7/10
Fine-tuning Streaming
Status

Current open-weight model

Qwen logo
Qwen Character

qwen-flash-character

Low-latency character dialogue, virtual companions, game NPCs, role-playing applications, IP character replication, and conversational smart devices

Type Lightweight
Context 33K
Reasoning 4/10
Speed 9/10
Web search Streaming
Status

Current; dynamically updated managed model

Input Regional pricing: USD 0.034 per 1M input tokens in Beijing and US Virginia; USD 0.05 per 1M input tokens in Singapore. Cached input is listed at USD 0.007 per 1M tokens in Beijing and US Virginia and USD 0.01 in Singapore. Alibaba Cloud also publishes loc
Output Regional pricing: USD 0.203 per 1M output tokens in Beijing and US Virginia; USD 0.40 per 1M output tokens in Singapore. Alibaba Cloud also publishes localized regional prices.
View model →
Qwen logo
Qwen Character

qwen-plus-character

Character role-play, virtual social applications, game NPCs, IP character replication, smart toys, in-car assistants, and empathetic conversational experiences

Type Other
Context 33K
Reasoning 3/10
Speed 8/10
Web search
Status

Current; dynamically updated model

Input $0.50 per 1 million tokens internationally; $0.115 per 1 million tokens in China (Beijing), Hong Kong, Germany, the United States, and Japan
Output $1.40 per 1 million tokens internationally; $0.287 per 1 million tokens in China (Beijing), Hong Kong, Germany, the United States, and Japan

Complex text-to-image layouts, multilingual typography, posters, menus, storyboards, interface mockups, product visuals, and image editing with one to three reference images

Type Image Generation
Context 5K
Reasoning 2/10
Speed 6/10
Multimodal Image input Media output
Status

Current and available

Input $0.003 per image in the international Singapore deployment
Output $0.04 per 1K image or $0.075 per 2K image in the international Singapore deployment; regional pricing varies
View model →

Instruction-based image editing, multi-image fusion, character-consistent compositions, object replacement, style transfer, poster and text editing, and product-image variation.

Type Image Editing
Reasoning 5/10
Speed 8/10
Multimodal Image input Media output
Status

Current; canonical model currently equivalent to qwen-image-edit-plus-2025-10-30

Input 0; image editing is billed by output image
Output $0.028671 per image in China (Beijing); $0.03 per image in Singapore/international
View model →
Qwen logo
Qwen Image

Qwen-Image-Max

Realistic text-to-image generation, creative concepts, marketing visuals, editorial artwork, product concepts, and general-purpose image creation.

Type Image Generation
Reasoning 1/10
Speed 5/10
Media output
Status

Current and available for text-to-image generation; the undated qwen-image-max identifier is currently equivalent to qwen-image-max-2025-12-30.

Input No separate input charge; text prompt input is included in per-image billing.
Output $0.075/image in Singapore international deployment; $0.071677/image in China (Beijing).
View model →
Qwen logo
Qwen Image 3.0

qwen-image-3.0

Fast text-to-image generation, text-heavy layouts, image editing, reference-image compositing, marketing graphics, and high-volume image workflows.

Type Other
Reasoning 3/10
Speed 8/10
Multimodal Image input Media output
Status

Current and available through Alibaba Cloud Model Studio

Input $0.003 per input image internationally; regional pricing varies
Output $0.03 per generated image internationally; regional pricing varies
View model →
Qwen logo
QwQ

qwq-plus

Mathematics, coding, research-style analysis, and complex multi-step reasoning

Type Reasoning
Context 131K
Reasoning 9/10
Speed 6/10
Tool use Web search Streaming
Status

Legacy; currently accessible as of October 7, 2026; scheduled to cease service on October 10, 2026 at 00:00 UTC+08, subject to the actual change time

Input $0.230 per 1M tokens in China (Beijing); $0.80 per 1M tokens in Singapore international deployment
Output $0.574 per 1M tokens in China (Beijing); $2.40 per 1M tokens in Singapore international deployment
Qwen logo
Tongyi Embedding Vision

tongyi-embedding-vision-plus

Cross-modal retrieval, image and video similarity search, multimodal semantic indexing, recommendation, and content classification

Type Multimodal Embedding
Reasoning 1/10
Speed 8/10
Multimodal Image input Video input
Status

Available

Input $0.09 per 1 million input tokens for text, image, and video in the international Singapore deployment
Output $0; embedding output is free
Qwen logo

Qwen3.8-Max-Prime