Cost-sensitive cross-modal retrieval, image and video search, multimedia catalog indexing, and vector search
Type
Multimodal
Context
1K
Reasoning
1/10
Speed
8/10
Multimodal
Image input
Video input
Status
Current and available through Alibaba Cloud Model Studio International deployment
Input
Image/video: $0.03 per 1 million input tokens; text: $0.09 per 1 million input tokens
Output
Free; embedding output is not charged
View model
→
Cross-modal retrieval, text-to-image search, image similarity, video search, semantic classification, clustering, and multimodal vector indexing.
Type
Multimodal Embedding
Context
512
Reasoning
0/10
Speed
0/10
Multimodal
Image input
Video input
Status
Current and accessible through Alibaba Cloud Model Studio in the China (Beijing) region; free trial pricing is listed.
View model
→
Agent-environment simulation, tool-interaction modeling, terminal and software-engineering trajectories, and research on language world models
Type
Other
Context
262K
Reasoning
8/10
Speed
7/10
Tool use
Streaming
Status
Current open-weight model
View model
→
Low-latency duplex voice assistants, real-time customer service, AI companions, and streamed speech-to-speech applications
Type
Multimodal
Context
41K
Reasoning
6/10
Speed
8/10
Multimodal
Audio input
Media output
Status
Current and accessible; standard-edition real-time duplex speech model
Input
Singapore: text $0.80/M tokens; audio $6.40/M tokens. China (Beijing): text $0.688/M tokens; audio $5.501/M tokens.
Output
Singapore: text $6.40/M tokens; audio $24/M tokens. China (Beijing): text $5.501/M tokens; audio $20.628/M tokens.
View model
→
Low-latency voice assistants, customer service, AI companions, full-duplex spoken interaction, and voice applications using tools or cloned voices.
Type
Multimodal
Context
262K
Reasoning
4/10
Speed
9/10
Multimodal
Audio input
Media output
Status
Current and available
Input
Singapore: $0.80 per 1M text-input tokens; $6.40 per 1M audio-input tokens. China (Beijing): $0.688 per 1M text-input tokens; $5.501 per 1M audio-input tokens.
Output
Singapore: $6.40 per 1M text-output tokens; $24 per 1M audio-output tokens. China (Beijing): $5.501 per 1M text-output tokens; $20.628 per 1M audio-output tokens.
View model
→
Low-latency voice assistants, real-time customer service, duplex speech conversations, interactive voice agents, and applications requiring streaming audio responses.
Type
Multimodal
Context
41K
Reasoning
6/10
Speed
9/10
Multimodal
Audio input
Media output
Status
Available; current API access confirmed, but newer models are recommended for some new projects
Input
International/Singapore: text input $0.23 per 1M tokens; audio input $0.93 per 1M tokens. China (Beijing): text input $0.413 per 1M tokens; audio input $4.126 per 1M tokens.
Output
International/Singapore: text-only output $0.70 per 1M tokens; combined text and audio output $1.87 per 1M tokens. China (Beijing): text-only output $4.126 per 1M tokens; combined text and audio output $13.752 per 1M tokens.
View model
→
Real-time multilingual speech transcription, live captions, meeting transcription, voice interfaces, streaming subtitles, and Chinese-dialect recognition.
Type
Speech Recognition
Context
8K
Reasoning
1/10
Speed
9/10
Audio input
Streaming
Status
Current and available through Alibaba Cloud Model Studio in the International/Singapore and China (Beijing) regions.
Input
International/Singapore: USD 0.93 per 1 million input tokens; China (Beijing): USD 0.848 per 1 million input tokens.
Output
International/Singapore: USD 0.70 per 1 million output tokens; China (Beijing): USD 0.636 per 1 million output tokens.
View model
→
Expressive text-to-speech, audiobooks, film and video dubbing, content creation, premium voice services, multilingual speech, dialect synthesis, and voice cloning
Type
Other
Reasoning
1/10
Speed
8/10
Media output
Streaming
Status
Current and available
Input
USD 0.20 per 10,000 input characters in Singapore/International; USD 0.19253 per 10,000 input characters in China (Beijing)
Output
Not separately billed; pricing is based on input characters
View model
→
Long-form offline transcription of meetings, interviews, calls, media files, and multilingual or dialect-rich recordings
Type
Other
Context
8K
Reasoning
2/10
Speed
7/10
Audio input
Status
Current and publicly available
Input
USD 0.15 per 1 million input tokens in Singapore; USD 0.113 per 1 million input tokens in China (Beijing)
Output
USD 0.47 per 1 million output tokens in Singapore; USD 0.382 per 1 million output tokens in China (Beijing)
View model
→
Autonomous-driving research, driving-scene VQA, 3D BEV perception, trajectory prediction, and embodied-AI experimentation
Type
Multimodal
Reasoning
6/10
Speed
5/10
Multimodal
Image input
Media output
Status
Current open-weight research release
View model
→
Text-to-image generation, image editing, text rendering in images, photorealistic scenes, creative design, and producing multiple image variants.
Type
Image Generation
Reasoning
3/10
Speed
8/10
Multimodal
Image input
Media output
Status
Current; accelerated model; functionally equivalent to qwen-image-2.0-2026-03-03
Input
$0 per input image; only generated output images are billed
Output
$0.035 per image internationally; $0.028671 per image in China (Beijing) and Singapore
View model
→
Qwen-Image
Qwen-Image-2.1
Local text-to-image generation, multi-reference composition, transparent asset creation, product visuals, and localized image editing
Type
Multimodal
Speed
7/10
Multimodal
Image input
Media output
Status
Current open-weight release
Model page unavailable
Cost-sensitive text-to-image generation, posters, marketing graphics, illustrations, and images containing Chinese or English text
Type
Image Generation
Speed
8/10
Media output
Status
Current and accessible; currently equivalent to qwen-image
Input
Not applicable; image generation is billed per output image
Output
$0.03/image internationally; $0.028671/image in China (Beijing)
View model
→
Qwen-Image-2.1
Qwen-Image-2.1-PE-I2I
Rewriting vague image-editing instructions, preserving source-image details, coordinating multi-image edits, and preparing prompts for Qwen-Image-2.1.
Type
Multimodal
Context
262K
Reasoning
5/10
Speed
5/10
Multimodal
Image input
Streaming
Status
Current; open-weight research release
Model page unavailable
Qwen-Image-2.1
Qwen-Image-2.1-PE-T2I
Expanding short or multilingual image requests into detailed prompts for Qwen-Image-2.1
Type
Other
Context
262K
Reasoning
5/10
Speed
5/10
Status
Current open-weight release
Model page unavailable
Natural-language single-image editing, bilingual text changes, object insertion or removal, style transfer, pose changes, and image fusion
Type
Other
Reasoning
2/10
Speed
6/10
Multimodal
Image input
Media output
Status
Current and accessible
Input
$0.045 per image in Singapore for international deployment
View model
→
High-quality image editing, multi-image composition, industrial design concepts, geometric transformations, character-consistent edits, and controlled visual revisions
Type
Other
Reasoning
2/10
Speed
6/10
Multimodal
Image input
Media output
Status
Current and available; canonical model ID is functionally equivalent to qwen-image-edit-max-2026-01-16
Input
$0.075 per output image in the International/Singapore region; $0.071677 per output image in China (Beijing)
Output
$0.075 per output image in the International/Singapore region; $0.071677 per output image in China (Beijing)
View model
→
Professional text-to-image generation, image editing, posters, infographics, multilingual in-image text, photorealistic scenes, and reference-based creative production
Type
Other
Reasoning
6/10
Speed
5/10
Multimodal
Image input
Media output
Status
Current rolling model; functionally equivalent to qwen-image-2.0-pro-2026-04-22
Output
$0.075 per image internationally; $0.071676 per image in China (Beijing)
View model
→
Complex multilingual text generation, coding, logical reasoning, creative writing, structured extraction, and enterprise applications
Type
General Purpose
Context
33K
Reasoning
7/10
Speed
6/10
Tool use
Streaming
Status
Current and accessible rolling-update mainline model
Input
US$1.60 per 1 million tokens for International deployment; US$0.345 per 1 million tokens in China mainland
Output
US$6.40 per 1 million tokens for International deployment; US$1.377 per 1 million tokens in China mainland
Model page unavailable
General-purpose text generation, long-context analysis, multilingual applications, structured business workflows, function-calling agents, and applications that need optional reasoning mode.
Type
General Purpose
Context
1M
Reasoning
7/10
Speed
7/10
Tool use
Web search
Streaming
Status
Current; qwen-plus currently resolves to the qwen-plus-2025-12-01 snapshot
Input
$0.40 per 1M input tokens for 0–256K input; $1.20 per 1M input tokens above 256K in the International Singapore pricing tier. International Global pricing is $0.115 per 1M tokens for up to 128K, $0.345 for 128K–256K, and $0.689 for 256K–1M.
Output
$1.20 per 1M non-thinking output tokens and $4.00 per 1M thinking output tokens for 0–256K input; $3.60 non-thinking and $12.00 thinking per 1M output tokens above 256K in the International Singapore pricing tier. International Global pricing differs by r
View model
→
High-volume customer support, simple to moderate question answering, summarization, rewriting, structured text extraction, and cost-sensitive applications
Type
Lightweight
Context
131K
Reasoning
6/10
Speed
9/10
Web search
Streaming
Status
Currently accessible; no longer being updated; Alibaba Cloud recommends switching to Qwen-Flash
Input
$0.05 per 1M tokens international non-thinking; $0.05 per 1M tokens international thinking; $0.044 per 1M tokens China (Beijing)
Output
$0.20 per 1M tokens international non-thinking; $0.50 per 1M tokens international thinking; $0.087 per 1M tokens China non-thinking; $0.431 per 1M tokens China thinking
Model page unavailable
Fast image and video understanding, visual question answering, document analysis, and multimodal extraction
Type
Multimodal
Context
33K
Reasoning
6/10
Speed
9/10
Multimodal
Image input
Video input
Model page unavailable
Complex image and video understanding, document analysis, chart interpretation, visual question answering, and structured extraction
Type
Multimodal
Context
131K
Reasoning
8/10
Speed
6/10
Multimodal
Image input
Video input
Status
Currently accessible legacy visual language model; current qwen-vl-max endpoint is functionally equivalent to qwen-vl-max-2025-08-13
Input
$0.229 per 1M tokens in China (Beijing) and Singapore; $0.80 per 1M tokens for international deployment
Output
$0.573 per 1M tokens in China (Beijing) and Singapore; $3.20 per 1M tokens for international deployment
View model
→
High-resolution image and video understanding, OCR-style text recognition, document analysis, visual question answering, and multimodal assistants
Type
Multimodal
Context
131K
Reasoning
6/10
Speed
7/10
Multimodal
Image input
Video input
Status
Current rolling model; functionally equivalent to qwen-vl-plus-2025-08-15
Input
$0.115 per 1M tokens in China (Beijing); $0.21 per 1M tokens in Singapore
Output
$0.287 per 1M tokens in China (Beijing); $0.63 per 1M tokens in Singapore
Model page unavailable
Self-hosted chat assistants, multilingual text generation, coding and mathematics assistance, long-context document work, structured text generation, and cost-sensitive private deployments
Type
General Purpose
Context
131K
Reasoning
7/10
Speed
7/10
Tool use
Fine-tuning
Streaming
Status
Current open-weight model; publicly available for download and self-hosted deployment
View model
→
Self-hosted multilingual assistants, document processing, RAG, coding support, structured extraction, and cost-conscious production deployments.
Type
General Purpose
Context
131K
Reasoning
7/10
Speed
7/10
Tool use
Fine-tuning
Streaming
Status
Legacy open-weight model; downloadable and self-hostable, while Qwen2.5 API models marked deprecated are no longer callable through current Alibaba Cloud Model Studio pricing documentation.
Input
Not applicable for the downloadable checkpoint; no current official per-token hosted price verified for this exact model.
Output
Not applicable for the downloadable checkpoint; no current official per-token hosted price verified for this exact model.
View model
→
Self-hosted assistants, multilingual text generation, long-document processing, coding support, structured extraction, RAG, and agent applications
Type
General Purpose
Context
131K
Reasoning
7/10
Speed
5/10
Tool use
Fine-tuning
Streaming
Status
Available; open-weight model
View model
→
Self-hosted multilingual assistants, coding, mathematics, document analysis, structured text generation, and long-context workloads
Type
General Purpose
Context
131K
Reasoning
8/10
Speed
4/10
Fine-tuning
Streaming
Status
Available as an open-weight model; legacy relative to newer Qwen generations
View model
→
Multimodal assistants, audio and video understanding, visual question answering, voice interaction, speech instruction following, and local multimodal AI research
Type
Multimodal
Context
33K
Reasoning
7/10
Speed
7/10
Multimodal
Image input
Audio input
Status
Current; open-weight model and available through Alibaba Cloud Model Studio
Input
International: $0.10 per 1M text tokens; $6.76 per 1M audio tokens; $0.28 per 1M image/video tokens. China Beijing: $0.087 per 1M text tokens; $5.448 per 1M audio tokens; $0.287 per 1M image/video tokens.
Output
International: $0.40 per 1M tokens for text-only input; $0.84 per 1M tokens for text after multimodal input; $13.51 per 1M tokens for text-and-audio output. China Beijing: $0.345 per 1M tokens for text-only input; $0.861 per 1M tokens for text after multi
View model
→
Local or self-hosted image and video understanding, OCR, document extraction, chart and diagram analysis, visual question answering, visual grounding, and multimodal research
Type
Multimodal
Context
33K
Reasoning
7/10
Speed
7/10
Multimodal
Image input
Video input
Status
Open-weight and downloadable; supported as a fine-tuning base model in Alibaba Cloud Model Studio; current hosted inference availability and standard pricing for this exact model are not clearly listed in the latest Model Studio inference-pricing catalog.
View model
→
High-quality image, document, chart, screenshot, OCR, visual-grounding, and video analysis; multimodal agents and self-hosted experimentation
Type
Multimodal
Context
131K
Reasoning
8/10
Speed
3/10
Multimodal
Image input
Video input
Status
Available open-weight model; older Qwen2.5-VL generation and still listed by Alibaba Cloud Model Studio
View model
→
Fast, high-volume text generation; long-context analysis; summarization; extraction; structured outputs; and applications needing optional reasoning.
Type
Lightweight
Context
1M
Reasoning
7/10
Speed
9/10
Tool use
Web search
Streaming
Status
Current and available; canonical qwen-flash identifier is functionally equivalent to qwen-flash-2025-07-28
Input
Tiered per 1M input tokens. China Beijing: CNY 0.15 up to 128K input, CNY 0.60 above 128K to 256K, CNY 1.20 above 256K to 1M. International Singapore: CNY 0.367 up to 256K, CNY 1.835 above 256K to 1M. Regional pricing and promotional offers may vary.
Output
Tiered per 1M output tokens. China Beijing: CNY 1.50 up to 128K input, CNY 6 above 128K to 256K, CNY 12 above 256K to 1M. International Singapore: CNY 2.936 up to 256K, CNY 14.678 above 256K to 1M. Regional pricing and promotional offers may vary.
View model
→
Lightweight local assistants, offline prototypes, embedded experimentation, education, simple text generation, and resource-constrained deployments
Type
Lightweight
Context
33K
Reasoning
3/10
Speed
9/10
Tool use
Fine-tuning
Streaming
Status
Current open-weight model
View model
→
Efficient local inference, edge applications, multilingual chat, lightweight reasoning, coding assistance, and tool-enabled agents
Type
Lightweight
Context
33K
Reasoning
7/10
Speed
9/10
Tool use
Fine-tuning
Streaming
Status
Current open-weight model
Input
No official hosted API price; open-weight model
Output
No official hosted API price; open-weight model
View model
→
Local and self-hosted chat, compact reasoning, coding assistance, multilingual applications, retrieval-augmented generation, and lightweight tool-using agents
Type
General Purpose
Context
33K
Reasoning
7/10
Speed
8/10
Tool use
Fine-tuning
Streaming
Status
Current open-weight model
View model
→
Cost-efficient reasoning, coding, multilingual assistants, local deployment, structured text generation, tool-enabled agents, and fine-tuned applications
Type
General Purpose
Context
131K
Reasoning
8/10
Speed
8/10
Tool use
Fine-tuning
Streaming
Status
Current and accessible; open-weight model available for self-hosting and Alibaba Cloud Model Studio API deployment
Input
$0.072 per 1 million tokens for Global deployment; $0.18 per 1 million tokens for International deployment
Output
$0.287 per 1 million tokens for non-thinking mode and $0.717 per 1 million tokens for thinking mode in Global deployment; International pricing is $0.70 non-thinking and $2.10 thinking per 1 million tokens
View model
→
Local deployment, multilingual assistants, reasoning, mathematics, coding, structured text generation, research, and cost-sensitive agent workflows.
Type
General Purpose
Context
131K
Reasoning
7/10
Speed
8/10
Tool use
Fine-tuning
Streaming
Status
Current open-weight model; also available as qwen3-14b through Alibaba Cloud Model Studio, with regional capability and pricing differences.
Input
Alibaba Cloud Model Studio: China (Beijing) $0.144 per 1 million input tokens; Singapore $0.35 per 1 million input tokens. Regional prices and deployment terms may vary.
Output
Alibaba Cloud Model Studio: China (Beijing) $0.574 per 1 million output tokens in non-thinking mode and $1.434 per 1 million output tokens in thinking mode; Singapore $1.4 per 1 million output tokens in non-thinking mode and $4.2 per 1 million output toke
View model
→
Self-hosted assistants, coding, mathematical and logical reasoning, multilingual applications, long-context document processing, and agentic tool-use systems
Type
Reasoning
Context
131K
Reasoning
8/10
Speed
9/10
Tool use
Streaming
Status
Current and accessible; open-weight release with hosted Alibaba Cloud Model Studio availability
Input
US$0.108 per 1 million input tokens in several global deployments; US$0.20 per 1 million input tokens for the international Singapore deployment
Output
US$0.431 per 1 million output tokens in several global deployments; US$0.80 per 1 million output tokens for the international Singapore deployment. Thinking output is separately priced at higher rates where exposed.
View model
→
Self-hosted reasoning assistants, coding agents, mathematics, multilingual applications, structured text generation, and tool-calling workflows
Type
Reasoning
Context
256K
Reasoning
9/10
Speed
6/10
Tool use
Fine-tuning
Streaming
Status
Current; open-weight model and available through Alibaba Cloud Model Studio
Input
$0.16 per 1 million tokens in international Singapore, Germany Frankfurt, and US Virginia deployments; $0.287 per 1 million tokens in China Beijing
Output
$0.64 per 1 million tokens in international Singapore, Germany Frankfurt, and US Virginia deployments; China Beijing: $1.147 per 1 million non-thinking tokens or $2.868 per 1 million thinking tokens
View model
→
Complex reasoning, mathematics, software development, multilingual applications, function calling, agentic workflows, research and self-hosted open-weight deployment
Type
Reasoning
Context
131K
Reasoning
9/10
Speed
6/10
Tool use
Fine-tuning
Streaming
Status
Current and accessible through Alibaba Cloud Model Studio; original Qwen3 open-weight release, with newer 2507 instruct and thinking variants available separately
Input
$0.287 per 1M tokens for standard input in the United States, Germany and China; $0.700 per 1M tokens in Singapore. Thinking-mode input is priced the same.
Output
$1.147 per 1M tokens for standard output in the United States, Germany and China; $2.868 per 1M tokens for thinking-mode output. Singapore pricing is $2.800 standard output and $8.400 thinking-mode output per 1M tokens.
View model
→
Multilingual semantic search, RAG candidate reranking, enterprise document retrieval, knowledge-base search, and improving search-result relevance.
Type
Other
Context
4K
Reasoning
2/10
Speed
8/10
Status
Current and available
Input
$0.10 per 1 million input tokens for the international Singapore deployment; pricing is region-dependent. Output is free.
Output
Free
View model
→
Low-cost multilingual speech-to-text, language identification, offline transcription, real-time streaming ASR, and high-throughput deployments
Type
Other
Context
66K
Reasoning
2/10
Speed
8/10
Multimodal
Audio input
Streaming
Status
Current, open-weight, Apache 2.0 licensed
View model
→
Multilingual speech transcription, language identification, long-audio processing, and self-hosted or streaming ASR applications
Type
Other
Reasoning
2/10
Speed
8/10
Multimodal
Audio input
Fine-tuning
Status
Current; open-weight and downloadable
Input
No official hosted API price identified for this exact open-weight checkpoint
Output
No separate output charge for the self-hosted checkpoint
View model
→
Qwen3-ASR
Qwen3-ForcedAligner-0.6B
Word- and character-level speech timestamp alignment, subtitle synchronization, transcript timing, and audio annotation
Type
Other
Reasoning
1/10
Speed
8/10
Multimodal
Audio input
Status
Current open-weight model
Model page unavailable
Repository-scale coding agents, code generation, code completion, debugging, refactoring, terminal workflows, and cost-sensitive self-hosted deployments.
Type
Coding
Context
262K
Reasoning
8/10
Speed
8/10
Tool use
Streaming
Status
Current; open-weight model and available through Alibaba Cloud Model Studio
Input
USD 0.144 per 1M tokens for China Beijing input up to 32K; USD 0.216 for 32K–128K; USD 0.359 for 128K–256K. International Singapore and Frankfurt pricing is USD 0.30, USD 0.50, and USD 0.80 per 1M input tokens across the same bands.
Output
USD 0.574 per 1M tokens for China Beijing output up to 32K; USD 0.861 for 32K–128K; USD 1.434 for 128K–256K. International Singapore and Frankfurt pricing is USD 1.50, USD 2.50, and USD 4.00 per 1M output tokens across the same bands.
View model
→
Large-codebase analysis, code generation, refactoring, debugging, documentation and long-context coding-agent workflows
Type
Coding
Context
1M
Reasoning
7/10
Speed
7/10
Tool use
Streaming
Status
Current; canonical qwen3-coder-plus currently equivalent to qwen3-coder-plus-2025-09-23
Input
US Virginia: $0.574 per 1M tokens up to 32K input; $0.861 for 32K-128K; $1.434 for 128K-256K; $2.868 for 256K-1M. Regional pricing varies.
Output
US Virginia: $2.294 per 1M tokens up to 32K input; $3.441 for 32K-128K; $5.735 for 128K-256K; $28.671 for 256K-1M. Regional pricing varies.
View model
→
Qwen3-Embedding
text-embedding-v3
Multilingual text embeddings, semantic search, RAG, recommendation, clustering, classification, and migration of existing v3 vector indexes
Type
Embedding
Context
8K
Speed
8/10
Status
Available; retained primarily for compatibility with existing v3 embedding indexes
Input
$0.07 per 1 million input tokens in the Singapore international deployment
Output
Free
Model page unavailable
Qwen3-Embedding
text-embedding-v4
Multilingual semantic search, RAG pipelines, vector databases, document retrieval, clustering, classification, recommendation, and code retrieval
Type
Other
Context
8K
Reasoning
1/10
Speed
8/10
Status
Current and accessible through Alibaba Cloud Model Studio
Input
$0.07 per 1M input tokens in Singapore/International and Hong Kong; $0.072 per 1M input tokens in China (Beijing). Beijing batch-file processing is listed at $0.036 per 1M tokens.
Output
Free; pricing is based on input tokens
Model page unavailable
Streaming translation of recorded or uploaded audio and video, multilingual subtitles, translated voice tracks, and applications requiring translated text or synthesized speech.
Type
Other
Context
53K
Reasoning
2/10
Speed
8/10
Multimodal
Audio input
Video input
Status
Current stable model
Input
Audio input and output are billed at 12.5 tokens per second, with audio shorter than one second billed as one second. Video usage additionally consumes video tokens based on sampled frames and resolution. Monetary rates depend on the applicable Alibaba Cl
Output
Audio output is billed at 12.5 tokens per second. Text output uses the applicable Model Studio token rate when charged separately. No standalone fixed monetary rate was verified for this exact model in the model documentation.
View model
→
Real-time multilingual speech interpretation, live voice translation, conference translation, streaming media, and audiovisual translation with text or synthesized speech output
Type
Other
Context
53K
Reasoning
1/10
Speed
8/10
Multimodal
Image input
Audio input
Status
Legacy; still available; no longer recommended for new use
Input
China (Beijing): CNY 64 per 1M input audio tokens and CNY 8 per 1M input image tokens. Singapore: CNY 73.392 per 1M input audio tokens and CNY 9.541 per 1M input image tokens.
Output
China (Beijing): CNY 64 per 1M output text tokens and CNY 240 per 1M output audio tokens. Singapore: CNY 73.392 per 1M output text tokens and CNY 278.891 per 1M output audio tokens.
View model
→
Complex reasoning, coding assistance, web-grounded agents, function calling, structured extraction, and long-context text analysis
Type
Reasoning
Context
262K
Reasoning
9/10
Speed
7/10
Tool use
Web search
Streaming
Input
International: $1.20 per 1M input tokens up to 32K; $2.40 per 1M above 32K to 128K; $3.00 per 1M above 128K to 256K. Regional prices vary.
Output
International: $6.00 per 1M output tokens up to 32K; $12.00 per 1M above 32K to 128K; $15.00 per 1M above 128K to 256K. Regional prices vary.
View model
→
Local visual assistants, OCR, document and chart analysis, image question answering, lightweight video understanding, and multimodal prototyping
Type
Multimodal
Context
256K
Reasoning
6/10
Speed
8/10
Multimodal
Image input
Video input
Status
Current open-weight model
View model
→
Local image and video understanding, OCR, document extraction, visual question answering, visual coding, and lightweight multimodal agents
Type
Multimodal
Context
262K
Reasoning
6/10
Speed
8/10
Multimodal
Image input
Video input
Status
Current open-weight model; available on Hugging Face and supported for supervised fine-tuning in Alibaba Cloud Model Studio
View model
→
Local or hosted image and video understanding, OCR, document extraction, visual question answering, spatial reasoning, screenshot analysis, multimodal agents, and structured data extraction
Type
Multimodal
Context
262K
Reasoning
7/10
Speed
8/10
Multimodal
Image input
Video input
Status
Current; open-weight model with hosted inference available through Alibaba Cloud Model Studio
Input
$0.072 per 1 million input tokens
Output
$0.287 per 1 million output tokens
View model
→
Image and video understanding, OCR, document analysis, spatial reasoning, visual coding, long-context multimodal tasks, and visual-agent applications
Type
Multimodal
Context
256K
Reasoning
8/10
Speed
8/10
Multimodal
Image input
Video input
Status
Current and available; open-weight release and Alibaba Cloud Model Studio API model
Input
$0.20 per 1 million tokens for the listed international deployment; regional pricing may differ
Output
$0.80 per 1 million tokens for the listed international deployment; regional pricing may differ
View model
→
Document intelligence, OCR, image and video understanding, spatial reasoning, visual coding, and visual-agent applications
Type
Multimodal
Context
131K
Reasoning
8/10
Speed
6/10
Multimodal
Image input
Video input
Status
Current; open-weight checkpoint and available through Alibaba Cloud Model Studio managed inference
Input
$0.16 per 1 million input tokens in Alibaba Cloud Model Studio US (Virginia) global deployment; China Beijing pricing is $0.287 per 1 million input tokens
Output
$0.64 per 1 million output tokens in Alibaba Cloud Model Studio US (Virginia) global deployment; China Beijing pricing is $1.147 per 1 million output tokens
View model
→
High-quality image and video understanding, OCR, document intelligence, visual coding, spatial reasoning, long-context multimodal analysis and visual-agent applications
Type
Multimodal
Context
131K
Reasoning
8/10
Speed
5/10
Multimodal
Image input
Video input
Status
Current and accessible; open-weight release and Alibaba Cloud Model Studio API availability
Input
$0.287 per 1 million tokens in the US Virginia global deployment; $0.400 per 1 million tokens in Singapore
Output
$1.147 per 1 million tokens in the US Virginia global deployment; $1.600 per 1 million tokens in Singapore
View model
→
Multimodal reranking, cross-modal search, image retrieval, video retrieval, image clustering, and multimodal RAG
Type
Multimodal
Context
120K
Reasoning
2/10
Speed
7/10
Multimodal
Image input
Video input
Status
Current and available through Alibaba Cloud Model Studio in the China (Beijing) region
Input
Text input: $0.10 per 1 million tokens; image input: $0.258 per 1 million tokens; China (Beijing)
Output
Free or not separately charged for reranking output
Model page unavailable
Qwen3-VL-Embedding
qwen3-vl-embedding
Multimodal vector search, cross-modal retrieval, image and video search, semantic clustering, tagging, and retrieval pipelines combining text with visual content.
Type
Multimodal
Context
32K
Reasoning
4/10
Speed
7/10
Multimodal
Image input
Video input
Status
Current and available through Alibaba Cloud Model Studio; open-weight 2B and 8B variants are also available through Qwen's official model repositories.
Input
Text: $0.10 per 1 million tokens; image/video: $0.258 per 1 million tokens in China (Beijing).
Output
No separate output-token price documented; the model returns vector embeddings.
Model page unavailable
Long-context multimodal analysis, document understanding, video and image interpretation, general reasoning, coding, tool-enabled assistants, and self-hosted deployment
Type
Multimodal
Context
262K
Reasoning
8/10
Speed
8/10
Multimodal
Image input
Video input
Status
Accessible; no longer recommended for new projects
Input
$0.086 per 1M input tokens for requests up to 128K input tokens in US Virginia; $0.258 per 1M input tokens for 128K-256K requests
Output
$0.688 per 1M output tokens for requests up to 128K input tokens in US Virginia; $2.064 per 1M output tokens for 128K-256K requests
View model
→
Efficient multimodal assistants, coding, reasoning, long-context analysis, local deployment, and tool-using agents
Type
Multimodal
Context
262K
Reasoning
8/10
Speed
9/10
Multimodal
Image input
Video input
Status
Available; open-weight Apache 2.0 model with hosted API access through Alibaba Cloud Model Studio
Input
USD 0.057 per 1 million input tokens for up to 128K input; USD 0.229 per 1 million input tokens for 128K–256K input in the Global deployment scope. International flat pricing is USD 0.25 per 1 million input tokens.
Output
USD 0.459 per 1 million output tokens for up to 128K input; USD 1.835 per 1 million output tokens for 128K–256K input in the Global deployment scope. International flat pricing is USD 2 per 1 million output tokens.
View model
→
Advanced multimodal reasoning, image and video understanding, coding, long-context analysis, document and chart interpretation, function-calling agents, and web-grounded workflows.
Type
Multimodal
Context
262K
Reasoning
9/10
Speed
8/10
Multimodal
Image input
Video input
Status
Current and available through Alibaba Cloud Model Studio; released globally on February 24, 2026.
Input
$0.115 per 1M tokens for input up to 128K; $0.287 per 1M tokens for input from 128K to 256K. Pricing varies by deployment scope and region.
Output
$0.917 per 1M tokens for requests up to 128K input; $2.294 per 1M tokens for requests from 128K to 256K input. Pricing varies by deployment scope and region.
View model
→
Advanced multimodal reasoning, coding, video and image understanding, long-context analysis, tool-using agents, and self-hosted open-weight deployments
Type
Multimodal
Context
262K
Reasoning
9/10
Speed
6/10
Multimodal
Image input
Video input
Status
Current; open-weight model and available through Alibaba Cloud Model Studio
Input
$0.172 per 1M tokens for input up to 128K; $0.43 per 1M tokens for input above 128K and up to 256K in Beijing, Frankfurt, and Virginia. Singapore: $0.60 per 1M tokens.
Output
$1.032 per 1M tokens for input up to 128K; $2.58 per 1M tokens for input above 128K and up to 256K in Beijing, Frankfurt, and Virginia. Singapore: $3.60 per 1M tokens.
View model
→
Fast long-context text, image and video understanding; structured extraction; tool-enabled agents; web-grounded applications; and high-volume multimodal workloads.
Type
Multimodal
Context
1M
Reasoning
8/10
Speed
9/10
Multimodal
Image input
Video input
Status
Current and available
Input
$0.029 per 1M tokens for 0–128K input; $0.115 per 1M tokens for 128K–256K input; $0.172 per 1M tokens for 256K–1M input in US Virginia/Global pricing
Output
$0.287 per 1M tokens for 0–128K input; $1.147 per 1M tokens for 128K–256K input; $1.72 per 1M tokens for 256K–1M input in US Virginia/Global pricing
View model
→
Long-context reasoning, multimodal document and video analysis, coding, structured enterprise automation, function-calling agents, and web-grounded research.
Type
Multimodal
Context
1M
Reasoning
8/10
Speed
8/10
Multimodal
Image input
Video input
Status
Current; canonical rolling model identifier currently equivalent to qwen3.5-plus-2026-02-15
Input
International: $0.40 per 1M input tokens for input up to 256K; $0.50 per 1M input tokens for input above 256K and up to 1M. Regional and global pricing varies.
Output
International: $2.40 per 1M output tokens for input up to 256K; $3.00 per 1M output tokens for input above 256K and up to 1M. Regional and global pricing varies.
View model
→
OCR, document parsing, text localization, table extraction, handwritten-text recognition, and key information extraction from images
Type
Multimodal
Context
66K
Reasoning
4/10
Speed
7/10
Image input
Status
Current and available through Alibaba Cloud Model Studio
Input
0.069 USD per 1 million tokens in China (Beijing)
Output
0.275 USD per 1 million tokens in China (Beijing)
Model page unavailable
Fast multimodal analysis, long audio understanding, audiovisual question answering, voice assistants and spoken-response applications
Type
Multimodal
Context
262K
Reasoning
7/10
Speed
9/10
Multimodal
Image input
Audio input
Status
Current; canonical model ID functionally equivalent to qwen3.5-omni-flash-2026-03-15
Input
Singapore international: $0.40 per 1M tokens for text/image/video input and $3.00 per 1M tokens for audio input; China Beijing: $0.30 per 1M tokens for text/image/video input and $2.48 per 1M tokens for audio input
Output
Singapore international: $2.20 per 1M tokens for text output and $11.90 per 1M tokens for text-and-audio output; China Beijing: $1.83 per 1M tokens for text output and $9.90 per 1M tokens for text-and-audio output
View model
→
Low-latency voice assistants, speech-to-speech applications, realtime multimedia analysis, interactive agents, and multimodal conversations.
Type
Multimodal
Context
262K
Reasoning
6/10
Speed
9/10
Multimodal
Image input
Audio input
Status
Current and available; rolling model identity functionally equivalent to qwen3.5-omni-flash-realtime-2026-03-15
Input
International: $0.55 per 1M tokens for text/image/video input; $4.50 per 1M tokens for audio input. China mainland: $0.45 per 1M tokens for text/image/video input; $3.71 per 1M tokens for audio input.
Output
International: $3.30 per 1M tokens for text output; $17.70 per 1M tokens for audio output. China mainland: $2.75 per 1M tokens for text output; $14.71 per 1M tokens for audio output.
View model
→
Multilingual voice assistants, speech-enabled multimodal applications, audio-visual analysis, spoken explanations, accessibility tools, and interactive media workflows.
Type
Multimodal
Context
262K
Reasoning
7/10
Speed
7/10
Multimodal
Image input
Audio input
Input
$1.40 per 1M tokens for text/image/video input; $11.00 per 1M tokens for audio input in international deployments. China (Beijing) snapshot pricing is $0.96 per 1M tokens for text/image/video input and $7.29 per 1M tokens for audio input.
Output
$8.30 per 1M tokens for text output; $44.00 per 1M tokens for text-and-audio output in international deployments. China (Beijing) snapshot pricing is $5.50 per 1M tokens for text output and $29.29 per 1M tokens for text-and-audio output.
View model
→
Real-time voice assistants, speech-to-speech applications, multimodal customer service, visual conversational agents, live multimedia analysis, and interactive applications requiring controllable speech output.
Type
Multimodal
Context
262K
Reasoning
6/10
Speed
9/10
Multimodal
Image input
Audio input
Status
Current and available; canonical model identifier functionally equivalent to snapshot qwen3.5-omni-plus-realtime-2026-03-15
Input
China (Beijing): $1.38 per 1M text/image/video tokens; $11 per 1M audio tokens. Singapore: $2.10 per 1M text/image/video tokens; $16.50 per 1M audio tokens.
Output
China (Beijing): $8.25 per 1M text-output tokens; $41.26 per 1M text-and-audio-output tokens. Singapore: $12.40 per 1M text-output tokens; $62 per 1M text-and-audio-output tokens.
View model
→
Coding agents, repository-level software engineering, visual document analysis, video understanding, STEM reasoning, long-context assistants, and self-hosted multimodal applications
Type
Multimodal
Context
262K
Reasoning
8/10
Speed
7/10
Multimodal
Image input
Video input
Status
Current and available; open-weight release and Alibaba Cloud Model Studio API model
Input
$0.60 per 1M tokens internationally; $0.412564 per 1M tokens in China Beijing and Singapore
Output
$3.60 per 1M tokens internationally; $2.475384 per 1M tokens in China Beijing and Singapore
View model
→
Agentic coding, long-context software engineering, multimodal analysis, tool-using agents, and self-hosted deployments
Type
Multimodal
Context
262K
Reasoning
8/10
Speed
8/10
Multimodal
Image input
Video input
Status
Current; open-weight and available through Alibaba Cloud Model Studio
Input
USD 0.248 per 1M tokens in US Virginia, Germany Frankfurt, and China Beijing; USD 0.375 per 1M tokens in Singapore
Output
USD 1.485 per 1M tokens in US Virginia, Germany Frankfurt, and China Beijing; USD 2.25 per 1M tokens in Singapore
View model
→
Fast multimodal assistants, coding agents, visual document analysis, video understanding, tool-using workflows, object localization, and large-context applications
Type
Multimodal
Context
1M
Reasoning
8/10
Speed
9/10
Multimodal
Image input
Video input
Input
$0.25 per 1M input tokens up to 256K input tokens; $1.00 per 1M input tokens above 256K and up to 1M, Singapore international pricing
Output
$1.50 per 1M output tokens up to 256K input tokens; $4.00 per 1M output tokens above 256K and up to 1M, Singapore international pricing
View model
→
Advanced coding agents, front-end development, long-context analysis, structured API workflows, and text generation with web search
Type
General Purpose
Context
262K
Reasoning
8/10
Speed
7/10
Tool use
Web search
Streaming
Status
Preview; currently accessible; scheduled for deprecation on October 10, 2026
Input
$1.30 per 1M tokens for input up to 128K; $2.00 per 1M tokens for input above 128K and up to 256K, international pricing
Output
$7.80 per 1M tokens for requests up to 128K input; $12.00 per 1M tokens for requests above 128K and up to 256K input, international pricing
View model
→
Long-context multimodal analysis, agentic coding, OCR, object localization, frontend development, visual reasoning, and tool-enabled enterprise assistants
Type
Multimodal
Context
1M
Reasoning
9/10
Speed
7/10
Multimodal
Image input
Video input
Status
Current; rolling model ID qwen3.6-plus is currently equivalent to qwen3.6-plus-2026-04-02
Input
$0.276 per 1M tokens for 0-256K input tokens; $1.101 per 1M tokens for 256K-1M input tokens in Global deployment
Output
$1.651 per 1M tokens for 0-256K input tokens; $6.602 per 1M tokens for 256K-1M input tokens in Global deployment
View model
→
Fast multimodal agents, visual coding, tool-use workflows, search agents, and long-context document or screen analysis
Type
Multimodal
Context
1M
Reasoning
8/10
Speed
9/10
Multimodal
Image input
Video input
Status
Current and available; rolling qwen3.7-flash identifier is currently equivalent to qwen3.7-flash-2026-07-15
Input
Global: $0.028 per 1M input tokens for 0-32K; $0.083 for 32K-256K; $0.165 for 256K-1M. Singapore international: $0.030, $0.100, and $0.200 per 1M input tokens respectively.
Output
Global: $0.110 per 1M output tokens for 0-32K; $0.330 for 32K-256K; $0.660 for 256K-1M. Singapore international: $0.130, $0.400, and $0.800 per 1M output tokens respectively.
View model
→
Complex reasoning, advanced coding, long-context analysis, tool-using agents, productivity automation, and long-horizon task execution
Type
Reasoning
Context
1M
Reasoning
9/10
Speed
7/10
Tool use
Web search
Status
Legacy; currently accessible in supported Alibaba Cloud Model Studio deployments
Input
$1.65 per 1M tokens in the US Virginia Global scope; $2.50 per 1M tokens in the US scope
Output
$4.951 per 1M tokens in the US Virginia Global scope; $7.50 per 1M tokens in the US scope
Model page unavailable
Long-context reasoning, multimodal document and video analysis, coding, tool-using agents, structured extraction, and enterprise productivity workflows
Type
Multimodal
Context
1M
Reasoning
9/10
Speed
7/10
Multimodal
Image input
Video input
Status
Current; rolling identifier currently equivalent to qwen3.7-plus-2026-05-26
Input
$0.40 per 1M tokens up to 256K input tokens and $1.20 per 1M tokens from 256K to 1M input tokens in US Virginia; regional and promotional pricing varies
Output
$1.60 per 1M tokens up to 256K input tokens and $4.80 per 1M tokens from 256K to 1M input tokens in US Virginia; thinking-mode output uses the same listed rates
View model
→
Multilingual semantic search, retrieval-augmented generation, code retrieval, recommendation, clustering, classification, and large-scale text vectorization
Type
Embedding
Context
131K
Speed
8/10
Status
Current and available through Alibaba Cloud Model Studio internationally
Input
$0.07 per 1 million input tokens in Singapore international deployment; China pricing may differ by region and billing mode
View model
→
Qwen3.7-Text-Embedding
qwen3.7-text-embedding-flash
Cost-sensitive, high-throughput multilingual text vectorization, semantic search, RAG, recommendation, clustering, and classification.
Type
Lightweight
Context
131K
Reasoning
1/10
Speed
9/10
Input
CNY 0.000125 per 1,000 input tokens in China (Beijing); batch pricing CNY 0.000063 per 1,000 input tokens. Regional pricing may differ.
Model page unavailable
Advanced reasoning, coding, scientific and professional research, long-context analysis, and long-horizon agent workflows
Type
Reasoning
Context
1M
Reasoning
10/10
Speed
4/10
Tool use
Web search
Streaming
Status
Current and available; open-weight release and hosted API model
Input
$2 per 1 million tokens internationally; $1.65 per 1 million tokens in China (Beijing)
Output
$6 per 1 million tokens internationally; $4.951 per 1 million tokens in China (Beijing)
View model
→
Coding assistants, repository analysis, visual document workflows, long-context research, office automation, multimodal agents, and tool-using applications
Type
Multimodal
Context
1M
Reasoning
8/10
Speed
6/10
Multimodal
Image input
Video input
Status
Active and currently available; open-weight release and Alibaba Cloud Model Studio API model
Input
$0.424 per 1M tokens in China Beijing; $0.50 per 1M tokens in Singapore. Implicit cache input is $0.085 per 1M tokens in Beijing and $0.10 per 1M tokens in Singapore. Explicit cache creation is $0.53 per 1M tokens in Beijing and $0.625 per 1M tokens in Si
Output
$1.696 per 1M tokens in China Beijing; $3 per 1M tokens in Singapore.
View model
→
Fast long-context reasoning, coding assistance, visual document and chart analysis, video understanding, function-calling agents, and high-concurrency applications
Type
Multimodal
Context
1M
Reasoning
8/10
Speed
9/10
Multimodal
Image input
Video input
Status
Current and accessible through Alibaba Cloud Model Studio
Input
$0.113 per 1 million input tokens; implicit cache input $0.014 per 1 million tokens; explicit cache creation $0.177 per 1 million tokens; explicit cache read $0.014 per 1 million tokens
Output
$0.382 per 1 million output tokens
View model
→
High-volume coding, agentic workflows, long-context document and codebase analysis, office automation, and multimodal text-image-video understanding
Type
Multimodal
Context
262K
Reasoning
8/10
Speed
9/10
Multimodal
Image input
Video input
Status
Current open-weight experimental preview
View model
→
Complex coding, autonomous software engineering, long-horizon agent workflows, professional document analysis, visual reasoning, long videos, and demanding research tasks.
Type
Multimodal
Context
1M
Reasoning
9/10
Speed
7/10
Multimodal
Image input
Video input
Status
Stable official release; currently available through Alibaba Cloud Model Studio
Input
US$1.65 per 1M input tokens and US$2.00 per 1M input tokens for International scope; regional pricing varies. Cached-input pricing starts at US$0.206 per 1M tokens for the listed regional/global scope and US$0.25 per 1M tokens for International scope.
Output
US$4.951 per 1M output tokens for the listed regional/global scope and US$6.00 per 1M output tokens for International scope.
View model
→
Real-time speech translation, multilingual meetings, live interpretation, translated voice communication, and audiovisual translation with low latency.
Type
Multimodal
Context
53K
Reasoning
4/10
Speed
9/10
Multimodal
Image input
Audio input
Status
Current stable model
Input
Singapore: audio input $7.50 per 1 million tokens; image input $0.55 per 1 million tokens. China (Beijing): audio input $5.653 per 1 million tokens; image input $0.466 per 1 million tokens.
Output
Singapore: text output $20 per 1 million tokens; audio output $30 per 1 million tokens. China (Beijing): text output $14.133 per 1 million tokens; audio output $22.613 per 1 million tokens.
View model
→
Long-form audio and video understanding, multimedia analysis, audio-visual agents, content summarization, and tool-using workflows
Type
Multimodal
Context
1M
Reasoning
8/10
Speed
8/10
Multimodal
Image input
Audio input
Status
Current and available
Input
USD 0.15 per 1 million input tokens for International deployment; USD 0.016 per 1 million cache-hit input tokens
Output
USD 0.47 per 1 million output tokens for International deployment
View model
→
Real-time voice assistants, speech-to-speech applications, interactive video agents, live media analysis, multimodal customer service, meeting and collaboration interfaces, and applications requiring tool or MCP integration.
Type
Multimodal
Context
197K
Reasoning
7/10
Speed
9/10
Multimodal
Audio input
Video input
Status
Current and available; international deployment in Singapore
Input
China (Beijing): CNY 1.5 per 1 million tokens for text/images/video input and CNY 6 per 1 million tokens for audio input. Singapore: CNY 1.677 per 1 million tokens for text/images/video input and CNY 6.781 per 1 million tokens for audio input.
Output
China (Beijing): CNY 4.5 per 1 million tokens for text output and CNY 12 per 1 million tokens for audio output. Singapore: CNY 5.104 per 1 million tokens for text output and CNY 13.636 per 1 million tokens for audio output.
View model
→
Qwen3Guard
Qwen3Guard-Stream-0.6B
Low-latency local moderation of streaming prompts and language-model responses
Type
Other
Context
33K
Reasoning
2/10
Speed
8/10
Streaming
Status
Current open-weight model
Model page unavailable
Real-time moderation of streamed prompts and language-model responses
Type
Other
Context
8K
Reasoning
3/10
Speed
8/10
Streaming
Status
Available open-weight model
View model
→
Qwen3Guard
Qwen3Guard-Stream-8B
Real-time multilingual moderation of prompts and streamed language-model responses
Type
Other
Context
8K
Reasoning
1/10
Speed
7/10
Fine-tuning
Streaming
Status
Current open-weight model
Model page unavailable
Low-latency character dialogue, virtual companions, game NPCs, role-playing applications, IP character replication, and conversational smart devices
Type
Lightweight
Context
33K
Reasoning
4/10
Speed
9/10
Web search
Streaming
Status
Current; dynamically updated managed model
Input
Regional pricing: USD 0.034 per 1M input tokens in Beijing and US Virginia; USD 0.05 per 1M input tokens in Singapore. Cached input is listed at USD 0.007 per 1M tokens in Beijing and US Virginia and USD 0.01 in Singapore. Alibaba Cloud also publishes loc
Output
Regional pricing: USD 0.203 per 1M output tokens in Beijing and US Virginia; USD 0.40 per 1M output tokens in Singapore. Alibaba Cloud also publishes localized regional prices.
View model
→
Qwen Character
qwen-plus-character
Character role-play, virtual social applications, game NPCs, IP character replication, smart toys, in-car assistants, and empathetic conversational experiences
Type
Other
Context
33K
Reasoning
3/10
Speed
8/10
Web search
Status
Current; dynamically updated model
Input
$0.50 per 1 million tokens internationally; $0.115 per 1 million tokens in China (Beijing), Hong Kong, Germany, the United States, and Japan
Output
$1.40 per 1 million tokens internationally; $0.287 per 1 million tokens in China (Beijing), Hong Kong, Germany, the United States, and Japan
Model page unavailable
Complex text-to-image layouts, multilingual typography, posters, menus, storyboards, interface mockups, product visuals, and image editing with one to three reference images
Type
Image Generation
Context
5K
Reasoning
2/10
Speed
6/10
Multimodal
Image input
Media output
Status
Current and available
Input
$0.003 per image in the international Singapore deployment
Output
$0.04 per 1K image or $0.075 per 2K image in the international Singapore deployment; regional pricing varies
View model
→
Instruction-based image editing, multi-image fusion, character-consistent compositions, object replacement, style transfer, poster and text editing, and product-image variation.
Type
Image Editing
Reasoning
5/10
Speed
8/10
Multimodal
Image input
Media output
Status
Current; canonical model currently equivalent to qwen-image-edit-plus-2025-10-30
Input
0; image editing is billed by output image
Output
$0.028671 per image in China (Beijing); $0.03 per image in Singapore/international
View model
→
Realistic text-to-image generation, creative concepts, marketing visuals, editorial artwork, product concepts, and general-purpose image creation.
Type
Image Generation
Reasoning
1/10
Speed
5/10
Media output
Status
Current and available for text-to-image generation; the undated qwen-image-max identifier is currently equivalent to qwen-image-max-2025-12-30.
Input
No separate input charge; text prompt input is included in per-image billing.
Output
$0.075/image in Singapore international deployment; $0.071677/image in China (Beijing).
View model
→
Fast text-to-image generation, text-heavy layouts, image editing, reference-image compositing, marketing graphics, and high-volume image workflows.
Type
Other
Reasoning
3/10
Speed
8/10
Multimodal
Image input
Media output
Status
Current and available through Alibaba Cloud Model Studio
Input
$0.003 per input image internationally; regional pricing varies
Output
$0.03 per generated image internationally; regional pricing varies
View model
→
Mathematics, coding, research-style analysis, and complex multi-step reasoning
Type
Reasoning
Context
131K
Reasoning
9/10
Speed
6/10
Tool use
Web search
Streaming
Status
Legacy; currently accessible as of October 7, 2026; scheduled to cease service on October 10, 2026 at 00:00 UTC+08, subject to the actual change time
Input
$0.230 per 1M tokens in China (Beijing); $0.80 per 1M tokens in Singapore international deployment
Output
$0.574 per 1M tokens in China (Beijing); $2.40 per 1M tokens in Singapore international deployment
Model page unavailable
Tongyi Embedding Vision
tongyi-embedding-vision-plus
Cross-modal retrieval, image and video similarity search, multimodal semantic indexing, recommendation, and content classification
Type
Multimodal Embedding
Reasoning
1/10
Speed
8/10
Multimodal
Image input
Video input
Input
$0.09 per 1 million input tokens for text, image, and video in the international Singapore deployment
Output
$0; embedding output is free
Model page unavailable
Model page unavailable