Qwen Image

Qwen-Image-Max

by Qwen · Current and available for text-to-image generation; the undated qwen-image-max identifier is currently equivalent to qwen-image-max-2025-12-30.

Qwen-Image-Max is Alibaba Cloud Model Studio’s hosted text-to-image model for realistic still-image generation. It returns one PNG image per request at up to 1664×928, costs $0.075 per image in the international Singapore deployment, and does not support image editing or image inputs.

Image generation Reasoning Coding
Qwen-Image-Max is a hosted text-to-image model from Alibaba Cloud Model Studio. It is aimed at users who want realistic, natural-looking still images from descriptive prompts rather than conversational text, coding, or image-editing capabilities. The model produces one PNG image per request at a maximum listed resolution of 1664×928, with regional pricing charged per generated image.
Outputs

What Qwen-Image-Max can produce

Image generation
Inputs

What it can understand

Text
Capabilities

Supported features

Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
5/10 Speed
4/10 Cost efficiency
Specifications

Technical details

Model family Qwen Image
Model type Image Generation
Release date 2025-12-30
Status Current and available for text-to-image generation; the undated qwen-image-max identifier is currently equivalent to qwen-image-max-2025-12-30.
Knowledge cutoff notes

Alibaba Cloud's public model documentation does not provide a separate knowledge-cutoff date for Qwen-Image-Max. As an image-generation model, the field is not directly applicable in the same way as it is for language models.

Model notes

The canonical current model ID is qwen-image-max. Alibaba Cloud states that it is currently equivalent to qwen-image-max-2025-12-30. The model supports text-to-image generation only, returns one PNG image per request, and lists a maximum resolution of 1664×928. It does not support image editing. Pricing is per generated image and varies by region. The Singapore international deployment is listed at $0.075/image, while China (Beijing) is listed at $0.071677/image. The international deployment has a documented rate limit of two requests per minute. Editorial scores are capability-oriented estimates for an image-generation model; reasoning and coding scores are not meaningful measures of its intended use.

Cost

Model pricing

Input No separate input charge; text prompt input is included in per-image billing.
Output $0.075/image in Singapore international deployment; $0.071677/image in China (Beijing).
Model guide

Qwen-Image-Max: Realistic Text-to-Image Generation Without Editing

Qwen-Image-Max is Alibaba Cloud Model Studio’s hosted Qwen Max image-generation model for producing one realistic PNG image from a text prompt. It supports text-to-image generation at up to 1664×928 resolution, is billed per generated image, and does not provide image editing or multi-image workflows.

What is Qwen-Image-Max?

Qwen-Image-Max is Alibaba Cloud Model Studio’s hosted model for generating still images from text prompts. A user supplies a natural-language description, and the service returns one generated PNG image. The model is part of the Qwen Image family and is positioned as a higher-quality option for realistic and natural-looking image generation.

The current canonical model identifier is qwen-image-max. Alibaba Cloud documentation lists this undated identifier as equivalent to the dated identifier qwen-image-max-2025-12-30. For applications using the provider’s current model catalog, the undated identifier is the preferred model name.

Unlike a general-purpose Qwen assistant, Qwen-Image-Max is not intended to answer questions, write code, return structured text, or perform tool calls. Its primary output is a generated image.

Core capabilities and output limits

Qwen-Image-Max supports text-to-image generation only. The provider’s documented limits make it a relatively straightforward prompt-to-image endpoint: one text prompt goes in, and a single PNG image comes out.

SpecificationDocumented behavior
Model identifierqwen-image-max
Primary taskText-to-image generation
Text inputSupported
Image inputNot supported for this model
OutputOne PNG image
Maximum output countOne image per request
Maximum listed resolution1664×928
Image editingNot supported
Billing unitGenerated image

The resolution limit and single-output behavior matter when planning a production workflow. A request cannot produce a set of variations in one call, and the model is not documented as supporting image-to-image transformation, inpainting, or other editing operations. If an application needs several alternatives, it must make separate generation requests.

Image quality and positioning

Qwen-Image-Max emphasizes realism and naturalness. In practical terms, it is best understood as a model for creating polished still-image concepts from written descriptions, with an emphasis on reducing visibly artificial results. These are provider positioning claims rather than a supplied independent benchmark result, so actual quality can vary with the prompt, subject, composition, and use case.

Within the Qwen Image lineup, Qwen-Image-Max is a hosted Max-tier generation option rather than a downloadable open-weight model. The supplied documentation describes newer Qwen Image 2.0 and 3.0 models as having advantages for more advanced generation and editing workflows. That makes Qwen-Image-Max most relevant when a simple realistic text-to-image request is more important than the broadest feature set or newest image capabilities.

Supported modalities and missing features

The model accepts text and returns an image. It does not natively return text, audio, video, embeddings, executable actions, or structured text responses. It also does not accept an image, audio clip, or video as an input according to the supplied model record.

Several familiar language-model features are therefore not applicable:

  • Reasoning: No separate reasoning mode or reasoning-token capability is documented.
  • Coding: Code generation is not a supported purpose of the model.
  • Tool use: Function calling, web search, and external actions are not supported model outputs.
  • Streaming: Streaming text generation is not relevant to this image-only output model.
  • JSON mode: The model does not generate structured JSON as its primary output.
  • Context length: No model context-window value is published in the supplied documentation.
  • Maximum output tokens: This is not applicable because the primary output is an image rather than generated text.

These limitations do not indicate a defect for the intended use. They simply distinguish Qwen-Image-Max from multimodal assistants and image models designed for editing or reference-image workflows.

Pricing and availability

Alibaba Cloud Model Studio lists Qwen-Image-Max at $0.075 per generated image for the international Singapore deployment. The China (Beijing) deployment is listed at $0.071677 per generated image. The model is billed by generated image rather than by input or output tokens, and the supplied research indicates that the text prompt is included in the per-image charge.

Prices and availability are deployment-specific and may change. The international deployment is documented with a rate limit of two requests per minute. That limit can be important for applications that need to generate many images, because a single request returns only one image and repeated variations may require multiple calls.

Qwen-Image-Max is available as a hosted service through Alibaba Cloud Model Studio. The supplied research does not identify it as a downloadable open-weight model. Users should therefore evaluate regional availability, account requirements, quotas, endpoint configuration, and current pricing in the relevant Alibaba Cloud documentation before committing to a production integration.

Strengths and trade-offs

The model’s main strength is focus. It provides a simple text-to-image path for users who want a realistic still image without adding editing, reference-image, or conversational features that they may not need. Per-image billing also makes the basic cost unit easy to understand: each generated result is charged as one image.

Its principal trade-offs are output flexibility and throughput. One image per request limits batch experimentation, while the 1664×928 maximum may not meet every high-resolution production requirement. The two-requests-per-minute international rate limit can also make large variation sets slower to produce. These constraints matter more for automated creative pipelines than for occasional concept generation.

The editorial capability scores in the supplied model record rate speed and cost favorably relative to the model’s intended image-generation role, but they are internal evaluations rather than provider-published benchmark results. They should not be interpreted as standardized measurements or as guarantees of response time, image quality, or total project cost.

Best use cases

Qwen-Image-Max is a sensible choice for applications that need one realistic image from a written description, including:

  • Concept art and early visual development
  • Marketing and campaign visuals
  • Editorial illustrations
  • Product concepts and mood boards
  • Social-media graphics
  • General creative image generation

For example, a designer could use a detailed prompt to create an initial product concept or a mood-board image, then review the result manually. A content workflow that needs multiple alternatives can submit separate requests, but should account for the single-image output limit and the international rate limit.

When to choose Qwen-Image-Max

Choose Qwen-Image-Max when the requirement is primarily realistic text-to-image generation and a single PNG result at up to 1664×928 is sufficient. It is especially suitable when straightforward per-image pricing and a hosted Alibaba Cloud deployment are more important than editing or reference-image support.

Another option may be more appropriate in several situations. Choose an image-editing model when the workflow needs inpainting, image-to-image transformation, or controlled changes to an existing picture. Choose a newer Qwen Image model when higher resolution, multiple outputs, multi-reference composition, or more advanced generation and editing features are required. The supplied research specifically identifies Qwen Image 2.0 and 3.0 as newer alternatives for those broader workflows, although their exact current capabilities and prices should be checked separately.

For text answers, coding, web search, multimodal conversation, or structured application responses, use a suitable language or multimodal model instead. Qwen-Image-Max is not a general assistant and should not be selected merely because it belongs to the broader Qwen family.

Limitations to check before use

  • Only one PNG image is returned per request.
  • The maximum listed resolution is 1664×928.
  • Image input and image editing are not supported by this model.
  • No context length or text-output limit is published because the model is image-focused.
  • There is no documented reasoning, coding, tool-use, JSON, audio, or video output capability.
  • The international deployment has a documented limit of two requests per minute.
  • Regional pricing, availability, quotas, and model identifiers can change.

Overall, Qwen-Image-Max is best viewed as a focused hosted image generator rather than an all-purpose multimodal model. Its value comes from producing a single realistic image from text with clear per-image billing. Its narrower input and output support makes it less suitable for editing pipelines, high-volume variation generation, or applications that need text and image functions in the same model.


Answers to Frequently Asked Questions

What are the main limitations of Qwen-Image-Max?
Qwen-Image-Max does not support image inputs, editing, multiple outputs per request, text responses, coding, tool use, JSON mode, audio, or video. The international deployment also has a documented limit of two requests per minute.
How much does Qwen-Image-Max cost?
Alibaba Cloud Model Studio lists Qwen-Image-Max at $0.075 per generated image for the international Singapore deployment and $0.071677 per generated image for the China (Beijing) deployment. Prices and availability may change by region.
What image resolution and output limits does Qwen-Image-Max have?
The model returns one PNG image per request, with a maximum listed resolution of 1664×928. Applications that need multiple variations must submit separate generation requests.
Does Qwen-Image-Max support image editing or image-to-image generation?
No. Qwen-Image-Max supports text-to-image generation only. It does not accept image inputs and does not support image editing, inpainting, or image-to-image transformation.
What is Qwen-Image-Max used for?
Qwen-Image-Max is a hosted Alibaba Cloud model for generating one realistic still image from a natural-language text prompt. It is suitable for concept art, marketing visuals, editorial illustrations, product concepts, mood boards, and social-media graphics.


Sources 4
Provider

About Qwen