Qwen-Image

Qwen-Image-2.1

by Qwen · Current open-weight release

Open-weight Qwen image model for text-to-image generation, editing with up to 10 reference images, transparent RGBA creation, subject extraction, and localized mask-based changes.

Image generation
Qwen-Image-2.1 combines image generation and image editing in one downloadable model. It accepts English or Chinese prompts, can work with up to 10 reference images, supports transparent RGBA output, and enables targeted edits through masks or visual annotations. Its open-weight distribution is aimed at local and self-hosted workflows rather than a directly hosted Qwen-Image-2.1 API endpoint.
Outputs

What Qwen-Image-2.1 can produce

Image generation
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Multimodal output
Model profile

Performance characteristics

7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Qwen-Image
Model type Multimodal
Release date 2026-09-21
Status Current open-weight release
Model notes

Qwen-Image-2.1 is a downloadable open-weight model with a 7B visual generation component. It supports text-to-image generation, image-to-image editing, transparent RGBA output, subject extraction, localized edits using masks or annotations, and up to 10 reference images. The official model card identifies the license as the Qwen Research License Agreement. Alibaba Cloud Model Studio separately offers qwen-image-2.1-pro; that hosted API model is a distinct model identity and its per-image pricing should not be assigned to Qwen-Image-2.1. The official Hugging Face model page states that the model is not deployed by an Inference Provider.

Model guide

Qwen-Image-2.1: Open-Weight Image Generation with Multi-Reference Editing

Qwen-Image-2.1 is Alibaba's open-weight 7B visual generation model for text-to-image creation, multi-image editing, transparent RGBA assets, subject extraction, and localized image changes.

What is Qwen-Image-2.1?

Qwen-Image-2.1 is an open-weight image generation and editing model from Alibaba's Qwen team. Its visual generation component has 7 billion parameters. The model is designed for people who need to create or modify images locally, rather than use only a managed image-generation service.

It supports two closely related workflows. In text-to-image generation, a prompt describes a new image. In image-to-image editing, one or more existing images provide visual information that the model should preserve, combine, or modify. This makes Qwen-Image-2.1 more suitable for iterative creative work than a generation-only system.

The model can produce illustrations, product imagery, posters, infographics, storyboards, panoramas, and other designed compositions. The official model materials also describe support for English and Chinese prompts.

Where Qwen-Image-2.1 fits in the Qwen lineup

Qwen-Image-2.1 is the open-weight model identity described by the Qwen model repository and related technical documentation. It is distributed as downloadable weights for local inference through compatible tooling such as Hugging Face Diffusers.

Alibaba Cloud Model Studio separately lists qwen-image-2.1-pro. That is a hosted API model identifier, not the same canonical entity as the downloadable Qwen-Image-2.1 model. Hosted API pricing or service behavior for qwen-image-2.1-pro should therefore not be treated as pricing or availability information for Qwen-Image-2.1.

This distinction matters in practice. Qwen-Image-2.1 gives a developer responsibility for downloading model files, providing suitable hardware, installing compatible software, and managing inference. A hosted service handles much of that operational work but is a separate product choice.

What can Qwen-Image-2.1 do?

Text-to-image generation

Qwen-Image-2.1 can generate an image from a written prompt. Its documented use cases include creative illustrations, product visuals, posters, infographics, storyboards, panoramas, and other compositions where layout and visual content need to be specified together.

Prompt-based generation is useful for creating a first draft, but the model's broader value is its ability to continue working with images after that first draft. A team can generate a concept, provide the result as an input, and make focused changes instead of starting over for every revision.

Multi-image editing and composition

The model supports up to 10 reference images in one editing workflow. Multiple references can provide separate subjects, products, people, or environmental elements that should appear in a single composition.

For example, several product photographs could be combined into a promotional scene, or multiple people and visual assets could be brought together for a group composition. Other documented use cases include virtual try-on and interior design. The reference-image limit is a practical advantage for workflows that would otherwise require repeated, separate edits.

Transparent RGBA generation

Qwen-Image-2.1 supports native RGBA image generation. RGBA adds an alpha channel to the usual red, green, and blue channels, allowing pixels to carry transparency information. The result can be used as a cutout or layered asset without first removing a background in another application.

The model can also edit transparent layers and extract a subject from an RGB photograph as an RGBA asset. This is particularly relevant to e-commerce graphics, character assets, compositing, packaging concepts, and design systems that reuse subjects across multiple backgrounds.

Localized edits with masks and annotations

Localized editing lets the user identify the part of an image that should change while leaving the rest of the composition intact. Qwen-Image-2.1 supports circles, painted annotations, and separate masks as ways to indicate an editing region.

Possible edits include replacing clothing, changing hair, removing an object, inserting a new subject, or modifying text in a defined area. Masks are useful when an external tool can precisely define a region; painted annotations and circles can be more convenient during exploratory work. In either case, the workflow is intended to reduce unwanted changes outside the selected area.

Technical profile and output considerations

The visual generation component uses 32 single-stream DiT layers. DiT refers to a diffusion-transformer architecture, in which transformer layers help process the representation used during image generation. The model also uses mixed-granularity attention and key-value cache reuse. According to the technical description, these design choices are intended to reduce memory use and improve efficiency when multiple reference images are processed.

Official examples document these output aspect ratios: 1:1, 4:3, 3:4, 3:2, 2:3, 16:9, and 9:16. Example resolutions reach up to 2752 by 2752 pixels, depending on the selected aspect ratio. The exact feasible resolution and generation speed will depend on the inference configuration and available GPU hardware.

The supplied specifications do not provide a language-model-style context window, maximum token output, or a universal generation-time guarantee. Those limits should not be inferred from the model's 7B parameter count. Image dimensions, number of steps, reference images, and hardware can all affect memory requirements and throughput.

Supported modalities and workflows

Qwen-Image-2.1 is an image model with text input and image input. Text prompts can request a new image, while existing images can be supplied for editing and composition. Its direct output is visual: generated or edited images, including transparent RGBA assets.

It is not a general-purpose text model. It does not provide text generation, coding, speech, video generation, embeddings, or general tool and function calling as core capabilities. It also should not be evaluated as a reasoning model using language-model reasoning or coding benchmarks.

The model is intended for inference workflows rather than a built-in agent platform. There is no verified claim in the supplied research that it provides structured JSON output, streaming responses, web search, or managed batch processing.

Pricing, license, and availability

Qwen-Image-2.1 has no verified recurring API price in the supplied research. It is distributed as downloadable open-weight model files through the official Qwen organization on Hugging Face. Downloading the weights is different from the total cost of operating the model: local users still need suitable GPU capacity, storage, compatible libraries, electricity, and engineering time.

The official model card identifies the license as the Qwen Research License Agreement. Anyone considering commercial deployment, redistribution, or a hosted service should review the license terms directly rather than assuming that open-weight availability means unrestricted commercial use.

Local deployment is documented through compatible diffusion tooling, including Diffusers. The official Hugging Face model page states that Qwen-Image-2.1 is not deployed by an Inference Provider. This means users should not assume that a one-click hosted endpoint is available under the model's canonical name.

Strengths and limitations

Strengths

  • One model for creation and editing: the same model supports new image generation, reference-based editing, and targeted changes.
  • Large reference-image allowance: up to 10 reference images can be used in one workflow, which helps with multi-subject composition and product-oriented tasks.
  • Transparency support: native RGBA generation and subject extraction can produce reusable assets without a separate background-removal step.
  • Local control: downloadable weights can suit teams that need self-hosting, data locality, workflow customization, or independence from a hosted API.
  • Useful composition formats: documented aspect ratios cover common portrait, landscape, square, and widescreen layouts.

Limitations

  • Operational burden: local inference requires appropriate hardware, model storage, software compatibility, and performance tuning.
  • No canonical hosted endpoint: the open-weight model is not the same as Alibaba Cloud's separately named qwen-image-2.1-pro API model.
  • No general language or agent abilities: it is not designed for text generation, coding, tool calls, speech, video, or embeddings.
  • Hardware-dependent performance: the supplied research does not establish a universal speed, memory, or resolution guarantee for every deployment.
  • License review is necessary: the Qwen Research License Agreement should be checked before production or commercial use.

When to choose Qwen-Image-2.1

Choose Qwen-Image-2.1 when the priority is local image generation with meaningful editing control. It is a strong fit for teams that need transparent assets, multi-reference composition, subject extraction, or mask-based revisions. It is also appropriate when data should remain within a self-managed environment or when developers want to customize the inference pipeline instead of relying on a fixed hosted interface.

The model is especially relevant to product design, e-commerce imagery, visual storytelling, posters, concept development, virtual try-on experiments, and interior-design visualization. Its up-to-10-image reference workflow can reduce the need to manually merge several assets before generation.

A hosted image API may be more appropriate when the main requirement is quick integration, predictable operational support, or avoiding GPU and model-management responsibilities. In that case, Alibaba Cloud's separately identified qwen-image-2.1-pro should be evaluated as its own service rather than treated as a direct deployment of this open-weight model.

A different model type is more suitable for applications centered on text, code, speech, video, embeddings, or autonomous tool use. Qwen-Image-2.1 should be selected for its visual generation and editing workflow, not as a general-purpose AI assistant.

Bottom line

Qwen-Image-2.1 is a downloadable 7B visual generation model that combines text-to-image creation with practical editing capabilities. Its most distinctive documented features are support for up to 10 reference images, transparent RGBA output, subject extraction, and localized edits using masks or annotations. Those capabilities make it more useful for iterative design and asset production than a model limited to creating isolated images from text.

Its trade-off is responsibility: users must provide the infrastructure and manage the local inference stack, and there is no verified canonical hosted endpoint or recurring price for this model in the supplied information. For teams that value local control and image-editing flexibility, that trade-off may be worthwhile. For teams that need a managed API or non-visual AI functions, another option will be more appropriate.


Answers to Frequently Asked Questions

What is Qwen-Image-2.1?
Qwen-Image-2.1 is an open-weight image generation and editing model from Alibaba's Qwen team. Its visual generation component has 7 billion parameters and supports text-to-image generation, image-to-image editing, multi-reference composition, and localized image changes.
How many reference images can Qwen-Image-2.1 use?
Qwen-Image-2.1 supports up to 10 reference images in a single editing workflow. These images can provide subjects, products, people, or other visual elements for creating a combined composition.
Does Qwen-Image-2.1 support transparent images?
Yes. Qwen-Image-2.1 supports native RGBA image generation, which includes an alpha channel for transparency. It can also edit transparent layers and extract subjects from RGB photographs as reusable RGBA assets.
Is Qwen-Image-2.1 the same as qwen-image-2.1-pro?
No. Qwen-Image-2.1 is a downloadable open-weight model intended for local inference, while qwen-image-2.1-pro is a separately identified hosted API model listed by Alibaba Cloud Model Studio. Pricing and service behavior for qwen-image-2.1-pro should not be assumed to apply to Qwen-Image-2.1.
What are the licensing and deployment requirements for Qwen-Image-2.1?
Qwen-Image-2.1 is distributed as downloadable model files through the official Qwen organization on Hugging Face and can be deployed locally with compatible tooling such as Diffusers. The model card identifies the Qwen Research License Agreement, so the terms should be reviewed before commercial use, redistribution, or hosted deployment. Local users also need suitable GPU capacity, storage, compatible software, and engineering resources.


Sources 3
Provider

About Qwen