Qwen-Image-Edit

qwen-image-edit

by Qwen · Current and accessible

Alibaba's Qwen-Image-Edit is an image-to-image model for editing existing images with natural-language instructions. It supports semantic and appearance edits, object insertion or removal, style transfer, image fusion, and Chinese or English text changes. The model accepts supported image files with text input, returns one PNG image per request, and has a documented maximum resolution of 1024 by 1024 pixels. International Singapore pricing is listed at $0.045 per generated image.

Image generation Reasoning Coding
Qwen-Image-Edit is an image-editing model from Alibaba's Qwen team, provided through Alibaba Cloud Model Studio. Instead of requiring layer-based editing skills, it lets users describe changes in natural language alongside one or more image inputs. The original model is aimed at focused, single-image transformations and returns one edited PNG image per request.
Outputs

What qwen-image-edit can produce

Image generation
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Multimodal output
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
6/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family Qwen-Image-Edit
Model type Other
Release date 2025-08-19
Status Current and accessible
Knowledge cutoff notes

No authoritative model-specific knowledge cutoff is published. As an image-editing model, its relevant behavior is determined by its trained visual and text-editing capabilities rather than a documented conversational knowledge cutoff.

Model notes

The original qwen-image-edit model is distinct from qwen-image-edit-plus, qwen-image-edit-max, Qwen-Image-Edit-2509, and Qwen-Image-Edit-2511. Alibaba Cloud documentation lists text and image input, image output, one output image per request, and a fixed maximum resolution of 1024x1024. The current image-model catalog lists qwen-image-edit as supported for editing but not text-to-image generation. Pricing is per generated image rather than token-based. Newer Qwen image-edit models offer higher resolutions, multiple outputs, and more controls.

Cost

Model pricing

Input $0.045 per image in Singapore for international deployment
Model guide

Qwen-Image-Edit: Natural-Language Image Editing with Bilingual Text Support

Qwen-Image-Edit is Alibaba's image-to-image model for editing images through text instructions. It supports semantic changes such as pose, viewpoint, style, and context modifications, as well as appearance edits including object insertion, removal, enhancement, image fusion, and Chinese or English text editing. The original model accepts text and image inputs and returns one PNG image at a documented maximum resolution of 1024 by 1024 pixels.

What is Qwen-Image-Edit?

Qwen-Image-Edit is an image-to-image model developed by Alibaba's Qwen team and made available through Alibaba Cloud Model Studio. It is designed to modify an existing image according to a written instruction rather than generate a completely new image from text alone.

For example, a user can provide a photograph and request that an object be removed, a subject's clothing be changed, the background be replaced, or the visual style be transformed. The model is also intended for more structured visual changes, including pose manipulation, viewpoint changes, image fusion, and edits to text embedded in an image.

Alibaba announced the original model on August 19, 2025. It extends the Qwen-Image family's emphasis on text rendering into image-editing workflows. The current catalog also includes newer variants, but this page concerns the original qwen-image-edit model rather than those later releases.

Primary purpose and core capabilities

The model's main purpose is natural-language image editing. Its capabilities can be grouped into two broad categories:

  • Semantic editing: changing what a subject is doing or how the scene is presented, such as altering pose, viewpoint, artistic style, or surrounding context.
  • Appearance editing: making localized changes, such as adding, removing, or modifying objects while attempting to preserve the rest of the image.

These capabilities make Qwen-Image-Edit suitable for practical retouching and creative transformations. A user might ask it to replace an object, change clothing, enhance details, transfer an artistic treatment, or combine visual elements from multiple images.

One of its more distinctive documented uses is editing text inside images. The original Qwen release highlighted Chinese and English text editing, including changes intended to preserve the apparent font, size, and style. This can be useful for adapting signs, labels, posters, mockups, and other graphics without recreating the entire image manually.

Inputs, outputs, and resolution limits

Qwen-Image-Edit accepts text instructions together with image input. Alibaba Cloud Model Studio documentation lists support for multi-image input and image fusion, although the exact input behavior and supported parameters should be checked in the current API documentation before building a production workflow.

Supported image formats include JPG, JPEG, PNG, BMP, TIFF, WEBP, and GIF. For an animated GIF, the first frame is processed rather than the animation as a sequence. The model produces image output in PNG format.

The original model supports one output image per request. Its documented maximum output resolution is 1024 by 1024 pixels, and the output resolution is fixed rather than freely customizable in the way available with some newer image-editing models. This limit is important for users working on print assets, large-format images, or workflows that require several variations from one request.

SpecificationDocumented behavior
ProviderAlibaba Cloud Model Studio
Model nameqwen-image-edit
InputText instructions and image input
Image formatsJPG, JPEG, PNG, BMP, TIFF, WEBP, and GIF
OutputOne PNG image per request
Maximum documented resolution1024 by 1024 pixels
Pricing reference$0.045 per generated image in Singapore for international deployment

Reasoning, coding, and tool support

Qwen-Image-Edit is an image-editing model, not a general-purpose conversational model. It accepts text as an editing instruction, but it does not provide normal text-generation output. Consequently, conventional language-model measures such as a conversational context window or maximum text-output token limit are not applicable in the supplied model documentation.

It should not be selected for coding, long-form writing, question answering, or general reasoning tasks. The model record identifies coding capability as unsupported for its intended role, and the model is not a text-generation endpoint.

Alibaba Cloud's model information lists function calling, structured outputs, web search, context caching, batch inference, and fine-tuning as unsupported for this model. These restrictions mean that the endpoint should be treated as a focused image transformation service rather than an agent that can call tools, return structured JSON, or coordinate a multi-step workflow by itself.

Pricing and practical cost considerations

Current international pricing documentation lists Qwen-Image-Edit at $0.045 per generated image in the Singapore region. The charge is image-based rather than token-based. Since the original model returns one image per request, each successful generation corresponds to one billed output under that pricing description.

The price can be attractive for applications that need straightforward single-image edits, especially when the fixed 1024-by-1024 output is sufficient. However, total project cost also depends on how many attempts are needed. Image editing often involves iterative prompting, and a low per-image price does not guarantee a low cost if a workflow requires many variations or repeated corrections.

The Singapore price is a regional reference, not a universal promise that every Alibaba Cloud region or service surface will use the same rate. Users should verify the current Model Studio pricing page, deployment region, account requirements, and applicable billing terms before committing to production usage.

Main strengths and trade-offs

Qwen-Image-Edit's strongest feature is the breadth of edits it attempts to express through ordinary language. It combines localized object changes with broader semantic transformations, so users can describe both what should change and how the overall image should look afterward.

  • Natural-language workflow: users can describe an edit instead of manually manipulating layers or masks.
  • Useful semantic changes: pose, viewpoint, style, and scene context are among the documented editing categories.
  • Localized appearance edits: object insertion, removal, modification, and detail enhancement are supported use cases.
  • Bilingual image-text editing: Chinese and English text changes are a notable part of the original model's positioning.
  • Image fusion: multi-image workflows can combine visual elements when the relevant Model Studio endpoint and parameters are available.
  • Focused image output: the model is specialized for returning an edited image rather than mixing image editing with unrelated chat behavior.

Its main trade-offs are the single-output design, the 1024-by-1024 maximum documented resolution, and the lack of tool or structured-output features. It also does not provide the higher-resolution, multiple-output, or additional controls associated with newer Qwen image-edit variants.

When to choose Qwen-Image-Edit

Choose Qwen-Image-Edit when the task is a natural-language edit of an existing image and a single PNG output at up to 1024 by 1024 pixels is acceptable. It is a reasonable fit for:

  • Removing or inserting objects in photographs
  • Replacing clothing, backgrounds, or other visible elements
  • Applying a style transformation to an existing image
  • Changing a subject's pose or visual context
  • Editing Chinese or English lettering inside graphics
  • Combining elements from multiple image inputs
  • Building an application around Alibaba Cloud Model Studio's image-generation API

It is especially relevant when the workflow values direct visual editing over manual layer control and does not need a series of output candidates from every request.

When another option may be more appropriate

A newer Qwen image-edit model may be a better choice when the project needs higher resolution, multiple outputs, or more advanced controls. The supplied catalog identifies qwen-image-edit-plus, qwen-image-edit-max, Qwen Image 2.0, and Qwen Image 3.0 as newer or alternative options, but their individual capabilities and pricing should be evaluated separately rather than assumed from the original model.

For a text assistant, coding system, research agent, or tool-using application, Qwen-Image-Edit is the wrong model type because it is not designed to generate text, call functions, browse the web, or return structured JSON. A general-purpose multimodal model would be more suitable for those tasks.

It is also a poor fit when the final asset must exceed the documented 1024-by-1024 limit, when the application needs several images per request, or when the editing pipeline depends on customizable output dimensions. In those cases, a newer image-editing endpoint or a conventional graphics workflow may offer better control.

Bottom line

Qwen-Image-Edit is a focused Alibaba Cloud image-editing model for changing existing images with text instructions. Its useful distinction is the combination of semantic editing, localized object manipulation, image fusion, and Chinese and English text editing in one image-to-image workflow.

The original model is best understood as a costed, single-output editing endpoint: it accepts text and images, returns one PNG image, and has a documented maximum resolution of 1024 by 1024 pixels. Those boundaries make it practical for many retouching and creative tasks, but they also clearly separate it from newer Qwen image-edit variants and from general-purpose language or multimodal models.


Answers to Frequently Asked Questions

When should I choose a different model instead of Qwen-Image-Edit?
A different model may be more appropriate when you need resolutions above 1024 by 1024 pixels, multiple outputs per request, customizable dimensions, advanced controls, tool calling, structured JSON, coding, general reasoning, or text-generation capabilities. Newer Qwen image-edit variants may provide some of these features.
How much does Qwen-Image-Edit cost?
Alibaba Cloud's international pricing documentation lists Qwen-Image-Edit at $0.045 per generated image in the Singapore region. Pricing may vary by region and service terms, so users should verify the current Model Studio pricing and billing requirements before production use.
What image formats and resolution does Qwen-Image-Edit support?
Qwen-Image-Edit accepts JPG, JPEG, PNG, BMP, TIFF, WEBP, and GIF images. For animated GIFs, only the first frame is processed. The model returns one PNG image per request and has a documented maximum output resolution of 1024 by 1024 pixels.
What is Qwen-Image-Edit used for?
Qwen-Image-Edit is an image-to-image model for modifying existing images with natural-language instructions. It can remove or add objects, change clothing or backgrounds, alter poses and viewpoints, transform visual styles, combine elements from multiple images, and edit text embedded in graphics.
Can Qwen-Image-Edit edit Chinese and English text inside images?
Yes. The original Qwen-Image-Edit model is designed to edit Chinese and English text within images while attempting to preserve characteristics such as the apparent font, size, and style. This can be useful for signs, labels, posters, and mockups.


Sources 5
Provider

About Qwen