Qwen-Image-2.1

Qwen-Image-2.1-PE-I2I

by Qwen · Current; open-weight research release

An open-weight Qwen3.5-VL 9B model that analyzes reference images and vague editing instructions, then returns precise prompts and aspect-ratio guidance for downstream Qwen-Image-2.1 workflows.

Text Reasoning Coding
Qwen-Image-2.1-PE-I2I is a specialized multimodal prompt-rewriting model from Qwen. It accepts text instructions and reference images, analyzes the requested changes, and returns a structured JSON record containing a detailed editing prompt plus aspect-ratio guidance. It does not generate or edit pixels itself; its role is to prepare better instructions for a downstream image-generation or image-editing model.
Outputs

What Qwen-Image-2.1-PE-I2I can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Streaming Structured output Batch API
Model profile

Performance characteristics

5/10 Reasoning
2/10 Coding
5/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Qwen-Image-2.1
Model type Multimodal
Context window 262K tokens
Maximum output 24K tokens
Release date 2026-09-20
Status Current; open-weight research release
Knowledge cutoff notes

No explicit model-specific knowledge-cutoff date was found in the authoritative model card or official Qwen repository. The checkpoint is primarily an image-editing prompt-rewriting model rather than a general knowledge model.

Model notes

This is a dedicated image-to-image prompt enhancer, not the Qwen-Image-2.1 pixel-generation model. It is described by Qwen as a fine-tuned Qwen3.5-VL 9B checkpoint. The official edit profile accepts text plus one or more images and returns a JSON contract containing rewritten_prompt, wh_ratio, and ratio_follow. The two ratio fields are mutually exclusive. The documented production profile uses max_new_tokens=24000, temperature=1.0, top_p=0.95, top_k=20, and required thinking. The repository supports local Transformers inference, vLLM offline batch processing, and vLLM server deployment. No official hosted API price was found for this checkpoint. The model uses the Qwen Research License.

Model guide

Qwen-Image-2.1-PE-I2I: Prompt Enhancement for Reference-Based Image Editing

Qwen-Image-2.1-PE-I2I is an open-weight Qwen3.5-VL 9B model fine-tuned to transform vague image-editing instructions and one or more reference images into precise, structured prompts for the Qwen-Image-2.1 editing pipeline.

What is Qwen-Image-2.1-PE-I2I?

Qwen-Image-2.1-PE-I2I is an open-weight image-to-image prompt enhancement checkpoint provided by the Qwen team. It is described as a fine-tuned Qwen3.5-VL 9B model, adapted for a specific task: turning short, ambiguous image-editing requests into detailed instructions that another model can execute.

The distinction between this checkpoint and an image generator is important. Qwen-Image-2.1-PE-I2I does not return an edited image. Instead, it examines the input instruction and source image or images, determines what should change and what should remain unchanged, and produces a more explicit prompt for a downstream editing pipeline such as Qwen-Image-2.1.

For example, a short request such as “make the jacket red and put the person in a rainy street” may need to preserve the subject’s identity, pose, camera framing, and other details. The prompt-enhancement model expands that request into a more operational description for the image-editing model.

Where it fits in the Qwen lineup

This model occupies a supporting role within the Qwen image-editing workflow. Qwen-Image-2.1 is the downstream image-generation and editing system, while Qwen-Image-2.1-PE-I2I prepares the instruction that helps that system interpret a difficult edit. It is therefore best understood as a task-specific companion rather than a smaller general-purpose version of the image model.

Its Qwen3.5-VL 9B lineage gives it both language and visual input capabilities, but the checkpoint has been fine-tuned around prompt rewriting for image editing. That specialization makes it more relevant to reference-image workflows than to ordinary chat, general visual question answering, or unrestricted content generation.

How the model processes an edit

The model receives a textual editing request together with one or more reference images. It interprets the visual content, identifies the requested transformation, and writes a detailed instruction describing the intended result. It can also specify details that should be preserved, which is useful when only part of an image should change.

Multiple reference images are supported in the documented edit profile. When several images are supplied, the generated instruction can distinguish them with placeholders such as <image1> and <image2>. This allows an editing workflow to describe relationships between sources, such as using the subject from one image and the background or object from another.

The model also provides framing guidance. It can recommend a new width-to-height ratio or indicate that the output should follow the composition of a particular input image. This is useful when an edit needs to move from a portrait source to a landscape canvas, or when preserving the original framing is more important than changing the output dimensions.

What the structured output contains

The official workflow returns a task-specific JSON record with three fields:

FieldPurpose
rewritten_promptThe expanded image-editing instruction for the downstream model.
wh_ratioA requested width-to-height ratio for a new output composition.
ratio_followAn instruction to follow the framing or ratio of a specified input image.

The two ratio fields are mutually exclusive. A workflow should use either wh_ratio or ratio_follow, depending on whether it wants a new composition ratio or wants to preserve the framing of a source image.

This structured record should not be confused with a general provider-hosted JSON mode. The model produces JSON because its prompt-rewriting task defines a particular output contract. The supplied research does not verify a separate general-purpose JSON-mode API for this checkpoint.

Supported inputs and outputs

Qwen-Image-2.1-PE-I2I supports text input and image input. The documented profile is designed for one or more reference images, making it suitable for image-to-image editing preparation and multi-image composition instructions. The research does not identify audio or video input support.

Its direct output is text, including the rewritten prompt and structured aspect-ratio fields. It does not directly produce images, video, audio, or other non-text media. The actual image output must come from a separate downstream image-generation or image-editing model.

Context, output limits, and deployment

The recorded context length is 262,144 tokens, with a maximum output setting of 24,000 tokens. These are substantial limits for a prompt-rewriting task, although practical usage will usually be much shorter because the model is intended to describe image edits rather than generate long documents.

The official materials describe local Transformers inference, vLLM offline batch processing, and vLLM server deployment. This makes the checkpoint suitable for teams that want to run prompt enhancement as part of a self-managed image pipeline. Local deployment also means that users must provide the inference environment and enough GPU memory for the model; the supplied research does not provide a specific hardware minimum.

The documented production-style profile uses max_new_tokens=24000, temperature 1.0, top_p=0.95, top_k=20, and required thinking. These settings describe the official workflow rather than a guarantee that every deployment must use identical parameters.

Main strengths and trade-offs

  • Specialized editing interpretation: The checkpoint is focused on turning vague editing requests into detailed, actionable instructions.
  • Reference-image awareness: It can inspect one or more images and describe which visual elements should change or remain consistent.
  • Multi-image coordination: Image placeholders allow the rewritten prompt to refer to different source images explicitly.
  • Composition guidance: The output includes a mechanism for selecting a new aspect ratio or following a source image’s framing.
  • Self-managed deployment: Downloadable weights and documented Transformers and vLLM workflows provide more control than a checkpoint available only through a hosted endpoint.

Those strengths come with clear limitations. The model is not the image renderer, so a complete application requires a second model and an orchestration step. It also requires local inference infrastructure unless a third party makes it available through a hosted service. No official hosted API price was found for this checkpoint.

Because it is specialized, it is not an obvious choice for general conversation, coding, audio or video generation, or tasks that only need a simple text prompt. Adding it to a pipeline may improve complex edit instructions, but it also adds latency and operational complexity compared with sending a short prompt directly to an image model.

Capability, speed, and cost profile

The available editorial assessment rates its reasoning capability at 5 out of 10, coding at 2 out of 10, speed at 5 out of 10, and cost at 8 out of 10. These are editorial scores, not Qwen-published benchmark results. They reflect the model’s narrow purpose: it is more useful for visual instruction interpretation than for software development, and its open-weight availability can be attractive for organizations that prefer to avoid per-token hosted pricing.

There is no verified public input or output token price for Qwen-Image-2.1-PE-I2I. Its financial profile depends on the cost of operating the required local or third-party infrastructure. For high-volume batch workloads, self-hosting may provide predictable infrastructure economics, while occasional users may find a hosted image-editing service or a simpler prompt workflow more convenient.

The research records tool use and function calling as unsupported. The model can return structured task output, but that is different from calling external tools, browsing the web, or executing application functions.

When to choose Qwen-Image-2.1-PE-I2I

Choose this model when the main difficulty is expressing an image edit precisely rather than rendering the final image. It is particularly suitable for:

  • Applications that accept short, user-written editing requests and need to expand them automatically.
  • Workflows where identity, pose, composition, or non-targeted image details should be preserved.
  • Tasks that combine several reference images and need explicit source-image references.
  • Batch systems that prepare large numbers of editing prompts before sending them to Qwen-Image-2.1 or another compatible image pipeline.
  • Teams that need downloadable weights and want to control deployment rather than depend on a published per-token API.

Another option may be more appropriate when the application needs direct image generation, a managed hosted API, real-time low-latency interaction, general-purpose visual conversation, or built-in tool calling. A downstream image model is also required if the desired result is an edited image rather than a rewritten instruction.

Availability and licensing

Qwen-Image-2.1-PE-I2I is distributed as downloadable weights through the Qwen organization on Hugging Face. The model uses the Qwen Research License. Commercial deployment, redistribution, and derivative use should therefore be checked against the license terms before the checkpoint is incorporated into a production system.

In practical terms, the model is best viewed as a specialized, open-weight planning component for image editing. It adds value before image synthesis by translating human editing intent and visual references into a more complete instruction, but it does not replace the image-generation stage itself.


Answers to Frequently Asked Questions

What is Qwen-Image-2.1-PE-I2I used for?
Qwen-Image-2.1-PE-I2I is used to convert short or ambiguous image-editing requests and reference images into detailed instructions for a downstream image-editing model such as Qwen-Image-2.1. It does not generate the edited image itself.
Does Qwen-Image-2.1-PE-I2I generate images directly?
No. Its direct output is text in a structured format, including a rewritten editing prompt and aspect-ratio guidance. A separate image-generation or image-editing model is required to produce the final image.
Can Qwen-Image-2.1-PE-I2I work with multiple reference images?
Yes. The documented workflow supports one or more reference images and can identify them with placeholders such as and . This enables prompts that combine elements from different source images, such as a subject from one image and a background from another.
What fields does the Qwen-Image-2.1-PE-I2I output contain?
The task-specific JSON output contains rewritten_prompt, wh_ratio, and ratio_follow. rewritten_prompt provides the expanded editing instruction, while wh_ratio specifies a new width-to-height ratio and ratio_follow instructs the downstream workflow to follow the framing of a specified input image. The two ratio fields are mutually exclusive.
How can Qwen-Image-2.1-PE-I2I be deployed and licensed?
The checkpoint is available as downloadable weights from the Qwen organization on Hugging Face and supports local Transformers inference, vLLM offline batch processing, and vLLM server deployment. It uses the Qwen Research License, so commercial use, redistribution, and derivative applications should be reviewed against the license terms.


Sources 7
Provider

About Qwen