Qwen-Image-2.1

Qwen-Image-2.1-PE-T2I

by Qwen · Current open-weight release

An open-weight Qwen prompt rewriting model for text-to-image workflows. It expands short multilingual requests into detailed English prompts, recommends an aspect ratio, and prepares input for Qwen-Image-2.1 rather than generating images itself.

Text Reasoning Coding
Qwen-Image-2.1-PE-T2I is a specialized model for preparing prompts rather than generating images. Described as a fine-tuned Qwen3.5-VL 9B model, it expands a brief image idea into a more detailed English description and returns a recommended width-to-height ratio. The resulting prompt can then be supplied to Qwen-Image-2.1 for rendering.
Outputs

What Qwen-Image-2.1-PE-T2I can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Structured output
Model profile

Performance characteristics

5/10 Reasoning
2/10 Coding
5/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Qwen-Image-2.1
Model type Other
Context window 262K tokens
Maximum output 16K tokens
Release date 2026-09-20
Status Current open-weight release
Knowledge cutoff notes

No authoritative knowledge-cutoff date is specified in the model card or repository documentation.

Model notes

This is a prompt rewriting model, not the Qwen-Image-2.1 image-generation model. It is described as a fine-tuned Qwen3.5-VL 9B model and returns a JSON object containing rewritten_prompt and wh_ratio after a reasoning block. The published example uses max_new_tokens=16256. The weights are approximately 18.8 GB and are released under the Qwen Research License Agreement. No official hosted pricing, batch API, fine-tuning support, streaming endpoint, or separate legacy JSON-mode capability was verified for this exact checkpoint.

Model guide

Qwen-Image-2.1-PE-T2I: Open-Weight Prompt Rewriting for Image Generation

Qwen-Image-2.1-PE-T2I is an open-weight prompt rewriting model from Qwen. It transforms short, including multilingual, image requests into detailed English prompts and recommends an aspect ratio for use with the Qwen-Image-2.1 image-generation model.

What is Qwen-Image-2.1-PE-T2I?

Qwen-Image-2.1-PE-T2I is an open-weight text-to-image prompt rewriting model provided by Qwen. Its name describes its role: it is a prompt enhancer for text-to-image workflows, not the image-generation model itself. The model takes a short request such as an idea for a scene, subject, or design and expands it into a detailed English prompt intended for image synthesis.

The model is positioned alongside the Qwen-Image-2.1 workflow. Qwen-Image-2.1-PE-T2I prepares the instruction, while Qwen-Image-2.1 performs the actual image generation or editing. This separation can be useful when a brief user request does not contain enough detail for consistent rendering.

According to the supplied model documentation, the checkpoint is a fine-tuned Qwen3.5-VL 9B model. It is available as downloadable weights through the Qwen organization on Hugging Face under the Qwen Research License Agreement.

How the prompt rewriting workflow works

A user supplies a concise image request in a supported language. Qwen-Image-2.1-PE-T2I produces a response containing a reasoning section followed by a JSON-style result. The result includes two principal fields:

  • rewritten_prompt: a more detailed English prompt describing the intended image.
  • wh_ratio: a recommended width-to-height aspect ratio, such as 1:1, 4:3, 16:9, or 9:16.

The rewritten prompt can be passed to the Qwen-Image-2.1 Diffusers pipeline, while the recommended ratio can guide the choice of output dimensions. For example, a landscape scene may receive a 16:9 recommendation, whereas a portrait-oriented composition may receive 9:16. The ratio is a recommendation for the downstream image-generation step, not an image output from this model.

Applications should parse and validate the final object rather than assuming that every token in the response is JSON. The published workflow includes a reasoning block before the structured result, so a production integration should identify the final object and check that both expected fields are present.

Main capabilities

  • Prompt expansion: Converts short image requests into richer English descriptions suitable for text-to-image generation.
  • Multilingual request handling: The model is intended to accept image requests written in multiple languages while producing an English rewritten prompt.
  • Composition guidance: Recommends an aspect ratio for the downstream image-generation task.
  • Structured result: Returns machine-readable fields for the rewritten prompt and recommended ratio after a reasoning section.
  • Local deployment: Its open weights can be used with Transformers, PyTorch, vLLM, or compatible tooling according to the supplied documentation.

These capabilities make the model most useful as an intermediate step between a user-facing image prompt and a rendering model. It can standardize the input to an image pipeline without requiring every user to write a long, technically detailed prompt.

Technical specifications and limits

SpecificationVerified detail
ProviderQwen
Model typeOpen-weight text-to-image prompt rewriting model
Underlying descriptionFine-tuned Qwen3.5-VL 9B model
Text context limit262,144 positions, according to the supplied configuration
Published maximum generation example16,256 newly generated tokens
Primary outputDetailed English prompt and recommended aspect ratio
Direct image outputNo
Hosted pricingNo official per-token or per-request price was verified
WeightsApproximately 18.8 GB of safetensor files, according to the supplied repository information

The 262,144-position figure is a model configuration limit for the text context, not a guarantee that every practical application will use the entire window efficiently. The published example uses a maximum of 16,256 new tokens, which is an implementation setting documented for generation rather than evidence that typical prompt-rewriting requests need that many tokens.

The model is distributed as weights rather than as a documented hosted API product with a public usage tariff. As a result, there is no verified input price, output price, monthly plan, or request-based pricing to compare. Local operating costs will depend on the hardware, deployment stack, and usage volume.

Input, output, and reasoning behavior

The supplied capability data classifies the model as text-input and text-output. Although its architecture is described with the Qwen3.5-VL designation, the documented use here is rewriting textual image requests. The checkpoint should not be treated as a direct image, audio, or video generator.

The response includes a reasoning block before the final structured object. This can help the model plan a richer description, but it also creates an integration consideration: applications should not send the complete unparsed response to a downstream component that expects strict JSON. The model documentation describes the result as containing rewritten_prompt and wh_ratio; external validation is still appropriate before those values are used.

There is no verified native tool or function-calling interface for this exact checkpoint. Likewise, streaming, fine-tuning support, caching, and a separate legacy JSON-mode capability were not verified in the supplied research. Its structured output should therefore be understood as part of the model's demonstrated response format, not automatically as a hosted API JSON-mode feature.

Strengths and trade-offs

The clearest strength of Qwen-Image-2.1-PE-T2I is specialization. Instead of asking a general-purpose language model to invent a detailed image prompt, a developer can use a model designed around the Qwen-Image-2.1 generation workflow. Its ratio recommendation adds a practical piece of composition metadata that a plain prompt rewrite might omit.

Open weights are another important distinction. They allow organizations to examine and deploy the checkpoint locally using the documented ecosystem, rather than relying exclusively on a hosted image or language API. This can be useful for private prompt-processing workflows or applications that already operate their own inference infrastructure.

The trade-off is that local deployment shifts responsibility to the operator. Hardware provisioning, model loading, response parsing, performance tuning, and license review are not handled by a managed endpoint. The approximately 18.8 GB of weights also represents a more substantial deployment footprint than a lightweight hosted prompt-rewriting service.

The model's editorial evaluation in the supplied data rates its reasoning at 5 out of 10, coding at 2 out of 10, speed at 5 out of 10, and cost at 9 out of 10. These are editorial scores, not provider-published benchmarks. The high cost score reflects the absence of a verified hosted usage price and the potential advantage of downloadable weights, but actual operating cost depends on deployment hardware and utilization.

Best use cases

  • User-friendly image interfaces: Let users enter short ideas while the application prepares a more descriptive prompt for rendering.
  • Multilingual image-request handling: Normalize requests written in different languages into detailed English prompts for the downstream Qwen image workflow.
  • Batch prompt preparation: Expand a collection of brief concepts before sending them to an image-generation pipeline.
  • Aspect-ratio selection: Use the returned wh_ratio value to help choose a landscape, portrait, or square canvas.
  • Local or controlled deployments: Run the prompt-rewriting stage with downloadable weights when a hosted endpoint is not appropriate.

A typical application could accept a request such as “a cyclist riding through a rainy neon city at night,” send it to Qwen-Image-2.1-PE-T2I, extract the detailed English rewritten_prompt and wh_ratio, map the ratio to supported dimensions, and then pass both the prompt and dimensions to Qwen-Image-2.1.

When to choose Qwen-Image-2.1-PE-T2I

Choose this model when the main problem is prompt quality or consistency before image generation. It is especially suitable when the target renderer is Qwen-Image-2.1, when multilingual short requests are common, or when local access to open weights is more important than a managed API.

A different option may be more appropriate when the application needs the final image directly, because this checkpoint does not synthesize images. Qwen-Image-2.1 is the relevant component for rendering and editing. A hosted prompt service may also be preferable when a team wants to avoid downloading approximately 18.8 GB of weights and managing inference infrastructure, although no competing service or price comparison is verified in the supplied research.

It is also not a strong fit for general chat, coding assistance, tool execution, or broad multimodal application control. Its purpose is narrower: improve a text-to-image request and provide a recommended aspect ratio.

Limitations and deployment notes

The most important limitation is the division between prompt preparation and image generation. Installing this checkpoint alone will not produce an image. A separate compatible renderer, such as the Qwen-Image-2.1 pipeline, is required for the final visual result.

Developers should also account for the reasoning text that precedes the final object. Robust extraction and schema validation are necessary if the output feeds an automated pipeline. The documented format identifies the expected fields, but applications should still handle missing, malformed, or unexpected values.

No official hosted pricing, public batch API, streaming endpoint, or verified fine-tuning service was identified for this exact model. Its availability and intended use are therefore best understood as an open-weight deployment option rather than a conventional pay-per-request model API.

Overall, Qwen-Image-2.1-PE-T2I is a focused utility model for making image-generation requests more detailed and operationally consistent. Its value comes from the prompt-rewriting stage and aspect-ratio recommendation; users seeking direct image output, general-purpose reasoning, or a turnkey hosted endpoint should evaluate a different type of option.


Answers to Frequently Asked Questions

What is Qwen-Image-2.1-PE-T2I used for?
Qwen-Image-2.1-PE-T2I is an open-weight prompt rewriting model for text-to-image workflows. It expands short image requests into detailed English prompts and recommends an aspect ratio for downstream image generation.
Does Qwen-Image-2.1-PE-T2I generate images directly?
No. Qwen-Image-2.1-PE-T2I prepares the prompt and returns a recommended width-to-height ratio, but a separate image-generation model such as Qwen-Image-2.1 is required to create or edit the final image.
What does the output of Qwen-Image-2.1-PE-T2I contain?
The documented output contains a reasoning section followed by a structured result with two main fields: rewritten_prompt, which provides a detailed English image prompt, and wh_ratio, which recommends an aspect ratio such as 1:1, 4:3, 16:9, or 9:16.
Can Qwen-Image-2.1-PE-T2I be deployed locally?
Yes. The model provides downloadable open weights and can be used with tools such as Transformers, PyTorch, vLLM, or compatible deployment stacks. The supplied repository information lists approximately 18.8 GB of safetensor files, so local deployment requires suitable hardware and operational management.
Who should choose Qwen-Image-2.1-PE-T2I?
It is suitable for applications that need to improve short or multilingual image requests before sending them to Qwen-Image-2.1 or another renderer. It is especially useful for user-facing image interfaces, batch prompt preparation, aspect-ratio selection, and controlled local deployments.


Sources 5
Provider

About Qwen