MAI-Image-2.5

MAI-Image-2.5-Flash

by Microsoft Copilot · Public preview; scheduled for retirement on 2026-10-01

Microsoft’s MAI-Image-2.5-Flash is a public-preview image model for text-to-image generation and image-to-image editing. It accepts text and up to five JPEG or PNG images, returns one PNG image, and is priced for faster, lower-cost visual workflows. Its preview status and scheduled October 1, 2026 retirement date are important limitations.

Image generation Reasoning Coding
MAI-Image-2.5-Flash is a Microsoft AI image-generation model optimized for faster, lower-cost visual creation and editing. Available in public preview through Microsoft Foundry, it can generate images from text and edit existing images using text instructions plus up to five JPEG or PNG inputs. Its main trade-off is clear: it is positioned for speed and economical image workflows, while remaining a preview service with an announced retirement date of October 1, 2026.
Outputs

What MAI-Image-2.5-Flash can produce

Image generation
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Multimodal output
Model profile

Performance characteristics

5/10 Reasoning
1/10 Coding
9/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family MAI-Image-2.5
Model type Multimodal
Context window 32K tokens
Maximum output 4K tokens
Release date 2026-06-02
Status Public preview; scheduled for retirement on 2026-10-01
Shutdown date 2026-10-01
Knowledge cutoff notes

Microsoft has not published a direct knowledge-cutoff date for this image-generation model. Its behavior is based on trained visual and textual data, but no authoritative cutoff date was verified.

Model notes

This is Microsoft's faster, lower-cost Flash variant in the MAI-Image-2.5 family. Microsoft Foundry documents text-to-image generation and image-to-image editing, with text and up to five JPEG or PNG images accepted for editing workflows and one PNG image returned. The model is currently accessible in public preview but has an official retirement date of October 1, 2026. Microsoft Foundry's quick-facts card lists a 131,072-token context window, while the current Microsoft Learn capability table lists 32,000 tokens; the model-specific capability table is used here. Image dimensions must be at least 768 pixels per side and the total pixel count cannot exceed 1,048,576.

Cost

Model pricing

Input $1.75 per 1M text input tokens; $1.75 per 1M image input tokens
Output $19.50 per 1M image output tokens
Model guide

MAI-Image-2.5-Flash: Microsoft’s Fast, Lower-Cost Image Generation Model

MAI-Image-2.5-Flash is Microsoft’s preview image model for fast, lower-cost text-to-image generation and image-to-image editing through Microsoft Foundry. It accepts text and image inputs, returns PNG images, and is aimed at high-volume creative and production workflows where speed and cost matter more than maximum image fidelity or long-term availability.

What is MAI-Image-2.5-Flash?

MAI-Image-2.5-Flash is Microsoft’s image-generation and image-editing model in the MAI-Image-2.5 family. Microsoft makes it available through Microsoft Foundry, where it is listed as a model sold directly by Azure. The model is currently in public preview.

The “Flash” designation describes its position within the family: it is intended to provide faster generation and lower pricing than the standard MAI-Image-2.5 model. That makes it a practical candidate for applications that create many images, need quick visual iteration, or cannot justify the cost of a higher-fidelity image model for every request.

This is not a general-purpose conversational model. Its primary output is an image, not a text response, code completion, audio file, video, embedding, or structured business record. Text is used to describe the desired image or edit, but the model returns one PNG image.

Core capabilities and supported inputs

MAI-Image-2.5-Flash supports two main workflows:

  • Text-to-image generation: a text prompt describes a new image to create.
  • Image-to-image editing: one or more source images are supplied with instructions describing the requested changes.

Microsoft Foundry documentation describes support for text and up to five JPEG or PNG images in editing workflows. Example tasks include replacing a particular object, changing a layout, updating text in an image, cleaning up motion blur, visualizing a concept, or adapting an existing composition for another use.

Input images must be at least 768 pixels on each side, and the total image size cannot exceed 1,048,576 pixels. This constraint matters when an application accepts user uploads: oversized images may need to be resized before they are submitted, while very small images may not meet the model’s documented requirements.

The output is one PNG image. The model does not provide a documented choice of multiple returned images in the supplied specifications, and its image output should not be confused with text output or a general multimodal chat response.

Technical specifications at a glance

SpecificationVerified detail
ProviderMicrosoft AI
PlatformMicrosoft Foundry
Model familyMAI-Image-2.5
Version2026-06-02
StatusPublic preview
InputText and up to five JPEG or PNG images for editing workflows
OutputOne PNG image
Context length32,000 tokens according to the model-specific capability table
Maximum output limit4,096 tokens
Tool or function callingNot supported
Web searchNot supported

Microsoft Foundry also displays a different 131,072-token context figure in a quick-facts card. The model-specific capability table, which documents this model’s capabilities in more detail, lists 32,000 tokens; that is the more conservative figure to use when planning workloads.

Pricing and cost positioning

Microsoft’s announced pricing is usage-based rather than a recurring consumer subscription. The listed rates are:

  • Text input: $1.75 per 1 million tokens
  • Image input: $1.75 per 1 million tokens
  • Image output: $19.50 per 1 million tokens

Image-output pricing is therefore the most important cost component for applications that generate substantial visual content. A system that repeatedly creates images should estimate output usage separately from prompt and image-input usage, because the rates apply to different parts of the request.

The Flash variant’s pricing and speed positioning make it more suitable for high-volume generation, rapid experimentation, automated creative variations, and other workflows where many requests are more important than achieving the highest possible quality on every individual image. The research does not provide a benchmark comparing its latency or visual quality with other models, so claims about exact speed improvements or ranking should be treated as unverified.

Main strengths and trade-offs

The clearest strength of MAI-Image-2.5-Flash is its focus. It combines image generation with image editing while emphasizing faster, lower-cost operation than the standard MAI-Image-2.5 model. Support for up to five source images is useful for editing workflows that need references, compositional elements, or several views of an existing design.

Its supported editing scenarios are also practical rather than limited to producing entirely new artwork. Targeted object replacement, text changes, layout adaptation, motion-blur cleanup, and production visualization can fit into design, marketing, prototyping, and content-production pipelines.

There are important trade-offs:

  • Preview status: the model is not presented as a generally available, fully stable service. Microsoft advises evaluating preview services carefully before relying on them for production workloads.
  • Limited output type: it returns images, not text, audio, video, code, or embeddings.
  • No documented tool calling: the model cannot independently invoke external tools or services as part of its model capability set.
  • Retirement risk: Microsoft’s retirement schedule lists October 1, 2026 as the model’s retirement date.
  • Image-quality trade-off: its Flash positioning favors speed and cost, so users seeking the maximum available fidelity may prefer a different image model or the standard MAI-Image-2.5 option, where available and appropriate.
  • Review requirements: generated images can contain visual inaccuracies, bias, or inappropriate content and should be checked before publication or use in sensitive contexts.

Reasoning, coding, and tool support

MAI-Image-2.5-Flash should not be evaluated like a reasoning or coding language model. Its documented role is visual generation and editing. The supplied model record assigns it a low coding score and does not identify general-purpose reasoning, coding assistance, web search, or tool use as supported capabilities. Those scores are editorial evaluations, not Microsoft-published benchmark results.

Although the model uses text prompts, that does not mean it provides ordinary text answers. The model record identifies text output as unsupported and image output as supported. Applications needing a written explanation, code generation, document analysis, web research, or function execution should use a model designed for those tasks, potentially alongside an image model rather than replacing it with MAI-Image-2.5-Flash.

Best use cases

MAI-Image-2.5-Flash is most appropriate when a workflow needs visual output quickly and at controlled usage cost. Suitable examples include:

  • High-volume image generation: creating many concept images, campaign variations, or product ideas.
  • Fast concept visualization: turning early descriptions into visual references for design and creative review.
  • Production design: exploring scenes, layouts, compositions, and other visual directions before final production.
  • Targeted editing: changing an object, adjusting text, adapting a layout, or cleaning up an existing image.
  • Creative-content pipelines: generating visual assets where fast iteration is more valuable than maximum fidelity for every request.
  • Image-editing applications: accepting a prompt and several reference images to produce one revised PNG.

For an application, a sensible workflow is to validate image dimensions, restrict accepted file types to JPEG and PNG, communicate that the service is in preview, and add human or automated review before publishing generated material. Because the model has a scheduled retirement date, production systems should also have a migration plan.

When should you choose MAI-Image-2.5-Flash?

Choose MAI-Image-2.5-Flash when the priority is economical image generation or editing at relatively high volume, and when the application can tolerate preview status and a limited service lifetime. It is especially attractive for rapid design exploration, automated visual variations, and edits that do not require a general-purpose AI assistant.

Another image model may be more appropriate when maximum fidelity, long-term availability, or a more established production lifecycle is more important than cost and speed. The standard MAI-Image-2.5 model is the most directly relevant sibling comparison, because Flash is positioned as its faster, lower-cost counterpart. The supplied research does not establish exact quality or latency differences, so the choice should be tested with representative prompts and source images rather than based on an assumed benchmark gap.

Use a general-purpose language model instead when the central task is conversation, coding, research, structured text generation, or tool orchestration. MAI-Image-2.5-Flash can be part of a larger workflow in those cases, but it should remain the visual-generation component rather than the system’s primary reasoning or automation engine.

Availability and lifecycle considerations

MAI-Image-2.5-Flash is available through Microsoft Foundry as a Microsoft-hosted model in public preview. Its listed version is 2026-06-02. The official retirement schedule gives October 1, 2026 as the retirement date, so availability should not be treated as indefinite.

That lifecycle is a major practical consideration for teams building an application around the model. Before adoption, teams should confirm regional and account availability, test the model against the intended image-editing cases, estimate image-output costs, and identify a replacement or fallback. Preview services can change, and a model with a fixed retirement date is particularly unsuitable as the only dependency for a long-lived production pipeline.

Bottom line

MAI-Image-2.5-Flash is a focused Microsoft image model for fast, cost-conscious generation and editing. Its support for text prompts, multiple JPEG or PNG inputs, and PNG image output makes it useful for visual production and iterative editing workflows. Its limitations are equally important: it is preview-only, does not provide general text or coding capabilities, lacks documented tool calling, and is scheduled for retirement on October 1, 2026.

It is best viewed as a practical short- to medium-term option for high-volume image work where speed and cost matter. Teams that need durable availability, maximum image quality, or broader AI capabilities should compare alternatives before making it a central production dependency.


Answers to Frequently Asked Questions

Is MAI-Image-2.5-Flash suitable for production use?
MAI-Image-2.5-Flash can support production workflows, but it is currently in public preview and has a scheduled retirement date of October 1, 2026. Teams should validate regional availability, review generated images, estimate output costs, and maintain a migration or fallback plan before making it a critical dependency.
What are the best use cases for MAI-Image-2.5-Flash?
The model is well suited to high-volume image generation, rapid concept visualization, automated creative variations, production design, targeted image edits, layout adaptation, object replacement, and motion-blur cleanup. It is most useful when speed and cost matter more than maximum fidelity for every image.
How much does MAI-Image-2.5-Flash cost?
Microsoft lists usage-based pricing of $1.75 per 1 million text-input tokens, $1.75 per 1 million image-input tokens, and $19.50 per 1 million image-output tokens. Image output is the primary cost factor for applications that generate many images.
What is MAI-Image-2.5-Flash?
MAI-Image-2.5-Flash is Microsoft’s image-generation and image-editing model in the MAI-Image-2.5 family. It is available through Microsoft Foundry in public preview and is designed to provide faster, lower-cost visual generation than the standard MAI-Image-2.5 model.
What inputs and outputs does MAI-Image-2.5-Flash support?
The model accepts text prompts for image generation and text instructions with up to five JPEG or PNG images for editing workflows. It returns one PNG image. Input images must be at least 768 pixels on each side and cannot exceed 1,048,576 total pixels.


Sources 6
Provider

About Microsoft Copilot