What is MAI-Image-2.5-Pro?
MAI-Image-2.5-Pro is Microsoft's image generation and editing model for high-fidelity creative work. It can create an image from a written description or modify an existing image using text instructions and reference images. Unlike a general-purpose conversational model, its primary output is visual: the documented output format is a PNG image returned as base64-encoded data.
The model is hosted through Microsoft Foundry, Microsoft's catalog and deployment environment for AI models. Its current status is public preview, and the supplied Microsoft research identifies the exact Foundry model version as 2026-06-19. The model entered public preview on July 23, 2026, according to the supplied Microsoft AI announcement.
Microsoft positions MAI-Image-2.5-Pro as the highest-quality model in the MAI-Image-2.5 family. Its intended audience includes creative professionals, designers, marketing teams, commercial-image producers, and developers building applications that need controlled image creation or editing rather than simple image experimentation.
Core image-generation and editing capabilities
MAI-Image-2.5-Pro supports two main workflows:
- Text-to-image generation: a written prompt describes the desired subject, composition, style, environment, lighting, materials, or typography.
- Image-to-image editing: one or more supplied images provide visual context, while the prompt describes what should be changed, preserved, removed, or adapted.
Microsoft's documentation describes the model as supporting visual reasoning, precise object edits, layout adaptation, text updates, motion-blur cleanup, object consistency, character consistency, material and physical-property accuracy, and spatial and geometric reasoning. In practical terms, this makes it suitable for requests such as changing the text on a sign while preserving its perspective, adapting a product image to a new layout, repairing an image with motion blur, or keeping the same character and objects consistent across a composed scene.
Reference-image editing is particularly important for production workflows. Microsoft Foundry documentation states that MAI-Image-2.5 models can accept up to five JPEG or PNG reference images for image edits. Those references can provide the appearance of a product, a character, a setting, or several visual elements that need to remain coordinated in the result.
Where it fits in Microsoft's model lineup
MAI-Image-2.5-Pro belongs to Microsoft's MAI-Image-2.5 image-model family and is described as its highest-quality option. The distinction matters because this is not presented as a general Microsoft Copilot feature or as a text model that happens to create pictures. It is a dedicated image system exposed through Microsoft Foundry and image-generation or image-editing APIs.
The supplied documentation distinguishes MAI-Image-2.5-Pro from MAI-Image-2.6 models in at least one important way: web grounding is documented for MAI-Image-2.6, but not for MAI-Image-2.5-Pro. This model should therefore be selected for visual generation and editing, not for tasks that require native web search, current-information retrieval, or source-grounded answers.
Inputs, outputs, and technical limits
The model is multimodal in the practical sense that it accepts both text and images, but its output is image-only. It does not provide documented text, audio, video, or structured-data output.
| Specification | Documented value |
|---|---|
| Provider | Microsoft |
| Model family | MAI-Image-2.5 |
| Status | Public preview |
| Text input | Supported |
| Image input | Supported |
| Maximum reference images for edits | Up to five JPEG or PNG images |
| Context length | 131,072 tokens |
| Maximum output limit | 4,096 tokens |
| Image output | Supported; output is always PNG |
| Minimum image dimensions | Width and height must each be at least 768 pixels |
| Maximum total output pixels | 1,048,576 pixels, equivalent to 1024×1024 |
| Web search or grounding | Not documented for this model |
| Tool or function calling | Not supported in the supplied specification |
The context-length and output-token values are catalog specifications, but image dimensions impose a separate practical constraint. For MAI-Image-2.5 models, each output dimension must be at least 768 pixels and the total output cannot exceed 1,048,576 pixels. A 1024×1024 image reaches that stated total-pixel limit. The supplied research does not establish that every possible aspect ratio is available within those constraints, so applications should validate requested dimensions against the Foundry service.
What the model is especially good at
The strongest case for MAI-Image-2.5-Pro is a visual task where fidelity and control matter more than minimum cost or maximum throughput. Microsoft specifically positions it for high-quality hero imagery, detailed editing, commercial imagery, production design, and scenes that require consistent objects, materials, spatial relationships, and in-image text.
Accurate typography is a notable use case. Many image systems struggle when a prompt requires legible words, labels, packaging copy, signage, or interface-like layouts. MAI-Image-2.5-Pro is documented as supporting precise in-image text rendering and text updates during editing. That does not mean every generated word will be perfect, so text-heavy deliverables still require review, but it identifies typography as an intended strength rather than an incidental capability.
The model is also suited to controlled variations. A team could begin with a reference image and request a new composition, updated layout, changed materials, or a targeted object edit while preserving important visual relationships. Character and object consistency can be useful for campaign concepts, storyboards, product presentations, and other work where several elements must remain recognizable.
Pricing and cost considerations
Microsoft's supplied pricing information lists separate rates for text input, image input, and image output:
- Text input: $5 per 1 million tokens.
- Image input: $8 per 1 million tokens.
- Image output: $106 per 1 million tokens.
These are usage-based model prices rather than a consumer subscription price. The output rate is substantially higher than the input rates, so the main cost driver for image-heavy workloads is likely to be generated output rather than short prompts. Actual application costs can also depend on how a client submits reference images and how the Microsoft Foundry service accounts for input and output usage. Teams should verify the live pricing page before committing to production volumes because the model is in preview and prices may change.
The pricing structure makes MAI-Image-2.5-Pro a better fit for high-value images where quality, editing precision, or consistency justifies the expense than for indiscriminate bulk generation. For large-scale thumbnail creation, rapid experimentation, or low-cost variations, a faster or less expensive image option may be more appropriate, although the supplied research does not identify a specific alternative with directly comparable pricing.
Reasoning, coding, and tool support
The model's documented reasoning strengths are visual rather than language-agent capabilities. Microsoft describes spatial and geometric reasoning, material and physical-property accuracy, and visual reasoning. These capabilities help it interpret relationships within an image, such as the placement of objects, the geometry of a scene, or the appearance of materials.
The catalog gives MAI-Image-2.5-Pro a comparative reasoning score of 7, coding score of 1, and speed score of 4. These are editorial estimates, not Microsoft-published benchmark results. The coding score should not be interpreted as a measure of the model's ability to write software: coding is not its purpose, and the model does not produce text as its primary output.
No tool use, function calling, streaming, fine-tuning, caching, or batch API support is documented in the supplied specification. There is also no native web grounding listed for this model. Developers should treat it as an image-generation and image-editing endpoint rather than an autonomous agent or general-purpose API model.
Limitations and trade-offs
The first limitation is specialization. MAI-Image-2.5-Pro is not the right choice for text-only analysis, software development, audio generation, video creation, or current-information research. It produces images rather than a conversational answer, and the supplied catalog marks text, audio, and video output as unsupported.
Second, the model is not described as a low-latency, high-volume image generator. Its comparative speed score is 4 and its cost score is 4, both editorial estimates on the supplied scale. Those scores are not formal performance measurements, but they reinforce the practical positioning: this is a quality-oriented model, not one that should automatically be used for every inexpensive or time-sensitive generation task.
Third, the model remains in public preview. Preview availability can involve changing limits, pricing, model versions, or service behavior. Production users should test the exact Foundry deployment, confirm regional availability, monitor image quality on their own prompts, and retain human review for commercial assets and text-sensitive images.
Finally, high-fidelity generation does not eliminate the need for verification. Generated typography, object details, proportions, and physical materials should be checked before publication. The model's documented capabilities describe its intended behavior and strengths, not a guarantee that every output will satisfy a production brief.
When to choose MAI-Image-2.5-Pro
Choose MAI-Image-2.5-Pro when the central requirement is a visually detailed image that needs controlled composition or editing. It is a strong candidate for:
- Commercial and photorealistic creative work.
- Hero images and production-design concepts.
- Product imagery involving specific materials, layouts, or spatial relationships.
- Image edits that must preserve objects or characters from references.
- Scenes with signs, labels, packaging, or other important in-image text.
- Applications that need up to five reference images for a coordinated edit.
Consider another type of model when the priority is low cost, high throughput, very fast response, text conversation, coding, audio or video output, or web-grounded research. A general-purpose multimodal assistant may be more suitable when the workflow requires discussion and analysis around images. A lower-cost image generator may be preferable for rough concepts or large batches. A model with documented web grounding may be better when current facts must influence the result. Those alternatives involve capability trade-offs, and the supplied research does not provide enough information to name or rank a specific competing model.
Bottom line
MAI-Image-2.5-Pro is a specialized Microsoft image model for users who value visual fidelity, controlled editing, typography, and consistency more than minimal generation cost. Its support for text prompts, reference-image editing, up to five input images, and detailed spatial and material control gives it a clear role in professional creative workflows. The main reasons to hesitate are its preview status, image-output cost, lack of documented web or tool support, and limited relevance outside visual generation and editing.

