What is MAI-Image-2.5?
MAI-Image-2.5 is a preview diffusion-based model from Microsoft AI. Its primary job is to create images from natural-language descriptions and to modify existing images while keeping important parts of the original composition intact. In practical terms, it can generate a new product scene from a prompt, replace an object in a supplied image, update text on packaging, or clean up a visual artifact such as motion blur.
The model is part of Microsoft’s MAI image-model lineup and is available through Microsoft Foundry, MAI Playground, and selected Microsoft product integrations. Microsoft Foundry provides the documented deployment and API route, while the playground and integrations provide more direct creative access where the model is enabled.
MAI-Image-2.5 is not a general-purpose conversational model. It does not produce native text, audio, video, embeddings, or executable actions as its primary output. Its main output is generated or edited imagery, returned as a PNG through the documented image APIs.
Generation and editing capabilities
The model supports both text-to-image and image-to-image workflows. Text-to-image generation starts with a natural-language prompt describing the desired scene, subject, style, composition, or design. Image editing starts with one or more JPEG or PNG reference images and an instruction describing the change.
Microsoft’s documented capabilities include:
- Generating photorealistic scenes, portraits, product imagery, branding assets, and commercial designs.
- Editing existing images through object removal, object replacement, attribute changes, and inpainting.
- Updating visible text in labels, posters, packaging, signs, and presentation graphics.
- Preserving layout and much of the surrounding composition during localized edits.
- Understanding relationships among objects, lighting, scale, and spatial placement.
- Cleaning up visual artifacts, including some forms of motion blur.
- Using up to five JPEG or PNG reference images with the edits API.
These capabilities make the model better suited to iterative production work than an image generator that is used only to create entirely new scenes. For example, a design team could begin with a product photograph, ask for a different background, replace a visible prop, and then adjust packaging text without recreating every other element from scratch.
Inputs, outputs, and technical limits
MAI-Image-2.5 accepts text input and image input. The text input can describe a new image or specify an edit to apply to supplied references. The image input is supported for editing workflows and can include up to five JPEG or PNG reference images through the edits API.
The documented API output is a PNG image. The model does not provide text, audio, video, music, embedding, structured-data, speech, or executable-action output. It also is not documented as supporting tool or function calling, streaming, fine-tuning, caching, or batch API access. Those unsupported or undocumented areas matter when evaluating it for a larger application: MAI-Image-2.5 should be treated as a specialized visual-generation endpoint rather than a multimodal agent.
Microsoft Foundry lists a 131,072-token context window and a 4,096-token output limit for the model catalog entry. However, the MAI image API documentation specifies a maximum prompt context of 32,000 tokens for API requests. These figures describe different layers of the offering, so developers should use the API-specific 32,000-token prompt limit when designing requests rather than assuming that the larger catalog context value applies unchanged to image-generation calls.
Image dimensions have their own restrictions. Each output dimension must be at least 768 pixels, while the total image area cannot exceed 1,048,576 pixels. The maximum area is approximately equivalent to a 1024 by 1024 image, although the permitted aspect ratio can vary within the stated dimension and area constraints. The API always returns PNG output.
Pricing and availability
Microsoft announced usage-based pricing of $5 per 1 million text input tokens, $8 per 1 million image input tokens, and $47 per 1 million image output tokens. These are model-usage charges and are separate from any applicable Microsoft Azure subscription, deployment, or infrastructure costs.
MAI-Image-2.5 was released on June 2, 2026, and the supplied documentation identifies it as a preview model. Preview availability can involve limited regions, quotas, or constrained capabilities. Microsoft’s preview guidance also means the service does not carry a service-level agreement suitable for production workloads. Teams considering the model for commercial production should verify current regional availability, quota policies, operational commitments, and pricing before committing to an architecture.
Where the model is strong—and where it is not
MAI-Image-2.5’s strongest use case is controlled visual creation. It combines new-image generation with localized editing, so it can serve both ideation and production-oriented workflows. Its emphasis on photorealism, identity preservation, text rendering, and fine-grained edits is particularly relevant to product marketing, advertising concepts, packaging mockups, presentation graphics, and visual design iteration.
The model is less appropriate when the task primarily involves conversation, code generation, document reasoning, real-time web research, audio, video, or autonomous actions. It has no documented tool or function support and does not return a text response as its principal result. A general-purpose language model or a dedicated audio or video system would be a more suitable choice for those tasks.
There is also a cost and speed trade-off. The supplied editorial evaluation rates its speed and cost position at 6 out of 10, but these are comparative editorial estimates, not Microsoft-published benchmarks. The model’s value is therefore less about being the cheapest or fastest possible image service and more about combining high-quality generation with controlled editing. Actual response times and total costs can vary with image inputs, request size, service capacity, deployment configuration, and usage limits.
Reasoning, coding, and tool support
MAI-Image-2.5 is not designed as a reasoning or coding model. The catalog data assigns it a comparative editorial reasoning score of 6 out of 10 and a coding score of 1 out of 10; these scores are editorial estimates rather than provider benchmarks. The reasoning score should not be interpreted as evidence that the model can perform general-purpose chain-of-thought reasoning or solve language tasks like a large language model.
Its useful form of reasoning is visual and compositional: interpreting relationships between objects, lighting, scale, spatial layout, and requested edits. It can use those visual relationships to make a localized change while attempting to preserve the rest of the image. That is different from supporting programming, external tools, structured function calls, or autonomous workflows.
Safety and implementation limitations
Generated images can contain inaccurate or misleading details and may reflect biases present in training data. Microsoft recommends reviewing outputs before using them in sensitive identity, legal, medical, financial, or news-related contexts. Users remain responsible for safety, privacy, copyright, and disclosure controls even when platform-level prompt and output filtering is applied.
Photorealistic edits involving minors are blocked by default in Microsoft Foundry, with access subject to Microsoft’s stated approval process. This restriction is important for teams working with portrait editing or identity-preserving workflows. The model’s preview status is another practical limitation: feature availability, quotas, regions, and operational guarantees may change.
Image editing also remains an inherently imperfect process. A request to preserve a person, product, or layout does not guarantee pixel-level identity or composition preservation. Outputs should be inspected for altered text, distorted objects, inconsistent lighting, inaccurate anatomy, and other visual changes before publication.
When to choose MAI-Image-2.5
Choose MAI-Image-2.5 when the workflow needs both image generation and controlled editing, especially when reference images, localized changes, photorealistic output, or rendered text are important. It is a reasonable candidate for:
- Product and advertising concept development.
- Marketing images and branded creative assets.
- Packaging, signage, poster, and presentation mockups.
- Portrait creation and identity-conscious visual editing, subject to applicable restrictions.
- Production design and iterative art direction.
- Object replacement, removal, inpainting, and artifact cleanup.
Consider another option when the priority is the lowest possible image cost, guaranteed production availability, a general-purpose text interface, coding, audio or video generation, web-grounded research, or tool-driven automation. MAI-Image-2.5 can be attractive for visual quality and edit control, but its preview status, specialized output format, API limits, and lack of documented agent features make it a poor fit for applications that need broader model capabilities.
Bottom line
MAI-Image-2.5 is a specialized Microsoft image model aimed at high-quality creation and precise visual modification rather than general-purpose AI assistance. Its combination of text-to-image generation, reference-image editing, layout awareness, object manipulation, and improved text rendering gives it a practical role in design and commercial content workflows. Developers should account for the 32,000-token API prompt limit, five-image edit limit, PNG-only output, 768-pixel minimum dimensions, 1,048,576-pixel maximum area, usage-based pricing, preview availability, and the need to review every output before sensitive or public use.

