MAI-Image

MAI-Image-2.5

by Microsoft Copilot · Preview

Microsoft MAI-Image-2.5 is a preview diffusion model for creating and editing images. It accepts text and up to five reference images through its editing API, returns PNG output, supports photorealistic visuals and text rendering, and is priced by input and output token usage.

Image generation Reasoning Coding
MAI-Image-2.5 is Microsoft AI’s image-generation and editing model for users who need more than a one-shot text-to-image result. It accepts text prompts and reference images, produces PNG images, and is designed for creative production, marketing assets, product visuals, presentation graphics, and localized edits that preserve the rest of an existing composition.
Outputs

What MAI-Image-2.5 can produce

Image generation
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Multimodal output
Model profile

Performance characteristics

6/10 Reasoning
1/10 Coding
6/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family MAI-Image
Model type Other
Context window 131K tokens
Maximum output 4K tokens
Release date 2026-06-02
Status Preview
Knowledge cutoff notes

Microsoft does not publish a direct knowledge-cutoff date for this image-generation model. Its behavior is based on trained visual and multimodal data rather than a documented conversational knowledge boundary.

Model notes

MAI-Image-2.5 is a diffusion-based image model with text-to-image and image-to-image capabilities. Microsoft Foundry lists a 131,072-token context window and 4,096-token output limit, while the MAI image API documentation specifies a 32,000-token maximum prompt context. API output is always PNG. MAI-Image-2.5 supports image edits with up to five JPEG or PNG reference images. For this model, image dimensions must each be at least 768 pixels and total output area must not exceed 1,048,576 pixels. Web grounding is documented for MAI-Image-2.6 models, not MAI-Image-2.5. Preview availability may involve limited regions, quotas, and no service-level agreement. Editorial scores are comparative estimates rather than Microsoft benchmarks.

Cost

Model pricing

Input $5 per 1 million text input tokens; $8 per 1 million image input tokens
Output $47 per 1 million image output tokens
Model guide

MAI-Image-2.5: Microsoft’s Preview Model for Precise Image Generation and Editing

Microsoft MAI-Image-2.5 is a preview diffusion-based image model for high-quality text-to-image generation and controlled image-to-image editing. It supports photorealistic visuals, portraits, product designs, text rendering, object replacement, inpainting, layout preservation, and artifact cleanup through Microsoft Foundry and related Microsoft experiences.

What is MAI-Image-2.5?

MAI-Image-2.5 is a preview diffusion-based model from Microsoft AI. Its primary job is to create images from natural-language descriptions and to modify existing images while keeping important parts of the original composition intact. In practical terms, it can generate a new product scene from a prompt, replace an object in a supplied image, update text on packaging, or clean up a visual artifact such as motion blur.

The model is part of Microsoft’s MAI image-model lineup and is available through Microsoft Foundry, MAI Playground, and selected Microsoft product integrations. Microsoft Foundry provides the documented deployment and API route, while the playground and integrations provide more direct creative access where the model is enabled.

MAI-Image-2.5 is not a general-purpose conversational model. It does not produce native text, audio, video, embeddings, or executable actions as its primary output. Its main output is generated or edited imagery, returned as a PNG through the documented image APIs.

Generation and editing capabilities

The model supports both text-to-image and image-to-image workflows. Text-to-image generation starts with a natural-language prompt describing the desired scene, subject, style, composition, or design. Image editing starts with one or more JPEG or PNG reference images and an instruction describing the change.

Microsoft’s documented capabilities include:

  • Generating photorealistic scenes, portraits, product imagery, branding assets, and commercial designs.
  • Editing existing images through object removal, object replacement, attribute changes, and inpainting.
  • Updating visible text in labels, posters, packaging, signs, and presentation graphics.
  • Preserving layout and much of the surrounding composition during localized edits.
  • Understanding relationships among objects, lighting, scale, and spatial placement.
  • Cleaning up visual artifacts, including some forms of motion blur.
  • Using up to five JPEG or PNG reference images with the edits API.

These capabilities make the model better suited to iterative production work than an image generator that is used only to create entirely new scenes. For example, a design team could begin with a product photograph, ask for a different background, replace a visible prop, and then adjust packaging text without recreating every other element from scratch.

Inputs, outputs, and technical limits

MAI-Image-2.5 accepts text input and image input. The text input can describe a new image or specify an edit to apply to supplied references. The image input is supported for editing workflows and can include up to five JPEG or PNG reference images through the edits API.

The documented API output is a PNG image. The model does not provide text, audio, video, music, embedding, structured-data, speech, or executable-action output. It also is not documented as supporting tool or function calling, streaming, fine-tuning, caching, or batch API access. Those unsupported or undocumented areas matter when evaluating it for a larger application: MAI-Image-2.5 should be treated as a specialized visual-generation endpoint rather than a multimodal agent.

Microsoft Foundry lists a 131,072-token context window and a 4,096-token output limit for the model catalog entry. However, the MAI image API documentation specifies a maximum prompt context of 32,000 tokens for API requests. These figures describe different layers of the offering, so developers should use the API-specific 32,000-token prompt limit when designing requests rather than assuming that the larger catalog context value applies unchanged to image-generation calls.

Image dimensions have their own restrictions. Each output dimension must be at least 768 pixels, while the total image area cannot exceed 1,048,576 pixels. The maximum area is approximately equivalent to a 1024 by 1024 image, although the permitted aspect ratio can vary within the stated dimension and area constraints. The API always returns PNG output.

Pricing and availability

Microsoft announced usage-based pricing of $5 per 1 million text input tokens, $8 per 1 million image input tokens, and $47 per 1 million image output tokens. These are model-usage charges and are separate from any applicable Microsoft Azure subscription, deployment, or infrastructure costs.

MAI-Image-2.5 was released on June 2, 2026, and the supplied documentation identifies it as a preview model. Preview availability can involve limited regions, quotas, or constrained capabilities. Microsoft’s preview guidance also means the service does not carry a service-level agreement suitable for production workloads. Teams considering the model for commercial production should verify current regional availability, quota policies, operational commitments, and pricing before committing to an architecture.

Where the model is strong—and where it is not

MAI-Image-2.5’s strongest use case is controlled visual creation. It combines new-image generation with localized editing, so it can serve both ideation and production-oriented workflows. Its emphasis on photorealism, identity preservation, text rendering, and fine-grained edits is particularly relevant to product marketing, advertising concepts, packaging mockups, presentation graphics, and visual design iteration.

The model is less appropriate when the task primarily involves conversation, code generation, document reasoning, real-time web research, audio, video, or autonomous actions. It has no documented tool or function support and does not return a text response as its principal result. A general-purpose language model or a dedicated audio or video system would be a more suitable choice for those tasks.

There is also a cost and speed trade-off. The supplied editorial evaluation rates its speed and cost position at 6 out of 10, but these are comparative editorial estimates, not Microsoft-published benchmarks. The model’s value is therefore less about being the cheapest or fastest possible image service and more about combining high-quality generation with controlled editing. Actual response times and total costs can vary with image inputs, request size, service capacity, deployment configuration, and usage limits.

Reasoning, coding, and tool support

MAI-Image-2.5 is not designed as a reasoning or coding model. The catalog data assigns it a comparative editorial reasoning score of 6 out of 10 and a coding score of 1 out of 10; these scores are editorial estimates rather than provider benchmarks. The reasoning score should not be interpreted as evidence that the model can perform general-purpose chain-of-thought reasoning or solve language tasks like a large language model.

Its useful form of reasoning is visual and compositional: interpreting relationships between objects, lighting, scale, spatial layout, and requested edits. It can use those visual relationships to make a localized change while attempting to preserve the rest of the image. That is different from supporting programming, external tools, structured function calls, or autonomous workflows.

Safety and implementation limitations

Generated images can contain inaccurate or misleading details and may reflect biases present in training data. Microsoft recommends reviewing outputs before using them in sensitive identity, legal, medical, financial, or news-related contexts. Users remain responsible for safety, privacy, copyright, and disclosure controls even when platform-level prompt and output filtering is applied.

Photorealistic edits involving minors are blocked by default in Microsoft Foundry, with access subject to Microsoft’s stated approval process. This restriction is important for teams working with portrait editing or identity-preserving workflows. The model’s preview status is another practical limitation: feature availability, quotas, regions, and operational guarantees may change.

Image editing also remains an inherently imperfect process. A request to preserve a person, product, or layout does not guarantee pixel-level identity or composition preservation. Outputs should be inspected for altered text, distorted objects, inconsistent lighting, inaccurate anatomy, and other visual changes before publication.

When to choose MAI-Image-2.5

Choose MAI-Image-2.5 when the workflow needs both image generation and controlled editing, especially when reference images, localized changes, photorealistic output, or rendered text are important. It is a reasonable candidate for:

  • Product and advertising concept development.
  • Marketing images and branded creative assets.
  • Packaging, signage, poster, and presentation mockups.
  • Portrait creation and identity-conscious visual editing, subject to applicable restrictions.
  • Production design and iterative art direction.
  • Object replacement, removal, inpainting, and artifact cleanup.

Consider another option when the priority is the lowest possible image cost, guaranteed production availability, a general-purpose text interface, coding, audio or video generation, web-grounded research, or tool-driven automation. MAI-Image-2.5 can be attractive for visual quality and edit control, but its preview status, specialized output format, API limits, and lack of documented agent features make it a poor fit for applications that need broader model capabilities.

Bottom line

MAI-Image-2.5 is a specialized Microsoft image model aimed at high-quality creation and precise visual modification rather than general-purpose AI assistance. Its combination of text-to-image generation, reference-image editing, layout awareness, object manipulation, and improved text rendering gives it a practical role in design and commercial content workflows. Developers should account for the 32,000-token API prompt limit, five-image edit limit, PNG-only output, 768-pixel minimum dimensions, 1,048,576-pixel maximum area, usage-based pricing, preview availability, and the need to review every output before sensitive or public use.


Answers to Frequently Asked Questions

What limitations should developers consider before using MAI-Image-2.5?
MAI-Image-2.5 is a specialized image model rather than a general-purpose conversational or agentic system. It has no documented tool or function calling, streaming, fine-tuning, caching, or batch API support, and it is available as a preview model without production-level service guarantees. Generated images should also be reviewed for inaccurate text, distorted objects, bias, privacy issues, and other visual errors.
How much does MAI-Image-2.5 cost?
Microsoft announced usage-based pricing of $5 per 1 million text input tokens, $8 per 1 million image input tokens, and $47 per 1 million image output tokens. These charges are separate from any applicable Azure subscription, deployment, or infrastructure costs.
What are the image size and prompt limits for MAI-Image-2.5?
Each output dimension must be at least 768 pixels, and the total image area cannot exceed 1,048,576 pixels. Although Microsoft Foundry lists a larger catalog context window, the image API documentation specifies a maximum prompt context of 32,000 tokens for API requests.
What is MAI-Image-2.5 used for?
MAI-Image-2.5 is a Microsoft preview model for generating images from text and editing existing images. It is designed for photorealistic scenes, product imagery, marketing assets, packaging mockups, object replacement, inpainting, text updates, and visual artifact cleanup.
What image inputs and outputs does MAI-Image-2.5 support?
The model accepts text prompts and JPEG or PNG reference images for editing workflows. The edits API supports up to five reference images, and the documented output is always a PNG image.


Sources 5
Provider

About Microsoft Copilot