MAI-Image

MAI-Image-2.6

by Microsoft Copilot · Public Preview

Microsoft MAI-Image-2.6 is a Public Preview image model for text-to-image generation and image-to-image editing. It accepts multiple visual references, supports optional Bing web grounding and automatic aspect-ratio selection, and returns PNG images up to a documented total area of 2,359,296 pixels through Microsoft Foundry.

Image generation Reasoning Coding
MAI-Image-2.6 is Microsoft AI’s flagship model in the MAI-Image family for creating and editing images. Rather than serving as a general-purpose text or reasoning model, it is designed for visual production: generating images from prompts, modifying supplied images, combining several references, and producing commercial, photorealistic, branded, and design-oriented visuals. The model is available in Microsoft Foundry Public Preview and through Microsoft AI’s MAI Playground.
Outputs

What MAI-Image-2.6 can produce

Image generation
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Web search Multimodal output
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family MAI-Image
Model type Multimodal
Release date 2026-08-10
Status Public Preview
Knowledge cutoff notes

No exact knowledge cutoff for MAI-Image-2.6 was disclosed in the authoritative sources reviewed. Web grounding is an optional runtime capability and does not establish or change the underlying training-data cutoff.

Model notes

Microsoft announced MAI-Image-2.6 on August 10, 2026, and listed it in Microsoft Foundry Public Preview with model version 2026-07-31. It supports text-to-image generation and image-to-image editing, including multi-image reference editing. The web_grounding option can retrieve current context from Bing Search, while auto_aspect_ratio lets the model select an output format. Outputs are PNG images returned as base64-encoded data in a JSON API response; the JSON wrapper does not represent native structured-text output. Width and height must each be at least 768 pixels, and the total image area cannot exceed 2,359,296 pixels. Editorial scores are comparative estimates for an image-generation model and are not Microsoft benchmarks. Public pricing for the exact 2.6 model was not verifiable in the current official pricing source.

Model guide

MAI-Image-2.6: Microsoft’s Public-Preview Model for High-Fidelity Image Creation and Editing

MAI-Image-2.6 is Microsoft AI’s public-preview image model for text-to-image generation and image-to-image editing. It supports multiple reference images, Bing web grounding, automatic aspect-ratio selection, improved text rendering, and outputs up to approximately 1.5K resolution through Microsoft Foundry.

What is MAI-Image-2.6?

MAI-Image-2.6 is an image-generation and image-editing model developed by Microsoft AI. It accepts natural-language instructions and, for editing workflows, one or more reference images. The model can create a new image from a description, alter an existing image, or combine visual elements from several inputs.

Microsoft positions MAI-Image-2.6 for professional creative work, commercial imagery, photorealistic scenes, product visualization, branding, portraits, 3D imagery, and design tasks that benefit from accurate composition and readable text inside the image. It belongs to Microsoft’s MAI-Image model family and is presented as a higher-capability successor to the MAI-Image-2.5 series, with support for higher resolutions and dynamic aspect ratios.

Microsoft announced MAI-Image-2.6 on August 10, 2026. It became available to developers through Microsoft Foundry Public Preview on September 4, 2026. The model is also accessible through Microsoft AI’s MAI Playground.

What the model can do

  • Text-to-image generation: Creates an image from a written description, such as a product scene, advertising concept, editorial illustration, or photorealistic environment.
  • Image-to-image editing: Uses supplied images as references for controlled changes rather than starting entirely from text.
  • Multi-image reference editing: Can combine several references, including people, products, styles, locations, and other visual elements.
  • Web grounding: When enabled, the web_grounding option can retrieve current context through Bing Search. This is intended for subjects involving real-world entities, places, events, or other information that may change over time.
  • Automatic aspect-ratio selection: With auto_aspect_ratio, the model can select an output format based on the prompt and supplied images.
  • Improved text rendering: Microsoft highlights stronger handling of text in images, which is relevant to posters, packaging, signs, advertisements, presentation graphics, and branded assets.

These capabilities make MAI-Image-2.6 more suitable for guided visual production than for simple one-off experimentation. A user can provide a product reference, a person reference, and a style reference, then ask for a coherent commercial scene while preserving important visual relationships.

Image size and technical limits

Microsoft Foundry lists MAI-Image-2.6 as a Global Standard preview model with model version 2026-07-31. The documented image-generation and image-editing APIs return PNG image data encoded as base64 inside a JSON response. The JSON wrapper is an API transport format; it does not mean that the model produces native structured text or JSON as an output modality.

Each output dimension must be at least 768 pixels. The total image area cannot exceed 2,359,296 pixels. That allows a maximum square output of approximately 1536 by 1536 pixels, although the exact width and height can vary when using other aspect ratios. Microsoft describes this as support for outputs up to approximately 1.5K resolution.

The supplied documentation does not specify a token context window, maximum text-token output, knowledge cutoff, audio capability, video capability, or a general-purpose reasoning limit. Those omissions are important: MAI-Image-2.6 should be evaluated as a visual model, not as a replacement for a text-based language model.

Supported inputs and outputs

AreaDocumented support
Text inputYes; natural-language prompts guide image generation and editing.
Image inputYes; one or more reference images can be used for editing and composition.
Image outputYes; generated PNG images are returned as base64-encoded data through the API.
Audio input or outputNot documented for this model.
Video input or outputNot documented for this model.
Text outputNot a primary model output; the API returns image data in a JSON response.
Structured text outputNot documented. The JSON response envelope should not be confused with a structured-output mode.

MAI-Image-2.6 therefore has multimodal input in the practical sense that it can accept both text and images, but its direct creative output is visual. It does not provide the text reasoning, code generation, speech, video, or embedding functions associated with broader model platforms.

Web grounding is not general tool use

The model supports an optional web-grounding feature that can use Bing Search to retrieve current information for an image request. This can help when a prompt refers to a current event, a real location, or a changing real-world subject. However, the supplied model record does not document general function calling, arbitrary tool execution, code execution, or autonomous multi-step action.

Web grounding also does not change the model’s underlying training-data cutoff. Microsoft has not disclosed an exact knowledge cutoff for MAI-Image-2.6. Grounding is a runtime retrieval option, so results should still be checked for accuracy and suitability before publication.

Strengths and practical uses

MAI-Image-2.6 is most distinctive when a visual task requires both generation quality and control over references. Its multi-image editing support can be useful when a result must preserve several kinds of information at once: the shape of a product, the appearance of a person, a particular visual style, and a target environment.

  • Product and brand imagery: Create campaign concepts, packaging scenes, product mockups, and advertising visuals while supplying product references.
  • Creative production: Explore compositions, art directions, character concepts, editorial graphics, and visual prototypes.
  • Photorealistic scenes: Generate realistic environments, portraits, and commercial compositions from detailed prompts.
  • Text-heavy graphics: Produce posters, signs, promotional layouts, and branded concepts where legible image text matters.
  • Controlled editing: Modify an existing image while using additional references to guide identity, style, objects, or setting.
  • Current-event or location concepts: Use Bing web grounding when the prompt depends on information that may have changed since training.

Microsoft also highlights portraits, 3D imagery, commercial design, and improved prompt fidelity. These are provider claims rather than independent benchmark results. The supplied research does not provide verified comparative benchmark scores, so output quality should be judged using the specific prompts and references in the intended workflow.

Limitations and trade-offs

MAI-Image-2.6 is a Public Preview model. Its availability, pricing, quotas, capabilities, and API behavior may change. Microsoft documents pricing through the Microsoft Foundry pricing system, but a verified public price for the exact 2.6 model was not available in the supplied sources. Cost should therefore be checked in the target Foundry region and deployment configuration before committing to production volume.

The 2,359,296-pixel area limit also matters for print-ready or very large-format work. Although approximately 1536 by 1536 pixels is useful for many digital assets, it may not be sufficient as a final production resolution for every print or display requirement. A separate upscaling or post-processing workflow may be needed, though such a workflow is not part of the documented model capability.

Image models can introduce incorrect visual details, distort text, change identity-related features, or misrepresent real people, places, products, and events. Web grounding can provide current context, but it does not guarantee that the generated image is factually or visually accurate. Human review is especially important for advertising, public communications, sensitive subjects, and images that depict identifiable people.

Reasoning, coding, speed, and cost profile

MAI-Image-2.6 is not intended for text reasoning or software development. Its editorial reasoning score is 2 and coding score is 1 on a comparative internal scale, while its editorial speed score is 7 and cost score is 8. These scores are subjective evaluations for an image-generation model, not Microsoft-published benchmarks or guaranteed service levels.

In practical terms, the model’s value comes from visual quality, reference control, and image-editing flexibility rather than from explaining complex problems or writing code. Its relative speed and cost ratings suggest an attractive profile for image workflows, but actual latency and spending depend on the Foundry service, image dimensions, request volume, region, and preview conditions. No exact price or response-time guarantee should be inferred from the editorial scores.

When to choose MAI-Image-2.6

Choose MAI-Image-2.6 when the main deliverable is an image and the workflow benefits from multiple visual references, controlled editing, readable text, photorealistic or commercial presentation, or flexible aspect ratios. It is particularly relevant for teams already using Microsoft Foundry or Microsoft AI services and for developers who need an image API rather than a conversational text model.

Another type of image model may be more appropriate when you need a documented production price, guaranteed non-preview behavior, substantially larger output dimensions, or specialized video and animation generation. A general-purpose language model is a better choice when the primary task is coding, long-form reasoning, document analysis, or structured text generation. MAI-Image-2.6 can be part of a larger workflow, but it should not be treated as an all-purpose AI model.

Availability and current positioning

Developers can use MAI-Image-2.6 through Microsoft Foundry with Azure authentication and a Microsoft Foundry project in a supported region. It is listed as a Global Standard preview model, and regional availability, quotas, pricing, and access conditions may change. The model is also offered through MAI Playground for users who want to explore its visual capabilities without building an API integration first.

Within Microsoft’s current catalog, MAI-Image-2.6 is a specialized member of the MAI-Image family rather than a general Copilot-style assistant. Its role is to generate and edit visual assets, with the combination of multi-image references, web grounding, dynamic aspect ratios, and approximately 1.5K output as its most relevant practical distinctions.


Answers to Frequently Asked Questions

How is MAI-Image-2.6 accessed and what are its limitations?
Developers can access MAI-Image-2.6 through Microsoft Foundry Public Preview using Azure authentication and a supported Microsoft Foundry region. It is also available through Microsoft AI’s MAI Playground. As a preview model, its pricing, quotas, availability, and API behavior may change. It is intended for visual generation and editing rather than coding, long-form reasoning, video, audio, or structured text generation.
What image sizes and formats does MAI-Image-2.6 support?
Each output dimension must be at least 768 pixels, and the total image area cannot exceed 2,359,296 pixels. This allows an approximately 1536 by 1536 pixel square image, with other aspect ratios also supported. Through the API, generated images are returned as PNG data encoded in base64 within a JSON response.
What is MAI-Image-2.6?
MAI-Image-2.6 is Microsoft AI’s image-generation and image-editing model for creating new images, modifying existing images, and combining elements from multiple reference images. It is designed for professional creative work, including commercial imagery, product visualization, branding, portraits, and design.
What can MAI-Image-2.6 do?
The model supports text-to-image generation, image-to-image editing, multi-image reference editing, improved text rendering, automatic aspect-ratio selection, and optional web grounding through Bing Search. It can be used for product scenes, advertising concepts, photorealistic environments, posters, packaging, signs, and other visual assets.


Sources 6
Provider

About Microsoft Copilot