MAI-Thinking

MAI-Thinking-1

by Microsoft Copilot · Public preview

Microsoft MAI-Thinking-1 is a public-preview, text-only reasoning model for complex mathematics, software engineering, quantitative enterprise analysis, and long-context document work. It uses a sparse Mixture-of-Experts architecture with 35 billion active parameters and approximately 1 trillion total parameters, supports a 256K-token context window and up to 64K output tokens, and is available through Microsoft Foundry Global Standard deployment. Numeric pricing was not exposed in the reviewed official pricing page, and the preview has no service-level agreement.

Text Reasoning Coding
Microsoft MAI-Thinking-1 is a text-only reasoning model for demanding tasks that benefit from deliberate, multi-step problem solving. It targets enterprise mathematics, coding, quantitative analysis, and long-context document work rather than image, audio, or video generation. The model is available through Microsoft Foundry in public preview, with Global Standard deployment and no service-level agreement.
Outputs

What MAI-Thinking-1 can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Streaming
Model profile

Performance characteristics

9/10 Reasoning
9/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family MAI-Thinking
Model type Reasoning
Context window 256K tokens
Maximum output 64K tokens
Release date 2026-08-12
Status Public preview
Knowledge cutoff notes

Microsoft’s reviewed model page, model card, deployment documentation, and technical report do not state an official knowledge-cutoff date for MAI-Thinking-1. The model’s web-search or retrieval integrations do not establish a knowledge cutoff.

Model notes

MAI-Thinking-1 is a sparse Mixture-of-Experts Transformer with 35B active parameters and approximately 1T total parameters. Microsoft reports that it was trained from scratch without distillation from third-party models. The exact Azure model version is 2026-06-01 and the available deployment type is GlobalStandard. Microsoft Foundry documentation states that the model has a 256K total token budget per request, with output capped at 64K tokens or the remaining context budget, whichever is smaller. Function calling is supported through the Chat Completions API, but the model has no native tool interface; the integrating application mediates tools and is responsible for the security boundary. PTU deployment is currently unsupported. The model is text-only and does not accept or produce image, audio, or video. Pricing is published as a Microsoft Foundry pricing category, but numeric rates were not exposed in the official pricing page reviewed. Editorial scores are comparative estimates based on Microsoft’s reported benchmark results, model positioning, context capacity, and preview availability, not official Microsoft ratings.

Model guide

MAI-Thinking-1: Microsoft’s Long-Context Reasoning Model for Math and Coding

MAI-Thinking-1 is Microsoft AI’s first dedicated reasoning model, built for complex mathematics, software engineering, quantitative analysis, and long-context enterprise workloads. It uses a sparse Mixture-of-Experts architecture with 35 billion active parameters and approximately 1 trillion total parameters, supports a 256K-token context window and up to 64K output tokens, and is available in public preview through Microsoft Foundry.

What is MAI-Thinking-1?

MAI-Thinking-1 is Microsoft AI’s first dedicated reasoning model. It is designed to spend more effort working through difficult problems before producing an answer, making it a better fit for complex mathematics, software engineering, quantitative analysis, and other tasks where a quick but shallow response may be insufficient.

The model is a sparse Mixture-of-Experts, or MoE, Transformer. In simple terms, it contains a very large collection of learned parameters but activates only part of them for each request. Microsoft describes MAI-Thinking-1 as having 35 billion active parameters and approximately 1 trillion total parameters. The sparse design is intended to provide substantial model capacity without requiring every parameter to be used for every piece of text.

MAI-Thinking-1 belongs to the MAI-Thinking model family and is offered through Microsoft Foundry rather than as a general consumer chatbot. Its documented access path is a Chat Completions-compatible API using Global Standard deployment. The model was released on August 12, 2026, and is listed as being in public preview.

Core specifications at a glance

SpecificationMAI-Thinking-1
ProviderMicrosoft AI
Model typeReasoning model
ArchitectureSparse Mixture-of-Experts Transformer
Active parameters35 billion
Total parametersApproximately 1 trillion
Context window256,000 tokens
Maximum output64,000 tokens, or the remaining context budget if smaller
Input and outputText in, text out
DeploymentMicrosoft Foundry Global Standard
AvailabilityPublic preview

The 256K figure is a total per-request token budget. The maximum output is capped at 64K tokens, but a request with a long input leaves less room for the response. The practical output limit is therefore the smaller of 64K tokens or the unused portion of the context budget.

Reasoning and coding capabilities

MAI-Thinking-1 is primarily intended for problems that require multiple connected steps. Suitable examples include proving or checking mathematical work, analyzing quantitative assumptions, reading a large codebase, proposing an implementation, debugging failures, and revising code after tests expose a problem.

Microsoft says the training and post-training process emphasized mathematical reasoning, software engineering, tool-use environments, instruction following, factuality, and reducing unnecessary refusals. The model was trained from scratch on clean, traceable, appropriately licensed data, and Microsoft states that it did not use distillation from third-party models.

For coding teams, the model’s value is not limited to generating a short code snippet. Its long context can be useful when a task involves multiple files, extensive logs, technical requirements, or a lengthy test history. It can also be used for code reading, editing, testing, debugging, and recovery from failed approaches. These are intended use cases rather than a guarantee that every generated solution will compile or be correct.

Microsoft reports results of 97.0% on AIME 2025, 94.5% on AIME 2026, 52.8% on SWE-Bench Pro, and 87.7% on LiveCodeBench v6. These figures are provider-reported results. They should be treated as evidence of the model’s positioning, not as universal predictions of application performance, because benchmark configurations, prompts, tool access, and evaluation conditions can differ.

A 256K context window for long enterprise tasks

MAI-Thinking-1’s 256K-token context window is one of its most practical specifications. A context window is the amount of text the model can consider in one request, including the prompt and generated response. A large window can reduce the need to split a long document, codebase extract, contract set, research collection, or incident history into many separate conversations.

Long context does not automatically mean perfect recall or analysis. Users still need to organize source material, identify the relevant sections, and verify important conclusions. The model can produce up to 64K output tokens, subject to the remaining context budget, but a very long answer is not always the most useful answer. For production applications, response length should be controlled according to the task and downstream processing requirements.

Text-only input and output, with application-managed tools

MAI-Thinking-1 accepts text and produces text. It does not natively accept or generate images, audio, or video. It is therefore not the appropriate choice for visual document understanding, image generation, speech interaction, video analysis, or media creation.

The model supports function calling through the Chat Completions API, but Microsoft’s documentation distinguishes this from having a native tool interface. An integrating application mediates the function calls and connects the model to external systems. That application is responsible for deciding which tools are available, validating arguments, handling returned data, and enforcing the security boundary.

This setup can support agent-style workflows such as querying an internal database, running a test command, retrieving documents, or passing structured results back to the model. It does not mean that MAI-Thinking-1 should be allowed to operate an unrestricted autonomous system. Microsoft says the model was not designed or evaluated as an autonomous decision-maker for consequential legal, financial, medical, employment, educational, housing, credit, or safety-critical decisions.

Availability, deployment, and preview status

MAI-Thinking-1 is available in public preview through Microsoft Foundry using Global Standard deployment. The exact Azure model version documented for the model is 2026-06-01. Provisioned Throughput Unit, or PTU, deployment is currently unsupported.

The preview status matters for organizations evaluating the model. Microsoft states that the preview does not include a service-level agreement and is not recommended for production workloads. Availability, limits, behavior, and pricing may change as the service develops. Teams can still use the preview for experiments, internal evaluation, benchmark testing, and carefully controlled prototypes, but should avoid treating the preview as a stable production dependency without an appropriate fallback.

Pricing and cost considerations

Microsoft lists MAI-Thinking-1 within the Microsoft Foundry pricing catalog, but the official pricing page reviewed for this profile did not expose numeric input or output rates. A verified per-token price therefore cannot be provided here.

The model’s sparse architecture may offer an efficiency advantage compared with a dense model of similar total parameter capacity, but the supplied documentation does not establish a specific price or speed advantage. Cost will depend on Microsoft Foundry’s current rates, request size, generated output, deployment conditions, and any applicable account or regional terms. Organizations should check the live Foundry pricing page before committing to a budget.

MAI-Thinking-1 should be viewed as a reasoning-focused option rather than automatically the cheapest or fastest model for every request. Its intended value is in difficult tasks where better multi-step analysis, coding performance, or long-context handling can justify additional latency or expense. A smaller, faster model may be more appropriate for simple classification, short extraction, routine chat, or high-volume low-complexity requests.

Best use cases

  • Mathematical and scientific reasoning: solving multi-step problems, checking derivations, and explaining quantitative conclusions.
  • Software engineering: reviewing code, debugging, writing tests, analyzing failures, and working across larger technical contexts.
  • Quantitative enterprise analysis: financial modeling, statistical analysis, market sizing, forecasting, and examination of assumptions.
  • Long-context document work: comparing large text collections, synthesizing technical material, and answering questions over extended project records.
  • Application-managed agent workflows: combining reasoning with approved functions, retrieval systems, databases, or testing tools under application control.

When to choose MAI-Thinking-1

Choose MAI-Thinking-1 when the central challenge is difficult text-based reasoning rather than media processing. It is a strong candidate for teams that need a large context window, substantial maximum output, coding assistance, quantitative analysis, or an application-managed workflow with function calling.

It may be preferable to a general-purpose fast model when a task requires several dependent reasoning steps, a large amount of source material, or repeated debugging. It may also be a better fit than a multimodal model when the workload is entirely textual and the priority is mathematical or software-engineering performance.

Another option is more appropriate when the task requires images, audio, video, or native media understanding; when the application needs a generally available production model with an SLA; or when predictable pricing and high-throughput speed matter more than advanced reasoning. A human review process is also necessary for high-impact decisions, regardless of the model’s benchmark performance.

Limitations to consider

The most important limitation is its public-preview status. There is no service-level agreement, PTU deployment is unsupported, and Microsoft does not recommend the model for production workloads in this stage. Feature behavior and pricing may change.

It is also strictly text-only, so files containing visual, audio, or video information require separate preprocessing or another model. Function calling is mediated by the host application rather than provided as an unrestricted native tool environment. That makes application design, permissions, validation, and monitoring particularly important.

Finally, strong reasoning benchmarks do not eliminate factual errors, coding defects, security problems, or unsuitable recommendations. Outputs should be tested against source material and evaluated by qualified people before they influence consequential decisions.

Bottom line

MAI-Thinking-1 is a specialized Microsoft reasoning model for complex text-based work. Its defining combination is a sparse 35-billion-active-parameter architecture, approximately 1 trillion total parameters, a 256K-token context window, and a 64K-token output ceiling. Those specifications make it relevant to enterprise mathematics, coding, quantitative analysis, and long-context workflows.

Its public-preview status, unavailable verified pricing, lack of PTU deployment, text-only design, and absence of a service-level agreement limit where it should be used today. For controlled evaluation and challenging reasoning tasks in Microsoft Foundry, it is a compelling option; for routine, multimodal, highly time-sensitive, or production-critical workloads, a different model type may be more suitable.


Answers to Frequently Asked Questions

Is MAI-Thinking-1 suitable for production workloads?
Microsoft does not recommend MAI-Thinking-1 for production workloads while it is in public preview. The preview has no service-level agreement, and its availability, limits, behavior, and pricing may change, so organizations should use appropriate testing, monitoring, and fallback plans.
Where can MAI-Thinking-1 be accessed?
MAI-Thinking-1 is available through Microsoft Foundry using a Chat Completions-compatible API and Global Standard deployment. It is currently listed as a public-preview model, while Provisioned Throughput Unit deployment is unsupported.
Does MAI-Thinking-1 support images, audio, or video?
No. MAI-Thinking-1 accepts text as input and produces text as output. Visual, audio, and video tasks require separate preprocessing or a multimodal model.
What is MAI-Thinking-1 designed for?
MAI-Thinking-1 is Microsoft AI’s reasoning-focused model for complex text-based tasks, including mathematics, software engineering, quantitative analysis, long-document analysis, and application-managed tool workflows.
How large is MAI-Thinking-1’s context window?
MAI-Thinking-1 has a 256,000-token context window. It can generate up to 64,000 output tokens, or less when the input uses much of the available context budget.


Sources 6
Provider

About Microsoft Copilot