Phi-3.5

Phi-3.5-MoE-instruct

by Microsoft Copilot · Retired from Azure Foundry on August 30, 2025; downloadable open-weight checkpoint remains available for self-hosted deployment.

Microsoft Phi-3.5-MoE-instruct is a multilingual, text-only open-weight mixture-of-experts model with approximately 42 billion total parameters, 6.6 billion active parameters, and a 128K-token context window. It supports reasoning, coding, mathematics, long-context processing, and local deployment, but its Azure Foundry endpoint was retired on August 30, 2025.

Text Reasoning Coding
Phi-3.5-MoE-instruct is Microsoft’s open-weight mixture-of-experts language model for text generation. It combines 16 experts and approximately 6.6 billion active parameters with a 128K-token context window, making it suitable for multilingual assistants, document analysis, retrieval-augmented generation, coding, mathematics, and other long-context workloads. Microsoft no longer offers the model through Azure Foundry after August 30, 2025, but the downloadable checkpoint remains available for compatible self-hosted deployments.
Outputs

What Phi-3.5-MoE-instruct can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Fine-tuning
Model profile

Performance characteristics

7/10 Reasoning
7/10 Coding
6/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Phi-3.5
Model type General Purpose
Context window 131K tokens
Maximum output 4K tokens
Knowledge cutoff October 2023
Release date August 2024
Status Retired from Azure Foundry on August 30, 2025; downloadable open-weight checkpoint remains available for self-hosted deployment.
Deprecation date August 30, 2025
Shutdown date August 30, 2025
Knowledge cutoff notes

The model card states that the static model was trained on an offline dataset with a cutoff date of October 2023 for publicly available data. Retrieval or search augmentation is needed for current information.

Model notes

The model has approximately 42B total parameters across 16 experts and approximately 6.6B active parameters when two experts are selected. It is text-only, supports a 128K-token context, and has an October 2023 cutoff for publicly available data. Microsoft lists Phi-4-mini-instruct as the Azure replacement after the August 30, 2025 retirement. Historical Azure pricing was approximately $0.16 per 1M input tokens and $0.64 per 1M output tokens. The open-weight checkpoint is MIT licensed and can be run with Transformers or vLLM. Editorial scores are comparative estimates, not vendor-provided ratings.

Cost

Model pricing

Input $0.16 per 1M input tokens historically on Azure; current hosted pricing unavailable after retirement.
Output $0.64 per 1M output tokens historically on Azure; current hosted pricing unavailable after retirement.
Model guide

Phi-3.5-MoE-instruct: Microsoft’s Long-Context Open-Weight Model for Local AI

Microsoft Phi-3.5-MoE-instruct is a multilingual, text-only instruction model with a 128K-token context window and 42 billion total parameters, of which approximately 6.6 billion are active during inference. Its mixture-of-experts design targets reasoning, mathematics, coding, long-context processing, and latency-sensitive applications. The Azure-hosted endpoint was retired on August 30, 2025, but the MIT-licensed checkpoint remains available for local or self-managed deployment.

What is Phi-3.5-MoE-instruct?

Phi-3.5-MoE-instruct is an instruction-tuned language model from Microsoft’s Phi-3.5 family. It accepts text prompts and generates text responses, so it is designed for tasks such as answering questions, following written instructions, summarizing documents, explaining code, solving mathematical problems, and supporting conversational applications.

The model uses a mixture-of-experts, or MoE, architecture. Instead of activating every parameter for every token, the model routes each part of a request through a subset of specialized expert networks. Phi-3.5-MoE-instruct contains approximately 42 billion total parameters distributed across 16 experts, while approximately 6.6 billion parameters are active when two experts are selected for an inference step. This reduces active computation compared with running a dense model containing all 42 billion parameters, although the complete checkpoint still has substantial memory requirements.

Microsoft released the model in August 2024 under the MIT license. The open-weight license and downloadable checkpoint make it relevant to organizations that want more control over deployment, infrastructure, and data handling than a closed hosted model normally provides.

Availability and current status

Phi-3.5-MoE-instruct is no longer an active Azure Foundry hosted model. Microsoft retired the Azure model endpoint on August 30, 2025, and identified Phi-4-mini-instruct as the suggested replacement for that hosted offering. This retirement concerns the Azure service rather than the model weights themselves.

The checkpoint remains available through Microsoft’s Hugging Face repository and can be used with compatible tools such as Transformers and vLLM. As a result, the model is best understood today as an open-weight model for self-managed or third-party-supported inference rather than as a current Microsoft-hosted API choice. Users should distinguish the historical Azure pricing and limits from the model’s continuing availability as downloadable software.

Technical specifications

SpecificationDetails
ProviderMicrosoft
Model familyPhi-3.5
ReleaseAugust 2024
ArchitectureDecoder-only Transformer mixture of experts
Total parametersApproximately 42 billion
Active parametersApproximately 6.6 billion when two experts are active
Context window131,072 tokens, commonly described as 128K tokens
Maximum Azure output4,096 tokens
InputText only
OutputGenerated text only
LicenseMIT
Knowledge cutoffOctober 2023 for publicly available data, according to the model card

The 128K-token context window is the model’s main architectural advantage for long prompts. It can provide room for large documents, conversation histories, codebases, or retrieved reference material in a single request. The practical limit still depends on the inference framework, available memory, prompt formatting, and the deployment environment.

Capabilities and best use cases

Microsoft designed Phi-3.5-MoE-instruct for instruction following, reasoning, mathematics, coding, multilingual generation, and general-purpose assistant tasks. Its 128K context makes it particularly relevant when a request depends on more source material than a smaller context window can comfortably contain.

  • Long-context document work: Summarization, question answering, extraction, and comparison across lengthy text collections.
  • Multilingual applications: The model supports Arabic, Chinese, Czech, Danish, Dutch, English, Finnish, French, German, Hebrew, Hungarian, Italian, Japanese, Korean, Norwegian, Polish, Portuguese, Russian, Spanish, Swedish, Thai, Turkish, and Ukrainian.
  • Code assistance: Code explanation, generation, transformation, and debugging-oriented conversations.
  • Mathematics and reasoning: Step-by-step problem solving and instruction-based analytical tasks, with the usual need to verify important results.
  • Retrieval-augmented generation: Combining the model with a search or retrieval system so that responses can use an organization’s documents or more current information.
  • Controlled local applications: Self-hosted assistants and text-generation systems where deployment control or an open license is important.

The model card and supplied research support these use cases as intended capabilities, but they do not guarantee a particular accuracy level for every domain. A model’s performance can vary with prompt design, language, document quality, retrieval quality, and the hardware and inference software used to run it.

Modalities, tools, and output control

Phi-3.5-MoE-instruct is text-only. It does not natively accept images, audio, or video, and it does not generate images, audio, music, or video. Applications can place descriptions or transcriptions into a text prompt, but that is different from native multimodal understanding.

The supplied model data does not identify native tool or function-calling support. It should therefore not be selected when an application requires a model with a documented built-in tool-use interface. Developers can still build application-level workflows around generated text, such as having software parse a response and then perform an action, but that is an external orchestration layer rather than a verified native model feature.

The model produces text and has no verified native structured-output or JSON-mode capability in the supplied research. A deployment can request a particular format through prompting or enforce formatting in application code, but reliable schema enforcement should not be assumed.

Reasoning, coding, speed, and cost trade-offs

Phi-3.5-MoE-instruct was positioned for reasoning, mathematics, and coding, and its MoE design activates fewer parameters than the total checkpoint size. In principle, that gives it a more favorable active-computation profile than a dense model with the same total parameter count. It does not, however, make the model lightweight to host: loading and serving a roughly 42-billion-parameter checkpoint can still require substantial memory and high-end hardware.

Editorial comparative assessments supplied with the model rate reasoning and coding at 7 out of 10, speed at 6 out of 10, and cost at 8 out of 10. These are subjective editorial scores, not Microsoft benchmark results or official performance guarantees. They reflect the model’s intended balance of capability, active computation, and open deployment economics rather than a universal ranking.

For local operators, the trade-off is therefore between capability and infrastructure. Phi-3.5-MoE-instruct offers a long context and a relatively efficient active path, but smaller dense models may be easier to run on limited hardware. A hosted model may also be more convenient when the priority is avoiding GPU procurement, model serving, and operational maintenance. Conversely, self-hosting can be preferable when data control, customization, or predictable access matters more than operational simplicity.

Limitations and risks

The most important functional limitation is its text-only design. It is not a direct choice for visual question answering, speech processing, audio generation, or video workflows. Those applications require additional modality-specific systems or a different model.

The model’s knowledge cutoff is October 2023 for publicly available data, according to its model card. It does not automatically know events, documents, or changes that occurred after that point. Current-information applications should connect it to retrieval or search augmentation and should still verify important answers.

Microsoft’s model documentation warns that the model can produce factual errors. Long context does not eliminate hallucinations: placing more text in a prompt can help ground an answer, but the model may still overlook relevant passages, misunderstand instructions, or generate unsupported conclusions. High-stakes uses should include retrieval checks, output validation, and human review.

Hardware is another practical constraint. Although only approximately 6.6 billion parameters are active in the stated two-expert configuration, the full checkpoint is approximately 42 billion parameters. Local deployment can therefore require substantial memory and capable accelerators. Microsoft documented testing on NVIDIA A100, A6000, and H100 GPUs, and notes that FlashAttention use depends on compatible hardware and software.

Pricing and deployment options

There is no current Azure Foundry hosted price for Phi-3.5-MoE-instruct because Microsoft retired the endpoint on August 30, 2025. Historical Azure pricing was approximately $0.16 per 1 million input tokens and $0.64 per 1 million output tokens. Those figures should be treated as historical reference only, not as an active quotation or an available billing plan.

Self-hosted use does not have a model-token fee from Microsoft in the supplied research, but it is not cost-free. Operators must account for GPU or cloud compute, storage, bandwidth, serving software, monitoring, and engineering time. Transformers support was added in version 4.46.0, and Microsoft’s repository includes guidance for Transformers, vLLM, and local inference applications.

When to choose Phi-3.5-MoE-instruct

Choose Phi-3.5-MoE-instruct when you need an MIT-licensed, downloadable text model with a very large context window and capabilities spanning multilingual generation, coding, mathematics, and reasoning. It is especially reasonable for teams that can operate their own infrastructure, need control over where inference runs, or want to build retrieval-augmented applications around an open checkpoint.

A different option may be more appropriate in several situations. Choose a current hosted model when you need an active managed endpoint, current pricing, and less infrastructure responsibility. Choose a multimodal model when image, audio, or video input or output is central to the application. Choose a smaller model when available hardware, response latency, or operating cost is more important than a 128K context window and the additional capacity of this checkpoint. For Azure users specifically, Microsoft lists Phi-4-mini-instruct as the suggested replacement for the retired hosted deployment, although that is a separate model and should be evaluated on its own requirements.

Overall, Phi-3.5-MoE-instruct is now primarily a self-managed long-context language model rather than a current Azure service. Its strongest reasons to use it are the open-weight MIT license, broad multilingual and text-generation scope, MoE architecture, and 128K context. Its main compromises are substantial deployment requirements, a pre-October-2023 knowledge base, text-only operation, lack of documented native tool support, and the loss of its Azure-hosted endpoint.


Answers to Frequently Asked Questions

What hardware and deployment costs are required to run Phi-3.5-MoE-instruct locally?
Although only about 6.6 billion parameters are active during an inference step, the complete checkpoint contains approximately 42 billion parameters and can require substantial memory and capable GPUs. Local operators must budget for accelerators or cloud compute, storage, bandwidth, serving software, monitoring, and engineering. There is no current Azure token price because the hosted endpoint was retired; historical Azure pricing should not be treated as an active quotation.
Can Phi-3.5-MoE-instruct process images or use native function calling?
No. Phi-3.5-MoE-instruct is a text-only model and does not natively process images, audio, or video. The supplied documentation also does not verify native tool calling, function calling, or structured JSON output. Developers can build external orchestration and formatting workflows, but these features should not be assumed to be built into the model.
What is the context window of Phi-3.5-MoE-instruct?
Phi-3.5-MoE-instruct supports a context window of up to 131,072 tokens, commonly described as 128K tokens. This makes it suitable for long documents, extensive conversation histories, codebases, and retrieval-augmented prompts, although the practical limit depends on the inference framework, memory, prompt format, and hardware.
What is Phi-3.5-MoE-instruct?
Phi-3.5-MoE-instruct is Microsoft’s instruction-tuned, open-weight language model for text generation. It uses a mixture-of-experts architecture with approximately 42 billion total parameters and about 6.6 billion active parameters when two experts are selected. It supports tasks such as summarization, coding, reasoning, mathematics, question answering, and multilingual text generation.
Does Phi-3.5-MoE-instruct still have an Azure hosted endpoint?
No. Microsoft retired the Azure Foundry endpoint for Phi-3.5-MoE-instruct on August 30, 2025. The model weights remain available through Microsoft’s Hugging Face repository and can be deployed with tools such as Transformers and vLLM. Microsoft identified Phi-4-mini-instruct as the suggested replacement for the retired hosted offering.


Sources 5
Provider

About Microsoft Copilot