What is Phi-3.5-MoE-instruct?
Phi-3.5-MoE-instruct is an instruction-tuned language model from Microsoft’s Phi-3.5 family. It accepts text prompts and generates text responses, so it is designed for tasks such as answering questions, following written instructions, summarizing documents, explaining code, solving mathematical problems, and supporting conversational applications.
The model uses a mixture-of-experts, or MoE, architecture. Instead of activating every parameter for every token, the model routes each part of a request through a subset of specialized expert networks. Phi-3.5-MoE-instruct contains approximately 42 billion total parameters distributed across 16 experts, while approximately 6.6 billion parameters are active when two experts are selected for an inference step. This reduces active computation compared with running a dense model containing all 42 billion parameters, although the complete checkpoint still has substantial memory requirements.
Microsoft released the model in August 2024 under the MIT license. The open-weight license and downloadable checkpoint make it relevant to organizations that want more control over deployment, infrastructure, and data handling than a closed hosted model normally provides.
Availability and current status
Phi-3.5-MoE-instruct is no longer an active Azure Foundry hosted model. Microsoft retired the Azure model endpoint on August 30, 2025, and identified Phi-4-mini-instruct as the suggested replacement for that hosted offering. This retirement concerns the Azure service rather than the model weights themselves.
The checkpoint remains available through Microsoft’s Hugging Face repository and can be used with compatible tools such as Transformers and vLLM. As a result, the model is best understood today as an open-weight model for self-managed or third-party-supported inference rather than as a current Microsoft-hosted API choice. Users should distinguish the historical Azure pricing and limits from the model’s continuing availability as downloadable software.
Technical specifications
| Specification | Details |
|---|---|
| Provider | Microsoft |
| Model family | Phi-3.5 |
| Release | August 2024 |
| Architecture | Decoder-only Transformer mixture of experts |
| Total parameters | Approximately 42 billion |
| Active parameters | Approximately 6.6 billion when two experts are active |
| Context window | 131,072 tokens, commonly described as 128K tokens |
| Maximum Azure output | 4,096 tokens |
| Input | Text only |
| Output | Generated text only |
| License | MIT |
| Knowledge cutoff | October 2023 for publicly available data, according to the model card |
The 128K-token context window is the model’s main architectural advantage for long prompts. It can provide room for large documents, conversation histories, codebases, or retrieved reference material in a single request. The practical limit still depends on the inference framework, available memory, prompt formatting, and the deployment environment.
Capabilities and best use cases
Microsoft designed Phi-3.5-MoE-instruct for instruction following, reasoning, mathematics, coding, multilingual generation, and general-purpose assistant tasks. Its 128K context makes it particularly relevant when a request depends on more source material than a smaller context window can comfortably contain.
- Long-context document work: Summarization, question answering, extraction, and comparison across lengthy text collections.
- Multilingual applications: The model supports Arabic, Chinese, Czech, Danish, Dutch, English, Finnish, French, German, Hebrew, Hungarian, Italian, Japanese, Korean, Norwegian, Polish, Portuguese, Russian, Spanish, Swedish, Thai, Turkish, and Ukrainian.
- Code assistance: Code explanation, generation, transformation, and debugging-oriented conversations.
- Mathematics and reasoning: Step-by-step problem solving and instruction-based analytical tasks, with the usual need to verify important results.
- Retrieval-augmented generation: Combining the model with a search or retrieval system so that responses can use an organization’s documents or more current information.
- Controlled local applications: Self-hosted assistants and text-generation systems where deployment control or an open license is important.
The model card and supplied research support these use cases as intended capabilities, but they do not guarantee a particular accuracy level for every domain. A model’s performance can vary with prompt design, language, document quality, retrieval quality, and the hardware and inference software used to run it.
Modalities, tools, and output control
Phi-3.5-MoE-instruct is text-only. It does not natively accept images, audio, or video, and it does not generate images, audio, music, or video. Applications can place descriptions or transcriptions into a text prompt, but that is different from native multimodal understanding.
The supplied model data does not identify native tool or function-calling support. It should therefore not be selected when an application requires a model with a documented built-in tool-use interface. Developers can still build application-level workflows around generated text, such as having software parse a response and then perform an action, but that is an external orchestration layer rather than a verified native model feature.
The model produces text and has no verified native structured-output or JSON-mode capability in the supplied research. A deployment can request a particular format through prompting or enforce formatting in application code, but reliable schema enforcement should not be assumed.
Reasoning, coding, speed, and cost trade-offs
Phi-3.5-MoE-instruct was positioned for reasoning, mathematics, and coding, and its MoE design activates fewer parameters than the total checkpoint size. In principle, that gives it a more favorable active-computation profile than a dense model with the same total parameter count. It does not, however, make the model lightweight to host: loading and serving a roughly 42-billion-parameter checkpoint can still require substantial memory and high-end hardware.
Editorial comparative assessments supplied with the model rate reasoning and coding at 7 out of 10, speed at 6 out of 10, and cost at 8 out of 10. These are subjective editorial scores, not Microsoft benchmark results or official performance guarantees. They reflect the model’s intended balance of capability, active computation, and open deployment economics rather than a universal ranking.
For local operators, the trade-off is therefore between capability and infrastructure. Phi-3.5-MoE-instruct offers a long context and a relatively efficient active path, but smaller dense models may be easier to run on limited hardware. A hosted model may also be more convenient when the priority is avoiding GPU procurement, model serving, and operational maintenance. Conversely, self-hosting can be preferable when data control, customization, or predictable access matters more than operational simplicity.
Limitations and risks
The most important functional limitation is its text-only design. It is not a direct choice for visual question answering, speech processing, audio generation, or video workflows. Those applications require additional modality-specific systems or a different model.
The model’s knowledge cutoff is October 2023 for publicly available data, according to its model card. It does not automatically know events, documents, or changes that occurred after that point. Current-information applications should connect it to retrieval or search augmentation and should still verify important answers.
Microsoft’s model documentation warns that the model can produce factual errors. Long context does not eliminate hallucinations: placing more text in a prompt can help ground an answer, but the model may still overlook relevant passages, misunderstand instructions, or generate unsupported conclusions. High-stakes uses should include retrieval checks, output validation, and human review.
Hardware is another practical constraint. Although only approximately 6.6 billion parameters are active in the stated two-expert configuration, the full checkpoint is approximately 42 billion parameters. Local deployment can therefore require substantial memory and capable accelerators. Microsoft documented testing on NVIDIA A100, A6000, and H100 GPUs, and notes that FlashAttention use depends on compatible hardware and software.
Pricing and deployment options
There is no current Azure Foundry hosted price for Phi-3.5-MoE-instruct because Microsoft retired the endpoint on August 30, 2025. Historical Azure pricing was approximately $0.16 per 1 million input tokens and $0.64 per 1 million output tokens. Those figures should be treated as historical reference only, not as an active quotation or an available billing plan.
Self-hosted use does not have a model-token fee from Microsoft in the supplied research, but it is not cost-free. Operators must account for GPU or cloud compute, storage, bandwidth, serving software, monitoring, and engineering time. Transformers support was added in version 4.46.0, and Microsoft’s repository includes guidance for Transformers, vLLM, and local inference applications.
When to choose Phi-3.5-MoE-instruct
Choose Phi-3.5-MoE-instruct when you need an MIT-licensed, downloadable text model with a very large context window and capabilities spanning multilingual generation, coding, mathematics, and reasoning. It is especially reasonable for teams that can operate their own infrastructure, need control over where inference runs, or want to build retrieval-augmented applications around an open checkpoint.
A different option may be more appropriate in several situations. Choose a current hosted model when you need an active managed endpoint, current pricing, and less infrastructure responsibility. Choose a multimodal model when image, audio, or video input or output is central to the application. Choose a smaller model when available hardware, response latency, or operating cost is more important than a 128K context window and the additional capacity of this checkpoint. For Azure users specifically, Microsoft lists Phi-4-mini-instruct as the suggested replacement for the retired hosted deployment, although that is a separate model and should be evaluated on its own requirements.
Overall, Phi-3.5-MoE-instruct is now primarily a self-managed long-context language model rather than a current Azure service. Its strongest reasons to use it are the open-weight MIT license, broad multilingual and text-generation scope, MoE architecture, and 128K context. Its main compromises are substantial deployment requirements, a pre-October-2023 knowledge base, text-only operation, lack of documented native tool support, and the loss of its Azure-hosted endpoint.

