Phi-3

Phi-3-mini-128k-instruct

by Microsoft Copilot · Retired from Microsoft Foundry on August 30, 2025; open-weight checkpoint remains downloadable and usable for self-hosted inference.

A practical overview of Microsoft's 3.8B Phi-3-mini-128k-instruct, covering its 128K context window, text-only design, reasoning and coding profile, historical hosted pricing, local deployment options, strengths, limitations, and Microsoft Foundry retirement.

Text Reasoning Coding
Phi-3-mini-128k-instruct is Microsoft's long-context Mini variant in the Phi-3 family. Released in 2024, it combines supervised fine-tuning and direct preference optimization with a compact 3.8-billion-parameter architecture. The model accepts text and produces text, and its 128K context window makes it suitable for long-document question answering, summarization, code understanding, and other workloads where a small deployable model is preferred.
Outputs

What Phi-3-mini-128k-instruct can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Fine-tuning
Model profile

Performance characteristics

6/10 Reasoning
6/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Phi-3
Model type Lightweight
Context window 131K tokens
Maximum output 4K tokens
Knowledge cutoff October 2023
Release date 2024-04-23
Status Retired from Microsoft Foundry on August 30, 2025; open-weight checkpoint remains downloadable and usable for self-hosted inference.
Deprecation date 2025-08-30
Shutdown date 2025-08-30
Knowledge cutoff notes

Microsoft's model card identifies October 2023 as the cutoff date for the static offline training dataset. The cutoff is separate from the April 2024 initial release date and the June 2024 updated instruction-tuning release.

Model notes

The canonical downloadable checkpoint is microsoft/Phi-3-mini-128k-instruct. Microsoft describes it as a 3.8-billion-parameter dense decoder-only Transformer trained on 4.9 trillion tokens, with an October 2023 knowledge cutoff. The model card states a 128K-token context length, while Microsoft Foundry documentation lists 131,072 input tokens and a 4,096-token output limit for the hosted chat-completion deployment. Microsoft retired the Foundry deployment on August 30, 2025 and recommended Phi-4-mini-instruct as the replacement. This does not remove the open-weight checkpoint from Hugging Face or prevent self-hosted inference. The model is text-only, MIT licensed, and supports local inference through Transformers, vLLM, SGLang, ONNX Runtime, and related tooling. Editorial scores are comparative estimates rather than Microsoft benchmark ratings.

Cost

Model pricing

Input $0.00013 per 1,000 input tokens, historical Azure Models-as-a-Service pricing; hosted offering retired
Output $0.00052 per 1,000 output tokens, historical Azure Models-as-a-Service pricing; hosted offering retired
Model guide

Phi-3-mini-128k-instruct: Microsoft’s Small Model for Long-Context Local AI

Microsoft Phi-3-mini-128k-instruct is a 3.8-billion-parameter, instruction-tuned open-weight language model with a 128,000-token context window. It is designed for efficient local, edge, and cloud deployment, with particular strengths in long-context comprehension, coding, mathematics, reasoning, and latency-sensitive applications.

What is Phi-3-mini-128k-instruct?

Microsoft Phi-3-mini-128k-instruct is a compact, instruction-tuned causal language model in the Phi-3 family. It has approximately 3.8 billion parameters and is distributed as an open-weight model under the MIT license. In practical terms, its weights can be downloaded and run through compatible inference software rather than requiring a Microsoft-hosted endpoint.

The model is designed for text generation and chat-style instruction following. Its defining feature is the 128,000-token context window, which allows a single request to contain substantially more source material than many small language models can handle. This makes it particularly relevant to long documents, transcripts, source-code repositories, and retrieval-augmented applications.

Microsoft released Phi-3-mini-128k-instruct in 2024. It is a separate model from the 4K-context Phi-3 Mini variant, so the two should not be treated as interchangeable: the 128K version is intended for workloads where a much larger prompt is important.

Core specifications at a glance

SpecificationDetails
ProviderMicrosoft
Model familyPhi-3
ParametersApproximately 3.8 billion
Release dateApril 23, 2024
Context length128,000 tokens for the model; 131,072-token input limit documented for the hosted deployment
Maximum hosted output4,096 tokens
Input and outputText input and text output
LicenseMIT
Knowledge cutoffOctober 2023
Hosted statusRetired from Microsoft Foundry on August 30, 2025

The context and output limits require some qualification. The model card describes a 128K-token context capability, while Microsoft Foundry documentation listed 131,072 input tokens and a 4,096-token maximum output for its hosted chat-completion deployment. Local runtimes may expose different generation settings or practical limits, so deployment-specific documentation should be checked before implementation.

Why the 128K context window matters

A context window is the amount of text a model can consider in one request, including instructions, source material, conversation history, and the requested response. With a 128,000-token window, Phi-3-mini-128k-instruct can process long documents or sizeable collections of related material without splitting everything into many small prompts.

Useful examples include asking questions about a lengthy technical manual, summarizing a meeting transcript, reviewing a large policy document, or examining multiple files from a codebase. It can also support retrieval-augmented generation workflows in which relevant passages are placed directly into the prompt before the model generates an answer.

Large context capacity is not the same as uniformly perfect recall. Quality can vary depending on where information appears in a very long prompt, how much irrelevant material is included, and how well the application formats the instructions. Production testing should measure retrieval accuracy, instruction persistence, latency, memory use, and answer quality at the prompt lengths the application expects.

Architecture and training focus

Phi-3-mini-128k-instruct is a dense decoder-only Transformer. Microsoft describes the Phi-3 models as being trained using filtered publicly available data and synthetic, textbook-like data emphasizing reasoning, mathematics, coding, general knowledge, and common-sense tasks. The model contains far fewer parameters than large frontier systems, which is central to its deployment value.

The instruction-tuned version uses supervised fine-tuning and direct preference optimization. These techniques are intended to improve the model's ability to follow requests, maintain multi-turn conversations, produce safer responses, and perform reasoning-oriented tasks. Microsoft reported that the June 2024 update improved structured-output behavior and long-context benchmark performance compared with the initial Phi-3 Mini release.

The documented knowledge cutoff is October 2023. The model therefore does not inherently know events, software changes, or documents created after that point. Current information must be supplied through the prompt or an external retrieval system.

Capabilities, modalities, and tools

This is a text-only model. It accepts text and returns generated text; it does not natively accept images, audio, or video, and it does not generate images, audio, or video. Its text-only design distinguishes it from multimodal systems and from other models in the broader Phi family that may support additional input types.

Phi-3-mini-128k-instruct can generate ordinary prose, summaries, explanations, code, and structured text when prompted appropriately. The supplied specifications do not identify native web search, function calling, tool execution, or action output as supported features. An application can still place the model inside a larger system with retrieval, software tools, or parsers, but those capabilities would come from the surrounding application rather than from a documented native tool-use interface.

Structured responses may be possible through careful prompting or application-side validation, but the supplied model record does not identify a dedicated JSON mode or structured-output guarantee. Systems that require machine-readable output should validate and, where necessary, repair the model's responses.

Reasoning and coding profile

Microsoft positions the Phi-3 family around reasoning, mathematics, coding, general knowledge, and common-sense tasks. Phi-3-mini-128k-instruct is therefore a reasonable candidate for lightweight code explanation, code completion experiments, mathematical problem solving, document classification, and step-by-step analytical assistance.

Its small size creates a clear trade-off. It can be faster and less expensive to run than much larger models, but it generally will not match larger open-weight or frontier models on difficult reasoning, complex software engineering, ambiguous instructions, or tasks requiring broad world knowledge. The editorial assessment for this model rates reasoning and coding at 6 out of 10, while speed is rated 8 out of 10 and cost efficiency 9 out of 10. These are comparative editorial scores, not Microsoft-published benchmark ratings.

For coding use, the model is best suited to constrained tasks such as explaining a function, generating short utilities, transforming code, drafting tests, or searching a supplied code context. Developers should test it carefully before relying on it for large refactors, security-sensitive code, or autonomous software changes.

Deployment options and current availability

The canonical downloadable checkpoint is microsoft/Phi-3-mini-128k-instruct on Hugging Face. It has also been made available in optimized formats and ecosystems associated with Transformers, vLLM, SGLang, ONNX Runtime, and compatible local inference applications. The exact hardware requirement depends on quantization, runtime configuration, batch size, and the amount of context processed.

Microsoft Foundry retired the hosted Phi-3-mini-128k-instruct offering on August 30, 2025 and recommended Phi-4-mini-instruct as a replacement. This retirement applies to the Microsoft-hosted model service, not to the downloadable open-weight checkpoint. Users can still self-host the model or use compatible third-party infrastructure, subject to the availability and support policies of those platforms.

Historical Azure Models-as-a-Service pricing was listed at $0.00013 per 1,000 input tokens and $0.00052 per 1,000 output tokens. These prices should not be treated as current purchasing options because the Microsoft Foundry hosted offering has been retired. Self-hosting instead shifts the cost to hardware, hosting, power, storage, and operational maintenance.

Main strengths and limitations

Strengths

  • Long context in a small model: the 128K context window is the main reason to choose this variant over many compact alternatives.
  • Deployment flexibility: MIT-licensed open weights support local, private, offline, and edge-oriented use cases.
  • Lower resource demands: approximately 3.8 billion parameters can make it more practical than larger models when latency, memory, or operating cost matters.
  • Useful general task coverage: the training and instruction tuning target coding, mathematics, reasoning, summarization, and conversational tasks.
  • Broad tooling compatibility: support for common open-model runtimes provides several options for local or self-managed inference.

Limitations

  • Not a frontier model: its compact size means lower expected capability on demanding reasoning, complex coding, and broad knowledge tasks than larger models.
  • Static knowledge: the October 2023 cutoff makes retrieval or another current-information source necessary for up-to-date answers.
  • Text only: it is not suitable for native image, audio, video, or speech workloads.
  • Long context can be expensive locally: processing the full 128K window can increase memory use and reduce throughput, even when the model itself is relatively small.
  • No active Microsoft-hosted endpoint: new users seeking a managed Microsoft deployment need to evaluate a supported successor or another hosting provider.
  • Multilingual use requires testing: the model is English-focused, so quality should be evaluated before production use in other languages.

When to choose Phi-3-mini-128k-instruct

Choose Phi-3-mini-128k-instruct when a project needs a relatively small text model with unusually large context support and the team can manage self-hosting or compatible infrastructure. Good fits include private document analysis, long-document summarization, codebase comprehension, offline assistants, retrieval-augmented generation, embedded applications, and latency- or cost-sensitive text workflows.

It is especially attractive when keeping data within a private environment matters more than obtaining the highest available reasoning quality. The MIT license and downloadable weights also provide more deployment control than a hosted-only model, although the operator remains responsible for infrastructure, security, monitoring, and output validation.

Another option may be more appropriate when the application needs current factual answers without building retrieval, native multimodal input, reliable function calling, high-end software engineering, or a currently supported Microsoft-managed endpoint. A larger model is generally preferable for difficult reasoning and complex coding; a multimodal model is preferable for image, audio, or video inputs; and a supported successor should be considered for new Microsoft Foundry deployments.

Bottom line

Phi-3-mini-128k-instruct is best understood as a compact, open-weight long-context model rather than a general replacement for larger current systems. Its combination of approximately 3.8 billion parameters, a 128K context window, text-only operation, MIT licensing, and local deployment support makes it useful for private and resource-conscious applications. Its principal trade-offs are lower peak capability, an October 2023 knowledge cutoff, limited native tooling, and the retirement of its Microsoft Foundry hosted offering.


Answers to Frequently Asked Questions

What are the main limitations of Phi-3-mini-128k-instruct?
Its main limitations are lower capability than larger frontier models on difficult reasoning and complex coding, an October 2023 knowledge cutoff, text-only input and output, no documented native web search or function-calling interface, increased local resource use for very long contexts, and weaker suitability for multilingual production workloads without testing.
Is Phi-3-mini-128k-instruct still available through Microsoft Foundry?
No. Microsoft retired the hosted Phi-3-mini-128k-instruct offering from Microsoft Foundry on August 30, 2025, and recommended Phi-4-mini-instruct as a replacement. The downloadable open-weight model remains available for self-hosting and compatible third-party infrastructure.
Can Phi-3-mini-128k-instruct run locally?
Yes. The downloadable open-weight checkpoint, listed as microsoft/Phi-3-mini-128k-instruct on Hugging Face, can be run locally or on self-managed infrastructure with compatible tools such as Transformers, vLLM, SGLang, ONNX Runtime, and other local inference applications. Hardware requirements depend on quantization, context length, batch size, and runtime configuration.
What is Phi-3-mini-128k-instruct?
Phi-3-mini-128k-instruct is a compact, instruction-tuned causal language model from Microsoft’s Phi-3 family. It has approximately 3.8 billion parameters, uses an MIT license, supports text input and output, and is designed to run locally or through compatible self-hosted inference software.
What is the context window of Phi-3-mini-128k-instruct?
Phi-3-mini-128k-instruct supports a context window of up to 128,000 tokens. This allows it to process long documents, transcripts, codebases, and retrieval-augmented prompts in a single request, although memory use, latency, and answer quality can vary at very large prompt sizes.


Sources 6
Provider

About Microsoft Copilot