Phi-3

Phi-3-small-128k-instruct

by Microsoft Copilot · Retired from Microsoft Foundry on 2025-08-30; open-weight checkpoint remains available for self-hosted deployment

Microsoft Phi-3-small-128k-instruct is a 7B open-weight, text-only model with a 131,072-token context window and MIT licensing. It supports instruction following, document analysis, coding, mathematics, and local deployment. Its Microsoft Foundry hosting was retired on August 30, 2025, while the Hugging Face checkpoint remains available for self-hosted use.

Text Reasoning Coding
Microsoft Phi-3-small-128k-instruct is a compact, open-weight language model designed to handle unusually large text inputs without requiring a frontier-scale model. It accepts text and generates text, supports a 131,072-token context window, and was released under the MIT license for uses including summarization, document analysis, question answering, coding assistance, and local AI applications. Microsoft Foundry historically offered hosted access, but that deployment was deprecated on June 30, 2025 and retired on August 30, 2025. The model remains relevant for self-hosted workloads through its Microsoft Hugging Face checkpoint.
Outputs

What Phi-3-small-128k-instruct can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Fine-tuning
Model profile

Performance characteristics

6/10 Reasoning
6/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Phi-3
Model type Lightweight
Context window 131K tokens
Maximum output 4K tokens
Knowledge cutoff October 2023
Release date 2024-05-21
Status Retired from Microsoft Foundry on 2025-08-30; open-weight checkpoint remains available for self-hosted deployment
Deprecation date 2025-06-30
Shutdown date 2025-08-30
Knowledge cutoff notes

Microsoft's model card explicitly identifies October 2023 as the offline training-data cutoff. This is separate from the model's May 21, 2024 release date and is not changed by retrieval, external context, or hosted tooling.

Model notes

Canonical Hugging Face identifier is microsoft/Phi-3-small-128k-instruct. The model has approximately 7B parameters, uses a dense decoder-only Transformer with alternating dense and block-sparse attention, and was trained on 4.8T tokens. Microsoft documents an October 2023 knowledge cutoff and an MIT license. The model card supports deployment with Transformers, vLLM, SGLang, ONNX Runtime, and other local inference tools. Microsoft Foundry marked the model legacy on June 9, 2025, deprecated it on June 30, 2025, and retired it on August 30, 2025, with Phi-4-mini-instruct listed as the suggested replacement. Editorial scores are comparative estimates, not vendor specifications.

Cost

Model pricing

Input USD 0.00015 per 1,000 input tokens historically on Azure Foundry; hosted Azure access is retired
Output USD 0.0006 per 1,000 output tokens historically on Azure Foundry; hosted Azure access is retired
Model guide

Phi-3-small-128k-instruct: Microsoft’s Open-Weight Long-Context Model

Microsoft Phi-3-small-128k-instruct is a 7-billion-parameter, text-only open-weight model built for instruction following, coding, mathematics, and long-context document workloads. Its 131,072-token context window was its defining feature, while its hosted Microsoft Foundry deployment has since been retired.

What is Phi-3-small-128k-instruct?

Phi-3-small-128k-instruct is a member of Microsoft’s Phi-3 family of small language models. It contains approximately 7 billion parameters and was released on May 21, 2024. The “128K” in its name refers to its maximum context length: the amount of text the model can consider in one request. Microsoft documents this limit as 131,072 tokens.

In practical terms, the long context makes the model suitable for working with large documents, extended conversations, source-code repositories, or multiple related files when those inputs fit within the model’s context window. It is an instruction-tuned model, meaning it was adapted to follow natural-language requests rather than merely predict unstructured text.

The model is text-only. It accepts text input and produces text output; it does not natively process images, audio, or video, and it does not generate those media types.

Model profile and verified specifications

SpecificationDetail
ProviderMicrosoft
Model familyPhi-3
ParametersApproximately 7 billion
ArchitectureDense decoder-only Transformer with alternating dense and block-sparse attention
Release dateMay 21, 2024
Context length131,072 tokens
Hosted maximum output4,096 tokens in Microsoft Foundry
Knowledge cutoffOctober 2023
LicenseMIT
Input and outputText input and text output

The context limit and output limit should not be confused. A 128K context does not mean that the model can generate a 128K-token answer. Microsoft Foundry documented a maximum output of 4,096 tokens, while the available output limit may depend on the inference software and deployment configuration used for self-hosting.

What the model is designed to do

Phi-3-small-128k-instruct was built for general-purpose text generation with emphasis on instruction following, common-sense and logical reasoning, mathematics, code generation, and long-context processing. Microsoft positioned the Phi-3 family as a collection of smaller models intended to provide useful capabilities with more manageable deployment requirements than much larger language models.

Its long-context capability is particularly relevant when the task depends on relationships across a large body of text. Examples include summarizing a lengthy report, answering questions about a collection of technical notes, reviewing a substantial source file, or extracting information from a long policy document. The model can also serve as a local assistant when an organization prefers to run an open-weight checkpoint rather than send prompts to a hosted service.

Microsoft reported favorable benchmark comparisons for the model at launch, but those claims should be interpreted in their original evaluation context. Benchmark results do not guarantee the same performance on a particular application, especially when prompts contain specialized terminology, ambiguous instructions, or information outside the model’s training cutoff.

Reasoning, coding, and long-context performance

The model can perform ordinary reasoning tasks such as following multi-step instructions, organizing information, answering questions, and carrying out mathematical transformations. Its reasoning capability should be understood as language-model reasoning rather than a guaranteed formal proof system: it can produce useful intermediate explanations but may still make arithmetic, logical, or factual mistakes.

For coding, Phi-3-small-128k-instruct can generate code, explain snippets, help transform existing code, and support programming questions. Its 128K context window can be useful when a task requires supplying a large module or several connected files. However, the model has no native tool-calling support in Microsoft Foundry, so it cannot independently invoke a compiler, execute a program, browse the web, or call an external function as part of the hosted model interface described in the supplied specifications.

The long context is an important practical advantage, but a large input window does not ensure that every detail will be used correctly. For production document analysis, outputs should be checked against the source material, particularly when the answer affects legal, financial, medical, security, or other high-impact decisions.

Deployment and current availability

The canonical open-weight checkpoint is microsoft/Phi-3-small-128k-instruct on Hugging Face. Microsoft’s model card documents deployment with Transformers, vLLM, SGLang, ONNX Runtime, and other compatible local inference tools. This gives developers several routes for running the model on their own infrastructure, subject to the hardware, quantization, memory, and serving configuration they select.

Microsoft Foundry was previously a hosted option. Its catalog documentation listed text input up to 131,072 tokens, text output up to 4,096 tokens, no tool calling, and text-only response formats. Microsoft later marked the model as legacy, deprecated it on June 30, 2025, and retired it from Foundry on August 30, 2025. Therefore, historical Azure API pricing should not be treated as a currently available hosted offer.

Microsoft listed Phi-4-mini-instruct as the suggested replacement for Azure users. That recommendation is a migration reference rather than evidence that the two models have identical behavior, limits, or pricing.

Historical pricing and cost trade-offs

Microsoft’s historical Azure pricing listed Phi-3-small-128k-instruct at USD 0.00015 per 1,000 input tokens and USD 0.0006 per 1,000 output tokens. These were hosted token rates, but the associated Foundry deployment has been retired, so they are historical rather than a current purchase option.

Self-hosted users do not pay Microsoft a per-token API charge for using the downloaded checkpoint. Instead, they pay for compute, storage, networking, serving software, maintenance, and operational support. The model’s approximately 7-billion-parameter size can make it more practical to deploy than much larger models, but the actual cost and speed depend on hardware and runtime choices. The supplied research does not establish a universal hardware requirement or a guaranteed latency figure.

Editorially, this model is best viewed as a cost-conscious option for text workloads rather than a replacement for every larger model. Its compact scale and open-weight availability favor controlled local deployment and high-volume processing, while larger or newer models may be preferable when the task demands stronger reliability, broader current knowledge, tool integration, or more advanced reasoning.

Modalities, structured output, and tools

  • Text input: supported.
  • Text output: supported.
  • Image, audio, and video input: not supported natively.
  • Image, audio, video, or music generation: not supported.
  • Tool or function calling: Microsoft Foundry listed it as unsupported.
  • Web search: not built into the model.
  • Streaming: supported in the supplied model record, depending on the serving implementation.
  • Fine-tuning: supported in the supplied model record, with the exact method depending on the deployment stack.

The model record does not verify a distinct native JSON mode or guaranteed structured-output feature. Developers can potentially constrain or parse generated text using external tooling, but that should not be presented as a provider-guaranteed structured-output capability.

Limitations to consider before deployment

The most important limitation is that Phi-3-small-128k-instruct is not a current-information system. Its documented knowledge cutoff is October 2023. It cannot inherently know events after that date unless later information is supplied in the prompt or connected through an external retrieval system. Adding retrieval can improve freshness, but retrieval itself is not a native capability of the model.

The model is also text-only and lacks native tool calling. Applications that need image understanding, speech, video processing, live web access, database actions, or reliable function execution should consider a model or system designed for those requirements.

As with other language models, it can produce incorrect facts, unsafe content, biased responses, or inconsistent reasoning. The MIT license provides broad rights to use and modify the checkpoint, but it does not remove the need for application-level testing, content safeguards, privacy review, and human oversight.

Its long context can also increase processing cost and latency when very large prompts are supplied. A smaller prompt containing only the relevant passages may be faster and more reliable than sending an entire document collection, even when the full collection fits within the stated limit.

When to choose this model

Choose Phi-3-small-128k-instruct when you need an open-weight, MIT-licensed text model with a very large context window and are prepared to manage deployment yourself. It is a reasonable candidate for:

  • long-document summarization and question answering;
  • local or private text assistants;
  • code explanation and code-generation support;
  • mathematics and general instruction-following tasks;
  • batch processing where a smaller model can reduce infrastructure demands;
  • research and commercial applications that benefit from an openly available checkpoint.

Another option may be more appropriate when you need native multimodal input, web search, current knowledge, tool calling, guaranteed structured responses, or stronger frontier-level accuracy. Azure users seeking a successor should evaluate Microsoft’s suggested Phi-4-mini-instruct, while recognizing that migration requires checking its current limits, availability, performance, and pricing separately.

Overall, Phi-3-small-128k-instruct’s clearest distinction is the combination of a small, open-weight model and a 131,072-token context window. It remains useful for self-hosted long-context text applications, but its retired hosted service, October 2023 knowledge cutoff, text-only design, and lack of native tools should guide deployment decisions.


Answers to Frequently Asked Questions

What are the main limitations of Phi-3-small-128k-instruct?
Its knowledge cutoff is October 2023, so it does not inherently know later events. It may also produce factual, mathematical, or logical errors, lacks native multimodal and tool-use capabilities, and can become slower or more expensive to run with very large prompts. Outputs should be verified for high-impact applications.
Is Phi-3-small-128k-instruct still available through Microsoft Foundry?
No. Microsoft marked Phi-3-small-128k-instruct as legacy, deprecated it on June 30, 2025, and retired it from Foundry on August 30, 2025. The open-weight checkpoint remains available for self-hosting through compatible tools such as Transformers, vLLM, SGLang, and ONNX Runtime.
Can Phi-3-small-128k-instruct process images, audio, or use tools?
No. The model is text-only and does not natively process images, audio, or video. It also does not have built-in web search or tool and function calling in the documented Microsoft Foundry interface.
How many tokens can Phi-3-small-128k-instruct process?
Phi-3-small-128k-instruct has a maximum context length of 131,072 tokens. This limit covers the input and conversation context the model can consider, not the length of the answer it can generate. Microsoft Foundry documented a maximum output of 4,096 tokens.
What is Phi-3-small-128k-instruct?
Phi-3-small-128k-instruct is Microsoft’s approximately 7-billion-parameter, instruction-tuned language model for text generation, reasoning, coding, mathematics, and long-context processing. It supports up to 131,072 tokens of context and is available as an open-weight checkpoint under the MIT license.


Sources 5
Provider

About Microsoft Copilot