What is Phi-3-small-128k-instruct?
Phi-3-small-128k-instruct is a member of Microsoft’s Phi-3 family of small language models. It contains approximately 7 billion parameters and was released on May 21, 2024. The “128K” in its name refers to its maximum context length: the amount of text the model can consider in one request. Microsoft documents this limit as 131,072 tokens.
In practical terms, the long context makes the model suitable for working with large documents, extended conversations, source-code repositories, or multiple related files when those inputs fit within the model’s context window. It is an instruction-tuned model, meaning it was adapted to follow natural-language requests rather than merely predict unstructured text.
The model is text-only. It accepts text input and produces text output; it does not natively process images, audio, or video, and it does not generate those media types.
Model profile and verified specifications
| Specification | Detail |
|---|---|
| Provider | Microsoft |
| Model family | Phi-3 |
| Parameters | Approximately 7 billion |
| Architecture | Dense decoder-only Transformer with alternating dense and block-sparse attention |
| Release date | May 21, 2024 |
| Context length | 131,072 tokens |
| Hosted maximum output | 4,096 tokens in Microsoft Foundry |
| Knowledge cutoff | October 2023 |
| License | MIT |
| Input and output | Text input and text output |
The context limit and output limit should not be confused. A 128K context does not mean that the model can generate a 128K-token answer. Microsoft Foundry documented a maximum output of 4,096 tokens, while the available output limit may depend on the inference software and deployment configuration used for self-hosting.
What the model is designed to do
Phi-3-small-128k-instruct was built for general-purpose text generation with emphasis on instruction following, common-sense and logical reasoning, mathematics, code generation, and long-context processing. Microsoft positioned the Phi-3 family as a collection of smaller models intended to provide useful capabilities with more manageable deployment requirements than much larger language models.
Its long-context capability is particularly relevant when the task depends on relationships across a large body of text. Examples include summarizing a lengthy report, answering questions about a collection of technical notes, reviewing a substantial source file, or extracting information from a long policy document. The model can also serve as a local assistant when an organization prefers to run an open-weight checkpoint rather than send prompts to a hosted service.
Microsoft reported favorable benchmark comparisons for the model at launch, but those claims should be interpreted in their original evaluation context. Benchmark results do not guarantee the same performance on a particular application, especially when prompts contain specialized terminology, ambiguous instructions, or information outside the model’s training cutoff.
Reasoning, coding, and long-context performance
The model can perform ordinary reasoning tasks such as following multi-step instructions, organizing information, answering questions, and carrying out mathematical transformations. Its reasoning capability should be understood as language-model reasoning rather than a guaranteed formal proof system: it can produce useful intermediate explanations but may still make arithmetic, logical, or factual mistakes.
For coding, Phi-3-small-128k-instruct can generate code, explain snippets, help transform existing code, and support programming questions. Its 128K context window can be useful when a task requires supplying a large module or several connected files. However, the model has no native tool-calling support in Microsoft Foundry, so it cannot independently invoke a compiler, execute a program, browse the web, or call an external function as part of the hosted model interface described in the supplied specifications.
The long context is an important practical advantage, but a large input window does not ensure that every detail will be used correctly. For production document analysis, outputs should be checked against the source material, particularly when the answer affects legal, financial, medical, security, or other high-impact decisions.
Deployment and current availability
The canonical open-weight checkpoint is microsoft/Phi-3-small-128k-instruct on Hugging Face. Microsoft’s model card documents deployment with Transformers, vLLM, SGLang, ONNX Runtime, and other compatible local inference tools. This gives developers several routes for running the model on their own infrastructure, subject to the hardware, quantization, memory, and serving configuration they select.
Microsoft Foundry was previously a hosted option. Its catalog documentation listed text input up to 131,072 tokens, text output up to 4,096 tokens, no tool calling, and text-only response formats. Microsoft later marked the model as legacy, deprecated it on June 30, 2025, and retired it from Foundry on August 30, 2025. Therefore, historical Azure API pricing should not be treated as a currently available hosted offer.
Microsoft listed Phi-4-mini-instruct as the suggested replacement for Azure users. That recommendation is a migration reference rather than evidence that the two models have identical behavior, limits, or pricing.
Historical pricing and cost trade-offs
Microsoft’s historical Azure pricing listed Phi-3-small-128k-instruct at USD 0.00015 per 1,000 input tokens and USD 0.0006 per 1,000 output tokens. These were hosted token rates, but the associated Foundry deployment has been retired, so they are historical rather than a current purchase option.
Self-hosted users do not pay Microsoft a per-token API charge for using the downloaded checkpoint. Instead, they pay for compute, storage, networking, serving software, maintenance, and operational support. The model’s approximately 7-billion-parameter size can make it more practical to deploy than much larger models, but the actual cost and speed depend on hardware and runtime choices. The supplied research does not establish a universal hardware requirement or a guaranteed latency figure.
Editorially, this model is best viewed as a cost-conscious option for text workloads rather than a replacement for every larger model. Its compact scale and open-weight availability favor controlled local deployment and high-volume processing, while larger or newer models may be preferable when the task demands stronger reliability, broader current knowledge, tool integration, or more advanced reasoning.
Modalities, structured output, and tools
- Text input: supported.
- Text output: supported.
- Image, audio, and video input: not supported natively.
- Image, audio, video, or music generation: not supported.
- Tool or function calling: Microsoft Foundry listed it as unsupported.
- Web search: not built into the model.
- Streaming: supported in the supplied model record, depending on the serving implementation.
- Fine-tuning: supported in the supplied model record, with the exact method depending on the deployment stack.
The model record does not verify a distinct native JSON mode or guaranteed structured-output feature. Developers can potentially constrain or parse generated text using external tooling, but that should not be presented as a provider-guaranteed structured-output capability.
Limitations to consider before deployment
The most important limitation is that Phi-3-small-128k-instruct is not a current-information system. Its documented knowledge cutoff is October 2023. It cannot inherently know events after that date unless later information is supplied in the prompt or connected through an external retrieval system. Adding retrieval can improve freshness, but retrieval itself is not a native capability of the model.
The model is also text-only and lacks native tool calling. Applications that need image understanding, speech, video processing, live web access, database actions, or reliable function execution should consider a model or system designed for those requirements.
As with other language models, it can produce incorrect facts, unsafe content, biased responses, or inconsistent reasoning. The MIT license provides broad rights to use and modify the checkpoint, but it does not remove the need for application-level testing, content safeguards, privacy review, and human oversight.
Its long context can also increase processing cost and latency when very large prompts are supplied. A smaller prompt containing only the relevant passages may be faster and more reliable than sending an entire document collection, even when the full collection fits within the stated limit.
When to choose this model
Choose Phi-3-small-128k-instruct when you need an open-weight, MIT-licensed text model with a very large context window and are prepared to manage deployment yourself. It is a reasonable candidate for:
- long-document summarization and question answering;
- local or private text assistants;
- code explanation and code-generation support;
- mathematics and general instruction-following tasks;
- batch processing where a smaller model can reduce infrastructure demands;
- research and commercial applications that benefit from an openly available checkpoint.
Another option may be more appropriate when you need native multimodal input, web search, current knowledge, tool calling, guaranteed structured responses, or stronger frontier-level accuracy. Azure users seeking a successor should evaluate Microsoft’s suggested Phi-4-mini-instruct, while recognizing that migration requires checking its current limits, availability, performance, and pricing separately.
Overall, Phi-3-small-128k-instruct’s clearest distinction is the combination of a small, open-weight model and a 131,072-token context window. It remains useful for self-hosted long-context text applications, but its retired hosted service, October 2023 knowledge cutoff, text-only design, and lack of native tools should guide deployment decisions.

