Phi-3

Phi-3-medium-128k-instruct

by Microsoft Copilot · Generally available in Microsoft Foundry; open-weight checkpoint available for download and self-hosting

Microsoft Phi-3-medium-128k-instruct is a 14B MIT-licensed text model with a 131,072-token context window and a 4,096-token Foundry output limit. It targets long documents, RAG, coding, mathematics, reasoning, and self-hosted or Azure inference, but lacks native multimodal input and verified tool calling.

Text Reasoning Coding
Phi-3-medium-128k-instruct is Microsoft's long-context Phi-3 Medium model for instruction-following text applications. It accepts text and returns text, supports a 131,072-token context window, and has a documented maximum output of 4,096 tokens in Microsoft Foundry. The open-weight checkpoint can be self-hosted under the MIT license or accessed through Microsoft's model catalog.
Outputs

What Phi-3-medium-128k-instruct can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Fine-tuning
Model profile

Performance characteristics

7/10 Reasoning
7/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Phi-3
Model type General Purpose
Context window 131K tokens
Maximum output 4K tokens
Knowledge cutoff October 2023
Release date 2024-05-21
Status Generally available in Microsoft Foundry; open-weight checkpoint available for download and self-hosting
Knowledge cutoff notes

Microsoft's model card states that the static model was trained on an offline dataset with a cutoff date of October 2023. Web search or retrieval integrations do not change the underlying cutoff.

Model notes

This is the 128K-context variant of Phi-3 Medium, distinct from Phi-3-medium-4k-instruct. Microsoft describes it as a 14B dense decoder-only Transformer trained on 4.8 trillion tokens, with supervised fine-tuning and direct preference optimization. The model's documented knowledge cutoff is October 2023. The checkpoint is MIT licensed and can be deployed with Transformers, vLLM, SGLang, ONNX Runtime, DirectML, and compatible quantized runtimes. Microsoft Foundry lists a 131,072-token context window and a 4,096-token output limit. Microsoft announced fine-tuning support for Phi-3-medium models in Azure, but structured output, prompt caching, batch processing, and legacy JSON-mode support are not separately verified for this exact model.

Cost

Model pricing

Input $0.17 per 1 million input tokens in Microsoft Foundry
Output $0.68 per 1 million output tokens in Microsoft Foundry
Model guide

Microsoft Phi-3 Medium 128K: A Long-Context Open Model for Local and Azure Use

Microsoft Phi-3-medium-128k-instruct is a 14-billion-parameter, MIT-licensed instruction model built for long-context text generation, coding, mathematics, reasoning, document processing, and retrieval-augmented generation. Its 131,072-token context window makes it suited to large documents and extended prompts, while open weights support local, edge, cloud, and Azure deployments.

What Phi-3 Medium 128K is

Microsoft Phi-3-medium-128k-instruct is a 14-billion-parameter instruction-tuned language model in the Phi-3 family. In practical terms, it is a text model that can follow written instructions, produce answers, summarize documents, generate and explain code, solve mathematical problems, classify information, and extract structured details from text.

The model is a dense decoder-only Transformer. “Dense” means that the model uses its full set of parameters for each generated response, rather than selecting only a subset as some mixture-of-experts systems do. The “instruct” designation indicates that it was adapted for conversational and instruction-following tasks rather than being offered only as a raw next-token prediction model.

Microsoft released the model weights on May 21, 2024, under the permissive MIT license. It is available through Microsoft's Hugging Face repository and is also listed in Azure AI Foundry as a generally available chat-completion model. This gives users two substantially different ways to use it: deploy the checkpoint themselves or consume a managed hosted version.

Why the 128K context matters

The defining feature of this version is its 131,072-token context window. A context window is the amount of input and conversation history the model can consider during one request. The 128K label is a rounded product name; the documented window is 131,072 tokens.

This capacity is useful when a task depends on more information than a typical short chat prompt can contain. Examples include reviewing a long technical document, combining many retrieved passages in a retrieval-augmented generation system, comparing several policy files, preserving a lengthy conversation, or providing extensive few-shot examples that demonstrate the expected output format.

A large context window does not guarantee perfect recall or equally strong reasoning over every part of a very long prompt. It also increases the memory and latency requirements of inference, particularly when the model is run locally. Developers should test whether the model actually uses distant information reliably for their workload instead of treating the context limit as a quality guarantee.

Capabilities and supported modalities

Phi-3-medium-128k-instruct is text-only. It accepts text input and produces text output. It does not natively process images, audio, or video, and it does not generate those media types. It is therefore appropriate for text workflows but not for visual question answering, speech processing, image creation, video generation, or multimodal document analysis unless a separate preprocessing system converts those inputs into text.

Its intended uses include:

  • Long-document summarization and question answering
  • Retrieval-augmented generation over large collections of retrieved text
  • Conversational applications and text assistants
  • Code generation, explanation, transformation, and debugging assistance
  • Mathematical reasoning and worked explanations
  • Text classification, extraction, and information organization
  • Few-shot prompting with many examples

The model has no separately verified native web-search or browsing capability. Its knowledge cutoff is October 2023, so it should not be treated as a source of current information without retrieval, external tools, or a maintained application layer.

Technical specifications and limits

SpecificationDocumented value
ProviderMicrosoft
Model familyPhi-3
Parameters14 billion
ArchitectureDense decoder-only Transformer
Context window131,072 tokens
Maximum output4,096 tokens in Microsoft Foundry
InputText
OutputText
LicenseMIT
Knowledge cutoffOctober 2023
Release dateMay 21, 2024

The 4,096-token output limit is specifically documented for Microsoft Foundry. A self-hosted deployment may expose different runtime controls, but the supplied research does not establish a universal output limit for every inference engine. The context window and maximum output should therefore be checked against the serving platform being used.

Reasoning, coding, and tool support

Microsoft positions Phi-3-medium-128k-instruct for reasoning, mathematics, coding, and general language understanding. Its instruction tuning makes it suitable for explaining intermediate steps, transforming code, answering technical questions, and working through text-based problems. These capabilities are practical strengths of the model's intended design, not a guarantee that every answer will be correct.

The available evaluation data gives the model an editorial reasoning score of 7 out of 10 and an editorial coding score of 7 out of 10. These scores are assessments in the supplied catalog data, not Microsoft-published benchmark results and should not be confused with standardized test measurements. They suggest a capable general-purpose small-to-mid-sized model, while newer or larger frontier models may be preferable for especially difficult reasoning, complex software engineering, or tasks requiring highly reliable multi-step planning.

Native tool or function calling is not separately verified for this exact model in the supplied research. The catalog records tool use as unsupported and structured output as unverified. Developers can still build an application around the model that parses responses or supplies retrieved text, but that is different from having a provider-guaranteed function-calling or schema-constrained output mode.

Deployment options

The open checkpoint can be downloaded and deployed with tools such as Transformers, vLLM, SGLang, ONNX Runtime, DirectML, and compatible quantized runtimes. This flexibility is important for organizations that need control over hosting, data movement, latency, or operating cost.

Local deployment does require suitable hardware and operational work. A 14-billion-parameter model with a 131,072-token context can consume substantial memory, and long prompts may increase latency even when the model itself is smaller than many frontier systems. Quantization and optimized runtimes may reduce resource requirements, but the research does not specify one universally suitable hardware configuration.

Azure AI Foundry offers a managed alternative. Hosted inference avoids the need to operate the model server, but introduces per-token charges and dependence on the availability, limits, and configuration of the Microsoft service. The model is also recorded as supporting fine-tuning in the catalog data; Microsoft has announced fine-tuning support for Phi-3-medium models in Azure.

Pricing

Microsoft Foundry pricing is listed at approximately $0.17 per 1 million input tokens and $0.68 per 1 million output tokens. These are usage-based inference prices rather than a recurring subscription fee. Input and output tokens are charged at different rates, so applications that generate long responses may cost more than applications that mostly classify or summarize text.

Self-hosting does not create provider token charges, but it is not free in an operational sense. Users must provide compute, storage, monitoring, maintenance, and capacity. For steady workloads with suitable infrastructure, self-hosting may offer predictable economics or greater control. For irregular workloads, managed Foundry access may be simpler because it avoids maintaining an always-available inference environment.

Prices and service conditions can change. The listed amounts should be checked in the relevant Microsoft Foundry catalog before deployment, especially when building a cost estimate for production traffic.

Main strengths and trade-offs

The strongest reason to choose this model is the combination of a large context window, open weights, and a relatively compact 14-billion-parameter size. A team can use the same model family for long text prompts, local experiments, private deployments, and managed Azure inference. The MIT license is also more permissive than licenses that restrict commercial use or require access through a hosted API.

Its long context can reduce the need to aggressively split documents or discard retrieved passages. Its text-only design keeps the serving problem narrower than a multimodal system, which can be an advantage when the application already works with text. The model is also intended for coding, mathematics, and reasoning rather than only short conversational replies.

The trade-offs are equally important. Phi-3-medium-128k-instruct is a 2024-generation model with an October 2023 knowledge cutoff. It has no intrinsic web search, and its 4,096-token Foundry output limit may be restrictive for tasks requiring long generated reports or large code outputs. The 14B size can be easier to deploy than a much larger model, but the 128K context still creates significant memory and latency demands.

There is also no separately verified native tool calling, structured output mode, prompt caching, or batch API for this exact model. Applications that depend on those features should verify that the chosen serving platform supplies them or select a model and provider with explicit support.

When to choose Phi-3-medium-128k-instruct

Choose this model when a text application needs unusually long prompts and benefits from the option to self-host. It is a strong candidate for document-heavy RAG, long technical or legal text review, internal knowledge assistants, code-oriented utilities, and experiments where an MIT-licensed open-weight model is important.

It is especially reasonable when the workload values control and deployment flexibility more than access to the newest frontier-model capabilities. A team can test locally, optimize the runtime, and move to Azure Foundry when managed serving is more convenient.

Another option may be more appropriate when the application requires image, audio, or video understanding; current information without an external retrieval layer; guaranteed function calling; highly reliable structured responses; very large generated outputs; or frontier-level performance on difficult reasoning and software-engineering tasks. For short prompts where context size is unimportant, a smaller or lower-cost model may also provide a better speed and cost trade-off.

Practical limitations to plan for

  • Validate important answers, calculations, and generated code; the model can still produce inaccurate or unsafe content.
  • Test long-context retrieval directly, including information placed near the beginning, middle, and end of a prompt.
  • Budget for higher memory use and latency when using very large contexts locally.
  • Do not assume that the MIT license removes all deployment obligations; application, data, and safety requirements still apply.
  • Check whether the selected runtime supports the desired streaming, fine-tuning, quantization, and output-control features.
  • For current facts, connect the model to a retrieval or search system rather than relying on its October 2023 training cutoff.

Overall, Phi-3-medium-128k-instruct is best understood as an open, text-only long-context model rather than a complete agent platform. Its appeal lies in combining 128K-class context capacity with local and managed deployment choices. Its limitations—especially the dated knowledge cutoff, lack of verified native tool support, text-only interface, and output limit in Foundry—should guide the decision about whether it fits a particular production system.


Answers to Frequently Asked Questions

How much does Phi-3 Medium 128K cost on Microsoft Foundry?
Microsoft Foundry pricing is listed at approximately $0.17 per 1 million input tokens and $0.68 per 1 million output tokens. Prices and service conditions can change, so current rates should be confirmed in the Microsoft Foundry catalog before deployment.
What are the main limitations of Phi-3-medium-128k-instruct?
The model is text-only, has an October 2023 knowledge cutoff, and has no separately verified native web search, function calling, or structured-output mode. Microsoft Foundry documents a maximum output of 4,096 tokens, while long contexts can require substantial memory and increase inference latency.
Can Phi-3 Medium 128K be run locally or through Azure?
Yes. The open model weights can be deployed locally using tools such as Transformers, vLLM, SGLang, ONNX Runtime, DirectML, and compatible quantized runtimes. Microsoft also offers managed inference through Azure AI Foundry, which avoids server maintenance but charges for token usage.
What does the 128K context window mean for Phi-3 Medium?
The model supports a documented context window of 131,072 tokens, allowing it to process long documents, extensive conversation history, multiple retrieved passages, or many few-shot examples in one request. A large context window does not guarantee perfect recall, and longer prompts can increase memory use and latency.
What is Microsoft Phi-3 Medium 128K?
Microsoft Phi-3-medium-128k-instruct is a 14-billion-parameter, instruction-tuned, text-only language model in the Phi-3 family. It can summarize documents, answer questions, generate and explain code, solve mathematical problems, classify text, and extract structured information.


Sources 6
Provider

About Microsoft Copilot