Phi-4

Phi-4-multimodal-instruct

by Microsoft Copilot · Available; open-weight; listed in Microsoft Foundry and Hugging Face

Microsoft's Phi-4-multimodal-instruct is a 5.6B-parameter open-weight model that accepts text, images, and audio and returns text. It supports OCR, chart interpretation, speech recognition, translation, summarization, and image question answering, with a 131,072-token context and 4,096-token hosted output limit. The model is available through Hugging Face and Microsoft Foundry, but has no native media generation, documented tool calling, or independent web access, and current numeric hosted pricing is unavailable.

Text Reasoning Coding
Microsoft Phi-4-multimodal-instruct combines a Phi-4-mini language backbone with vision and speech components. It accepts text, images, and audio in a context window of up to 131,072 tokens and produces text responses. The model is available as an open-weight checkpoint through Hugging Face and is also listed in Microsoft Foundry, giving developers a choice between local inference and hosted access. Its main appeal is multimodal understanding at a relatively small model size, while its main limitations are text-only output, no documented native tool calling, static knowledge, and potentially demanding local hardware requirements.
Outputs

What Phi-4-multimodal-instruct can produce

Text
Inputs

What it can understand

Text Images Audio Multimodal input
Capabilities

Supported features

Streaming Fine-tuning
Model profile

Performance characteristics

7/10 Reasoning
7/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Phi-4
Model type Multimodal
Context window 131K tokens
Maximum output 4K tokens
Knowledge cutoff June 2024
Release date 2025-02-26
Status Available; open-weight; listed in Microsoft Foundry and Hugging Face
Knowledge cutoff notes

The Microsoft model card describes the checkpoint as a static model trained on offline datasets and states that the cutoff for publicly available data is June 2024. External retrieval or web search would not change this underlying cutoff.

Model notes

The exact model has approximately 5.6 billion parameters and uses a Phi-4-mini language backbone with vision and speech encoders and adapters. It accepts text, images, and audio and returns generated text only. The Hugging Face model card documents local Transformers inference and supervised fine-tuning examples for vision and speech. Microsoft Foundry lists the exact model with a 131,072-token input context, 4,096-token output limit, text/image/audio input, text response format, and no tool calling. The model card identifies June 2024 as the cutoff for publicly available data. Current Azure pricing pages show placeholder dollar signs rather than numeric rates; launch materials listed historical Foundry rates separately for text/image and audio input, but those figures should not be treated as current pricing. Local inference may require high-memory GPUs and Flash Attention 2 or an eager-attention fallback.

Cost

Model pricing

Input Not publicly listed as a current numeric price; Azure pricing page displays $-
Output Not publicly listed as a current numeric price; Azure pricing page displays $-
Model guide

Phi-4-multimodal-instruct: A Compact Open-Weight Model for Image and Audio Understanding

Phi-4-multimodal-instruct is Microsoft's 5.6-billion-parameter open-weight model for understanding text, images, and audio. It returns text rather than generating images, video, or audio, making it suited to OCR, chart interpretation, speech recognition, translation, document analysis, and compact private deployments.

What is Phi-4-multimodal-instruct?

Phi-4-multimodal-instruct is Microsoft's multimodal member of the Phi-4 model family. It is an approximately 5.6-billion-parameter transformer designed to interpret several input types in one application: text, images, and audio. In practical terms, a developer can submit a written question alongside a photograph, document page, chart, or audio recording and ask the model to produce a textual answer.

The model uses a Phi-4-mini-based language backbone together with vision and speech encoders and associated adapters. Microsoft distributes the checkpoint through its Hugging Face organization, and the exact model is also listed in Microsoft Foundry. That combination makes it relevant both to developers experimenting with local open-weight inference and to teams that prefer a managed Microsoft-hosted route.

Its role is understanding rather than media creation. Phi-4-multimodal-instruct can analyze an image or audio recording, but its response is text. It does not natively generate an image, video, music track, speech recording, or other audio output.

Supported inputs and output

The model accepts text, image, and audio input. Its documented multimodal tasks include visual question answering, optical character recognition (OCR), chart and table interpretation, image comparison, speech recognition, speech translation, spoken-query question answering, summarization, and broader audio understanding.

CapabilitySupport
Text inputYes
Image inputYes
Audio inputYes
Video inputNot documented for this model
Text outputYes
Image, video, or audio outputNo

This input-output distinction is important when selecting the model. It can turn a spoken question into a written answer or extract information from an image, but it is not a generative media model. An application that needs synthesized speech, image creation, or video generation would need a different model or an additional generation service.

Context window and technical specifications

Microsoft Foundry lists a 131,072-token input context and a maximum hosted output of 4,096 tokens. A token is a small unit of text used by the model; the context limit represents the material the model can consider in one request, including text and the representation of supported multimodal inputs. The large context window is useful for long documents, multiple images, extended transcripts, or combinations of written and spoken evidence, although actual usable capacity can depend on the deployment and input format.

  • Parameters: approximately 5.6 billion
  • Model family: Phi-4
  • Input context: 131,072 tokens
  • Maximum hosted output: 4,096 tokens
  • Release date: February 26, 2025
  • Knowledge cutoff: June 2024, according to the model card
  • Text languages: more than 20 languages are documented, including English, Chinese, French, German, Japanese, Spanish, Portuguese, Arabic, and Ukrainian
  • Vision language: English
  • Audio languages: English, Chinese, German, French, Italian, Japanese, Spanish, and Portuguese

The June 2024 cutoff applies to the underlying offline-trained checkpoint. Phi-4-multimodal-instruct does not provide current web knowledge by itself, so current events, live prices, changing documentation, and other time-sensitive information require an external retrieval system if the deployment supports one.

What the model does well

The strongest use cases are tasks where the input is not purely text and the required answer is a compact written explanation. For example, an application can use it to read text from a photographed form, describe the key trend in a chart, summarize a meeting recording, translate supported speech, or answer a question about several images.

Its relatively small parameter count is an advantage compared with much larger frontier multimodal models. A smaller model can be easier to deploy privately, run closer to the data, or integrate into an application with tighter infrastructure and cost constraints. The supplied editorial assessment rates its speed and cost efficiency favorably, but those are comparative editorial judgments rather than Microsoft-published benchmark results.

The model also supports supervised fine-tuning according to the Hugging Face model card. This can be useful when a team has task-specific examples for a vision or speech workflow and wants to adapt the open-weight checkpoint rather than rely only on prompting. Fine-tuning requirements, hardware needs, and resulting quality will depend on the dataset and deployment setup.

Reasoning, coding, and tool use

Phi-4-multimodal-instruct can follow written instructions and produce explanations, classifications, summaries, and code-related text. It is therefore suitable for applications that combine visual or audio evidence with ordinary language-model tasks, such as asking for a structured description of a screenshot or generating a short script based on information extracted from a document.

However, the supplied research does not establish frontier-level reasoning or coding performance through benchmark results. The model should be evaluated on the specific tasks that matter to an application rather than assumed to match larger general-purpose models. Its strength is the combination of compactness and multimodal understanding, not a documented specialized reasoning mode.

Microsoft Foundry documentation for this exact model does not list native tool calling. The model can describe an action or generate code for an external tool, but an application should not assume that it can reliably invoke functions, browse the web, or perform transactions without an orchestration layer supplied by the developer. The model's web-search capability is recorded as unsupported, and its knowledge remains tied to the static checkpoint unless external retrieval is added.

Deployment options and practical constraints

There are two main deployment paths. Developers can download the open-weight checkpoint and run it locally with Transformers and Microsoft's multimodal implementation. Microsoft also provides official ONNX Runtime GenAI examples for local inference. Alternatively, the model can be consumed through Microsoft Foundry where the listed hosted interface and limits apply.

Local deployment provides greater control over data handling and can be attractive for private document or audio workloads. It is not necessarily lightweight in an absolute sense, however. The model card notes that local inference may require substantial GPU memory. Depending on the hardware, Flash Attention 2 may be needed, or the eager-attention implementation may be used as a fallback. Hardware compatibility and memory requirements should therefore be tested before committing to an on-device architecture.

Hosted deployment avoids much of the infrastructure work, but availability and pricing are separate questions. The exact model is listed as available in Microsoft Foundry and Hugging Face. Current Azure pricing material supplied for this model does not provide a numeric public rate; the pricing page displays placeholder dollar signs. Historical launch materials contained rates for particular Foundry input categories, but those figures should not be treated as current pricing. This model should therefore be costed using a current Microsoft quotation or deployment estimate rather than an assumed per-token number.

Limitations to consider

  • Text-only output: it cannot directly produce images, video, speech, music, or other audio.
  • No documented native tool calling: function execution, browsing, and external actions require application-level integration.
  • Static knowledge: the model card gives June 2024 as the public-data cutoff, so it cannot independently answer reliably about later events.
  • Language asymmetry: the documented vision language is English, while the listed audio-language support is narrower than the broader text-language coverage.
  • Local hardware demands: open-weight access does not mean that the model will run comfortably on every laptop or edge device.
  • Output limit: Microsoft Foundry lists a 4,096-token maximum hosted response, which may require staged processing for very long generated reports.
  • Quality variability: OCR, chart interpretation, transcription, translation, and reasoning outputs can still contain errors and should be checked in high-impact workflows.

When to choose Phi-4-multimodal-instruct

Choose Phi-4-multimodal-instruct when the application needs to understand images or audio but only needs a written response. It is a particularly reasonable candidate for private document processing, OCR, chart and table analysis, speech transcription, speech translation, image question answering, and compact multimodal assistants. Its open-weight distribution also makes it more suitable than a closed hosted-only model when local deployment, customization, or data locality is important.

It may be preferable to a much larger multimodal model when response speed, infrastructure cost, or deployment control matters more than maximum general reasoning quality. The supplied editorial scores rate reasoning and coding at 7 and speed at 8, but these scores are internal comparative assessments, not provider-published guarantees or benchmark results.

Another option is more appropriate when the application needs generated media, native function calling, dependable web-connected answers, advanced autonomous coding, or the strongest available frontier reasoning. It is also worth considering a larger or specialized speech, vision, or retrieval system when accuracy requirements exceed what a compact general-purpose multimodal checkpoint can provide.

Overall assessment

Phi-4-multimodal-instruct occupies a practical middle ground: it is substantially more capable than a text-only compact model for applications involving photographs, documents, charts, and recordings, while remaining much smaller and more controllable than the largest multimodal systems. Its defining trade-off is clear. It provides broad multimodal understanding and text generation, but not multimodal generation, built-in current information, or documented native tools.

For teams that can supply retrieval and tool orchestration externally, and that are willing to validate local hardware or hosted costs, it offers a focused way to add image and audio understanding without adopting a much larger model. For media creation or highly autonomous workflows, its text-only, static, tool-free design makes it the wrong component unless paired with additional systems.


Answers to Frequently Asked Questions

Does Phi-4-multimodal-instruct support web search and native tool calling?
No native web search or tool calling is documented for this model. Browsing, function execution, transactions, and other external actions require an application-level orchestration or retrieval layer. Its underlying knowledge cutoff is June 2024, so it cannot independently provide reliable current information.
Can Phi-4-multimodal-instruct be run locally?
Yes. Developers can download the open-weight checkpoint from Hugging Face and run it locally with Transformers, Microsoft's multimodal implementation, or official ONNX Runtime GenAI examples. Local inference may require substantial GPU memory, and hardware compatibility should be tested before deployment.
Can Phi-4-multimodal-instruct generate images, audio, or video?
No. Phi-4-multimodal-instruct is designed for media understanding, not media creation. It can analyze images and audio recordings and return a written response, but it does not natively generate images, video, speech, music, or other audio.
What are the context window and language capabilities of Phi-4-multimodal-instruct?
Microsoft Foundry lists a 131,072-token input context and a maximum hosted output of 4,096 tokens. The model supports text in more than 20 documented languages, while its documented vision language is English and its audio support includes English, Chinese, German, French, Italian, Japanese, Spanish, and Portuguese.
What is Phi-4-multimodal-instruct?
Phi-4-multimodal-instruct is Microsoft's approximately 5.6-billion-parameter multimodal model for understanding text, images, and audio. It accepts multimodal inputs and produces text responses for tasks such as OCR, chart interpretation, speech recognition, translation, summarization, and visual question answering.


Sources 6
Provider

About Microsoft Copilot