Phi-4

Phi-4-mini-reasoning

by Microsoft Copilot · Preview; currently available through Microsoft Foundry and open-weight distribution

Microsoft Phi-4-mini-reasoning is a 3.8B open-weight model specialized in multi-step mathematical reasoning. It offers a 128K-token context window, strong provider-reported math benchmark results, MIT licensing, and deployment options across local runtimes, ONNX Runtime, Azure AI Foundry, and edge-oriented hardware.

Text Reasoning Coding
Phi-4-mini-reasoning is Microsoft’s compact model for mathematical reasoning when a larger model would create unnecessary memory, cost, or latency requirements. It accepts and produces text, supports a 128K-token context window, and is available as an open-weight checkpoint under the MIT license. Its strongest use cases include mathematical tutoring, symbolic problem solving, proof assistance, educational software, and private or edge deployment.
Outputs

What Phi-4-mini-reasoning can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Fine-tuning
Model profile

Performance characteristics

8/10 Reasoning
6/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Phi-4
Model type Reasoning
Context window 128K tokens
Maximum output 128K tokens
Release date 2025-04-30
Status Preview; currently available through Microsoft Foundry and open-weight distribution
Knowledge cutoff notes

No explicit authoritative knowledge-cutoff date for the exact Phi-4-mini-reasoning model was verified in the available Microsoft model card, Foundry catalog, or launch material. The model was trained on synthetic mathematical data and is not suitable for assuming current factual knowledge.

Model notes

Phi-4-mini-reasoning is a 3.8B-parameter dense decoder-only Transformer in the Phi-4 family. Microsoft reports AIME 2024 57.5%, MATH-500 94.6%, and GPQA Diamond 52.0%. The model is optimized for mathematical reasoning and was trained with synthetic reasoning data, supervised fine-tuning, preference optimization, and reinforcement learning with verifiable rewards. It is distributed under the MIT license as an open-weight checkpoint. Microsoft Foundry lists the exact model as Preview with text input, text output, a 128K context window, and no tool calling. Fine-tuning refers primarily to self-hosted or checkpoint-based customization; a managed fine-tuning offering for this exact model was not verified. Microsoft pricing material does not expose a fixed public per-token price for this model, so input and output prices are left unknown. The model's factual knowledge is limited by its compact size and it was designed and tested primarily for math reasoning.

Model guide

Phi-4-mini-reasoning: Microsoft’s Compact Model for Mathematical Problem Solving

Microsoft Phi-4-mini-reasoning is a 3.8-billion-parameter open-weight reasoning model specialized in multi-step mathematical and logic problems. It combines a 128K-token context window, reported strong mathematics benchmarks, MIT licensing, and deployment options spanning local runtimes, ONNX Runtime, Azure AI Foundry, and compatible inference frameworks.

What is Phi-4-mini-reasoning?

Phi-4-mini-reasoning is a 3.8-billion-parameter language model from Microsoft’s Phi-4 family. It is designed specifically for multi-step mathematical reasoning rather than broad, unrestricted assistant behavior. In practical terms, the model is intended to work through problems, show intermediate reasoning-oriented responses, and produce solutions for mathematical or logic-intensive tasks.

The model is based on the Phi-4-Mini architecture and uses a dense decoder-only Transformer design. “Dense” means that the model uses its full network for each generated response, while “decoder-only” describes the architecture commonly used for text generation. Its relatively small parameter count makes it more practical to run on local computers, edge hardware, and latency-sensitive systems than much larger reasoning models.

Microsoft introduced Phi-4-mini-reasoning on April 30, 2025, alongside Phi-4-reasoning and Phi-4-reasoning-plus. The model is available as an open-weight checkpoint through Hugging Face and is also listed in Microsoft Foundry, where the exact model is marked as Preview.

Technical specifications

SpecificationVerified detail
ProviderMicrosoft
Model familyPhi-4
Parameters3.8 billion
ArchitectureDense decoder-only Transformer
Primary specializationMathematical and logic reasoning
InputText
OutputGenerated text
Context window128,000 tokens
Maximum output listed in Microsoft FoundryUp to 128,000 tokens
LicenseMIT
Hosted statusPreview in Microsoft Foundry

The 128K context window allows the model to receive a large amount of text in one request. This can be useful for long problem statements, collections of definitions, mathematical notes, or multi-step exercises. Microsoft Foundry lists up to 128K output tokens for its hosted model entry, although the practical output limit can vary with the runtime, deployment configuration, memory availability, and serving framework.

Reasoning performance and training

Phi-4-mini-reasoning was trained through a multi-stage process involving synthetic reasoning-data distillation, supervised fine-tuning, preference optimization, and reinforcement learning with verifiable rewards. Microsoft reports that the training data included more than one million synthetic mathematical problems covering levels from middle school to doctoral mathematics.

On Microsoft’s reported evaluation set, the model achieved 57.5% on AIME 2024, 94.6% on MATH-500, and 52.0% on GPQA Diamond. Microsoft reports the base Phi-4-Mini model at 10.0%, 71.8%, and 36.9% on those respective benchmarks. These figures are provider-reported results, not an independent guarantee of performance on every problem or deployment.

The results suggest that Phi-4-mini-reasoning delivers unusually strong mathematical performance for a model of its size. However, mathematics benchmarks should not be treated as a general measure of factual accuracy, writing quality, or real-world reliability. Microsoft notes that the model’s limited size also limits the amount of factual knowledge it can store. Applications requiring current or broad factual information may therefore need retrieval and external verification.

Modalities and supported features

Phi-4-mini-reasoning is text-only. It accepts text input and produces text output; the supplied model specifications do not identify native image, audio, video, speech, or music support. It is therefore not the right choice for interpreting diagrams, screenshots, photographs, recordings, or video unless another system first converts those materials into text.

The available Microsoft Foundry information does not list native tool calling or web-search grounding for this model. It should not be assumed to browse the web, retrieve current information, call functions, or interact with external systems by itself. Developers can potentially place it inside a larger application that supplies tools or retrieval, but those capabilities would come from the surrounding application rather than from a verified built-in model feature.

The model’s coding score in the supplied evaluation data is an editorial assessment rather than a Microsoft-published specification. Phi-4-mini-reasoning can be useful for mathematical code, symbolic calculations, and programming exercises, but the model’s documented specialization is mathematical reasoning, not general software engineering.

Where can it be deployed?

The canonical open-weight checkpoint is microsoft/Phi-4-mini-reasoning on Hugging Face. Its MIT license permits broad use subject to the license terms, making the model suitable for experimentation, private inference, embedded applications, and customized deployments.

It can be loaded with compatible Transformers-based tooling and served through inference frameworks that support the model. The supplied research identifies vLLM, SGLang, Ollama, llama.cpp, and other Phi-compatible runtimes as possible deployment environments. Exact compatibility, quantization support, memory requirements, and performance depend on the runtime and hardware configuration.

Microsoft also publishes ONNX Runtime variants for CPU, GPU, and NPU-oriented deployment. These versions are intended for environments such as desktops, servers, mobile hardware, and supported edge devices. Microsoft has additionally listed Phi-4-mini-reasoning in Azure AI Foundry and Foundry Local catalogs, giving users a choice between managed hosting and local execution.

Pricing and access

There is no verified fixed public per-token price for Phi-4-mini-reasoning in the supplied Microsoft pricing material. Its input and output prices are therefore unknown rather than zero. When run from the open-weight checkpoint, the financial cost depends on the user’s hardware, hosting arrangement, electricity, storage, and operational requirements. When deployed through Azure or another managed service, the applicable price may depend on the selected deployment and service terms.

The model is available through open-weight distribution and appears in Microsoft Foundry as a Preview model. Preview availability means that catalog status, regional access, deployment behavior, and service terms may change. Local availability can also depend on whether the chosen runtime and hardware support the checkpoint or its optimized formats.

Strengths and limitations

Strengths

  • Strong mathematics performance for its size: Microsoft’s reported benchmark results show a substantial improvement over the base Phi-4-Mini model on the cited evaluations.
  • Lower deployment burden: With 3.8 billion parameters, it is positioned for environments where a larger reasoning model would require too much memory, compute, or response time.
  • Long context: The 128K-token context window can accommodate lengthy problem statements, reference material, and multi-step work.
  • Open-weight access: The MIT license and Hugging Face checkpoint support local experimentation and application-specific deployment.
  • Hardware flexibility: Microsoft provides deployment paths involving standard inference runtimes and ONNX variants for CPU, GPU, and NPU hardware.

Limitations

  • Narrow specialization: It is primarily trained and evaluated for mathematical reasoning, so it should not automatically be treated as a general-purpose replacement for larger assistant models.
  • No native multimodal input: The model does not natively process images, audio, or video.
  • Limited general knowledge: Its compact size can lead to factual errors and weaker coverage of broad world knowledge.
  • No verified built-in web search or tool calling: Current information retrieval and external actions require an application layer if they are needed.
  • Preview and variable hosting terms: Microsoft Foundry availability and managed-service pricing may change, and no fixed public token price was verified.
  • Reasoning is not a guarantee: A detailed-looking solution can still contain an incorrect assumption, arithmetic error, or invalid proof step.

Best use cases

Phi-4-mini-reasoning is a good fit for applications where mathematical reasoning quality matters more than broad conversational coverage. Examples include educational tutors, step-by-step problem-solving assistants, mathematics practice tools, lightweight proof assistance, symbolic calculation workflows, and research into small reasoning models.

Its local deployment options make it especially relevant when prompts or solutions should remain on privately controlled infrastructure. It can also suit edge or embedded products that need a reasoning-capable model without the resource demands of a frontier-scale system. The long context window is useful when an application needs to provide extensive instructions, a large set of definitions, or several related problems in one request.

For production use, outputs should be checked with deterministic calculations, a symbolic mathematics system, a proof checker, retrieval, or human review where errors matter. The model’s mathematical specialization does not eliminate the need to verify solutions.

When to choose Phi-4-mini-reasoning

Choose Phi-4-mini-reasoning when the main task is mathematical or logic-heavy, the deployment budget is constrained, and local or edge inference is valuable. It is particularly attractive when a 3.8-billion-parameter model can provide sufficient reasoning quality while offering lower expected infrastructure and latency requirements than a much larger model.

Choose a larger general-purpose reasoning model instead when the application requires stronger broad factual knowledge, more reliable general coding, richer language coverage, native multimodal processing, web research, or built-in tool use. Choose a multimodal model when the inputs include images, diagrams, audio, or video. Choose a retrieval-enabled system when current information is central to the task.

Compared with the base Phi-4-Mini model, Phi-4-mini-reasoning is the more appropriate choice for deliberate, multi-step mathematical problem solving according to Microsoft’s reported benchmark comparison. The trade-off is that its reasoning-focused training does not make it a universal assistant, and its compact design still places limits on general knowledge and reliability.

Bottom line

Phi-4-mini-reasoning is a focused small reasoning model rather than a general AI platform. Its key proposition is the combination of mathematical reasoning performance, a 128K-token context window, open-weight MIT licensing, and deployment options that extend from local inference to Microsoft Foundry. Those characteristics make it useful for education, mathematics software, proof-oriented experimentation, and resource-constrained deployments.

Its limitations are equally important: text-only operation, no verified native web search or tool calling, unknown hosted token pricing, preview status in Microsoft Foundry, and weaker general factual coverage than larger models. It is best selected for verified mathematical workloads where compact deployment matters, not as a standalone source of current facts or a universal multimodal assistant.


Answers to Frequently Asked Questions

Where can Phi-4-mini-reasoning be deployed?
The open-weight checkpoint is available as microsoft/Phi-4-mini-reasoning on Hugging Face and can be used with compatible Transformers tooling and inference runtimes such as vLLM, SGLang, Ollama, and llama.cpp. Microsoft also provides ONNX Runtime variants and lists the model in Azure AI Foundry and Foundry Local catalogs.
Can Phi-4-mini-reasoning process images, use web search, or call tools?
Phi-4-mini-reasoning is text-only and does not natively process images, audio, or video. The available specifications also do not verify built-in web search or tool calling. These capabilities can be added by a surrounding application using retrieval or external tools.
How well does Phi-4-mini-reasoning perform on mathematical benchmarks?
Microsoft reports scores of 57.5% on AIME 2024, 94.6% on MATH-500, and 52.0% on GPQA Diamond. These are provider-reported benchmark results and do not guarantee accuracy on every mathematical problem or real-world deployment.
What are the main technical specifications of Phi-4-mini-reasoning?
The model has 3.8 billion parameters, a 128,000-token context window, text input and output, and a dense decoder-only Transformer architecture. It is released under the MIT license and is listed as a Preview model in Microsoft Foundry.
What is Phi-4-mini-reasoning?
Phi-4-mini-reasoning is Microsoft’s 3.8-billion-parameter language model designed for multi-step mathematical and logic reasoning. It uses a dense decoder-only Transformer architecture and is intended for tasks such as solving math problems, symbolic calculations, and educational assistance.


Sources 7
Provider

About Microsoft Copilot