What is Phi-4-mini-reasoning?
Phi-4-mini-reasoning is a 3.8-billion-parameter language model from Microsoft’s Phi-4 family. It is designed specifically for multi-step mathematical reasoning rather than broad, unrestricted assistant behavior. In practical terms, the model is intended to work through problems, show intermediate reasoning-oriented responses, and produce solutions for mathematical or logic-intensive tasks.
The model is based on the Phi-4-Mini architecture and uses a dense decoder-only Transformer design. “Dense” means that the model uses its full network for each generated response, while “decoder-only” describes the architecture commonly used for text generation. Its relatively small parameter count makes it more practical to run on local computers, edge hardware, and latency-sensitive systems than much larger reasoning models.
Microsoft introduced Phi-4-mini-reasoning on April 30, 2025, alongside Phi-4-reasoning and Phi-4-reasoning-plus. The model is available as an open-weight checkpoint through Hugging Face and is also listed in Microsoft Foundry, where the exact model is marked as Preview.
Technical specifications
| Specification | Verified detail |
|---|---|
| Provider | Microsoft |
| Model family | Phi-4 |
| Parameters | 3.8 billion |
| Architecture | Dense decoder-only Transformer |
| Primary specialization | Mathematical and logic reasoning |
| Input | Text |
| Output | Generated text |
| Context window | 128,000 tokens |
| Maximum output listed in Microsoft Foundry | Up to 128,000 tokens |
| License | MIT |
| Hosted status | Preview in Microsoft Foundry |
The 128K context window allows the model to receive a large amount of text in one request. This can be useful for long problem statements, collections of definitions, mathematical notes, or multi-step exercises. Microsoft Foundry lists up to 128K output tokens for its hosted model entry, although the practical output limit can vary with the runtime, deployment configuration, memory availability, and serving framework.
Reasoning performance and training
Phi-4-mini-reasoning was trained through a multi-stage process involving synthetic reasoning-data distillation, supervised fine-tuning, preference optimization, and reinforcement learning with verifiable rewards. Microsoft reports that the training data included more than one million synthetic mathematical problems covering levels from middle school to doctoral mathematics.
On Microsoft’s reported evaluation set, the model achieved 57.5% on AIME 2024, 94.6% on MATH-500, and 52.0% on GPQA Diamond. Microsoft reports the base Phi-4-Mini model at 10.0%, 71.8%, and 36.9% on those respective benchmarks. These figures are provider-reported results, not an independent guarantee of performance on every problem or deployment.
The results suggest that Phi-4-mini-reasoning delivers unusually strong mathematical performance for a model of its size. However, mathematics benchmarks should not be treated as a general measure of factual accuracy, writing quality, or real-world reliability. Microsoft notes that the model’s limited size also limits the amount of factual knowledge it can store. Applications requiring current or broad factual information may therefore need retrieval and external verification.
Modalities and supported features
Phi-4-mini-reasoning is text-only. It accepts text input and produces text output; the supplied model specifications do not identify native image, audio, video, speech, or music support. It is therefore not the right choice for interpreting diagrams, screenshots, photographs, recordings, or video unless another system first converts those materials into text.
The available Microsoft Foundry information does not list native tool calling or web-search grounding for this model. It should not be assumed to browse the web, retrieve current information, call functions, or interact with external systems by itself. Developers can potentially place it inside a larger application that supplies tools or retrieval, but those capabilities would come from the surrounding application rather than from a verified built-in model feature.
The model’s coding score in the supplied evaluation data is an editorial assessment rather than a Microsoft-published specification. Phi-4-mini-reasoning can be useful for mathematical code, symbolic calculations, and programming exercises, but the model’s documented specialization is mathematical reasoning, not general software engineering.
Where can it be deployed?
The canonical open-weight checkpoint is microsoft/Phi-4-mini-reasoning on Hugging Face. Its MIT license permits broad use subject to the license terms, making the model suitable for experimentation, private inference, embedded applications, and customized deployments.
It can be loaded with compatible Transformers-based tooling and served through inference frameworks that support the model. The supplied research identifies vLLM, SGLang, Ollama, llama.cpp, and other Phi-compatible runtimes as possible deployment environments. Exact compatibility, quantization support, memory requirements, and performance depend on the runtime and hardware configuration.
Microsoft also publishes ONNX Runtime variants for CPU, GPU, and NPU-oriented deployment. These versions are intended for environments such as desktops, servers, mobile hardware, and supported edge devices. Microsoft has additionally listed Phi-4-mini-reasoning in Azure AI Foundry and Foundry Local catalogs, giving users a choice between managed hosting and local execution.
Pricing and access
There is no verified fixed public per-token price for Phi-4-mini-reasoning in the supplied Microsoft pricing material. Its input and output prices are therefore unknown rather than zero. When run from the open-weight checkpoint, the financial cost depends on the user’s hardware, hosting arrangement, electricity, storage, and operational requirements. When deployed through Azure or another managed service, the applicable price may depend on the selected deployment and service terms.
The model is available through open-weight distribution and appears in Microsoft Foundry as a Preview model. Preview availability means that catalog status, regional access, deployment behavior, and service terms may change. Local availability can also depend on whether the chosen runtime and hardware support the checkpoint or its optimized formats.
Strengths and limitations
Strengths
- Strong mathematics performance for its size: Microsoft’s reported benchmark results show a substantial improvement over the base Phi-4-Mini model on the cited evaluations.
- Lower deployment burden: With 3.8 billion parameters, it is positioned for environments where a larger reasoning model would require too much memory, compute, or response time.
- Long context: The 128K-token context window can accommodate lengthy problem statements, reference material, and multi-step work.
- Open-weight access: The MIT license and Hugging Face checkpoint support local experimentation and application-specific deployment.
- Hardware flexibility: Microsoft provides deployment paths involving standard inference runtimes and ONNX variants for CPU, GPU, and NPU hardware.
Limitations
- Narrow specialization: It is primarily trained and evaluated for mathematical reasoning, so it should not automatically be treated as a general-purpose replacement for larger assistant models.
- No native multimodal input: The model does not natively process images, audio, or video.
- Limited general knowledge: Its compact size can lead to factual errors and weaker coverage of broad world knowledge.
- No verified built-in web search or tool calling: Current information retrieval and external actions require an application layer if they are needed.
- Preview and variable hosting terms: Microsoft Foundry availability and managed-service pricing may change, and no fixed public token price was verified.
- Reasoning is not a guarantee: A detailed-looking solution can still contain an incorrect assumption, arithmetic error, or invalid proof step.
Best use cases
Phi-4-mini-reasoning is a good fit for applications where mathematical reasoning quality matters more than broad conversational coverage. Examples include educational tutors, step-by-step problem-solving assistants, mathematics practice tools, lightweight proof assistance, symbolic calculation workflows, and research into small reasoning models.
Its local deployment options make it especially relevant when prompts or solutions should remain on privately controlled infrastructure. It can also suit edge or embedded products that need a reasoning-capable model without the resource demands of a frontier-scale system. The long context window is useful when an application needs to provide extensive instructions, a large set of definitions, or several related problems in one request.
For production use, outputs should be checked with deterministic calculations, a symbolic mathematics system, a proof checker, retrieval, or human review where errors matter. The model’s mathematical specialization does not eliminate the need to verify solutions.
When to choose Phi-4-mini-reasoning
Choose Phi-4-mini-reasoning when the main task is mathematical or logic-heavy, the deployment budget is constrained, and local or edge inference is valuable. It is particularly attractive when a 3.8-billion-parameter model can provide sufficient reasoning quality while offering lower expected infrastructure and latency requirements than a much larger model.
Choose a larger general-purpose reasoning model instead when the application requires stronger broad factual knowledge, more reliable general coding, richer language coverage, native multimodal processing, web research, or built-in tool use. Choose a multimodal model when the inputs include images, diagrams, audio, or video. Choose a retrieval-enabled system when current information is central to the task.
Compared with the base Phi-4-Mini model, Phi-4-mini-reasoning is the more appropriate choice for deliberate, multi-step mathematical problem solving according to Microsoft’s reported benchmark comparison. The trade-off is that its reasoning-focused training does not make it a universal assistant, and its compact design still places limits on general knowledge and reliability.
Bottom line
Phi-4-mini-reasoning is a focused small reasoning model rather than a general AI platform. Its key proposition is the combination of mathematical reasoning performance, a 128K-token context window, open-weight MIT licensing, and deployment options that extend from local inference to Microsoft Foundry. Those characteristics make it useful for education, mathematics software, proof-oriented experimentation, and resource-constrained deployments.
Its limitations are equally important: text-only operation, no verified native web search or tool calling, unknown hosted token pricing, preview status in Microsoft Foundry, and weaker general factual coverage than larger models. It is best selected for verified mathematical workloads where compact deployment matters, not as a standalone source of current facts or a universal multimodal assistant.

