What is Phi-4-reasoning?
Phi-4-reasoning is an open-weight language model from Microsoft Research and part of Microsoft's Phi-4 family. It was released on April 30, 2025, and is based on Phi-4 but further trained for reasoning-heavy tasks. The model is intended to work through problems in areas such as mathematics, science, coding, logic, instruction following, and planning.
At approximately 14 billion parameters, it is smaller than many flagship language models. That size makes it relevant for developers and researchers who need a capable reasoning model that can be deployed on their own infrastructure or on a separately priced managed service. It is not a general-purpose multimodal assistant: its documented interface accepts text and returns text.
The model's responses are designed to include a reasoning block followed by a solution or summary block. This can make the output useful for inspecting how a problem was approached, although a visible reasoning trace should not be treated as a guarantee that every intermediate step is correct.
Where it fits in Microsoft’s catalog
Phi-4-reasoning is a specialized member of the Phi-4 model family rather than a consumer chatbot or a fixed feature of Microsoft Copilot. Microsoft distributes the model through its Hugging Face repository under the MIT license and lists it as a preview model in Microsoft Foundry.
This positioning gives users two broad deployment paths. Teams can download the open weights and run the model locally or on their own servers, subject to the model's license and their hardware capabilities. Alternatively, they can use a managed deployment when available through Microsoft Foundry. The managed catalog listing is described as a preview chat-completion model, so availability and service conditions may change.
Microsoft Copilot is a separate product ecosystem that can expose different models and capabilities depending on the plan, region, and workload. Phi-4-reasoning should therefore not be assumed to be the model behind every Copilot response.
Technical specifications
| Specification | Verified detail |
|---|---|
| Provider | Microsoft |
| Model family | Phi-4 |
| Release date | April 30, 2025 |
| Parameters | Approximately 14 billion |
| Architecture | Dense decoder-only Transformer |
| Input | Text |
| Output | Text, including reasoning and solution sections |
| Context window | 32,768 tokens |
| Maximum output | 32,768 tokens |
| License | MIT |
| Catalog status | Open-weight; listed as a preview model in Microsoft Foundry |
| Primary language focus | English |
A token is a unit of text used by the model; it may represent a word, part of a word, punctuation, or another short text fragment. The 32,768-token context limit covers the material supplied to the model and the generated response according to the relevant service or runtime's accounting. The model card also recommends allowing up to 32,768 new tokens for difficult problems so that extended reasoning has room to develop.
Reasoning and benchmark performance
Phi-4-reasoning was trained with supervised fine-tuning and reinforcement learning on reasoning-oriented data. In practical terms, this means it is intended to spend more computation producing a structured solution than a model optimized primarily for short conversational replies.
Microsoft reports results across mathematics, general reasoning, science, coding, instruction following, and planning evaluations. The documented results include 75.3 on AIME 2024, 62.9 on AIME 2025, 76.6 on OmniMath, 65.8 on GPQA-Diamond, and 53.8 on LiveCodeBench for the evaluation ranges described in the model card. These are provider-reported benchmark results, not guarantees for an individual application. Scores can depend on prompting, sampling settings, evaluation versions, and the amount of allowed reasoning.
The model is most naturally suited to tasks such as solving a multi-step algebra problem, explaining an algorithm, reviewing code logic, deriving a result from scientific information, or converting a loosely stated requirement into a structured plan. It can still make arithmetic, factual, logical, or programming mistakes, so generated answers should be tested or independently checked when accuracy matters.
Coding, tools, and structured output
Coding is one of Phi-4-reasoning's intended strengths. It can generate code, reason about algorithms, explain implementation choices, and help analyze programming problems. Its value is greatest when the task requires planning or careful multi-step reasoning rather than simply producing a short boilerplate snippet.
The supplied catalog information does not document native tool calling or function calling for this model. It should not be treated as an autonomous agent that can browse the web, execute code, access files, or call external services by itself. The model also does not provide a documented native JSON or structured-output mode. Applications that need strict machine-readable output may have to enforce a format externally and validate the result.
Streaming is supported in the supplied model data, but streaming only changes how generated text is delivered. It does not add browsing, code execution, tool use, or multimodal input.
Modalities and important limitations
Phi-4-reasoning is text-only. It accepts text input and produces text output; it does not natively accept images, audio, or video, and it does not generate images, audio, music, or video. A multimodal application would need a separate model or an orchestration layer that converts non-text material into text before sending it to Phi-4-reasoning.
The model is primarily trained and evaluated for English-language reasoning. It may not be the best choice for multilingual applications, especially when consistent quality across many languages is a core requirement. Its long reasoning traces can also increase response time and token consumption. A short-answer model may be more efficient for simple classification, extraction, summarization, or routine chat.
As with other language models, Phi-4-reasoning can generate unsafe or inaccurate content. The supplied model information specifically cautions against relying on it without additional evaluation in high-impact areas such as healthcare, law, employment, finance, and public services.
Deployment options and recommended settings
The model is available from Microsoft's Hugging Face repository and can be run with Transformers. The supplied documentation also identifies compatible serving options including vLLM, SGLang, Ollama, and llama.cpp-compatible deployments. Exact hardware requirements depend on precision, quantization, batch size, context length, and serving configuration; the supplied research does not specify a single minimum hardware configuration.
Microsoft's model guidance recommends a temperature of 0.8, top-k of 50, top-p of 0.95, and sampling enabled. For complex problems, allowing a large maximum output gives the model room for extended reasoning, but it can also increase latency and resource usage. In production, teams should test shorter output limits for routine tasks and reserve the full allowance for problems that genuinely require long reasoning.
Pricing and cost trade-offs
No official numeric per-token price is published for Phi-4-reasoning in the supplied Microsoft pricing information. Because it is open-weight, self-hosting costs are determined by hardware, electricity, storage, engineering effort, quantization, concurrency, and the chosen inference stack. A managed Microsoft Foundry deployment, where available, is separately priced and should not be assumed to be free merely because the weights are openly distributed.
The open-weight license can make Phi-4-reasoning attractive for organizations that want more control over deployment or predictable infrastructure ownership. However, self-hosting shifts operational responsibility to the user. A hosted model may be simpler for low-volume experimentation, while local deployment can become more economical or more private at sustained volume if the organization already has suitable hardware.
Editorially, the model presents a favorable capability-to-size and capability-to-cost profile for reasoning workloads, but this is an evaluation rather than a Microsoft-published price or performance guarantee. Long responses may reduce the apparent cost advantage when a simpler model could solve the same task with fewer tokens.
When to choose Phi-4-reasoning
Choose Phi-4-reasoning when the central task is text-based reasoning and you value open weights, local deployment, or a relatively compact model. Good candidates include:
- Mathematical problem solving and worked explanations.
- Algorithm design, code generation, debugging, and code review.
- Scientific or technical question answering based on supplied text.
- Logic puzzles, structured analysis, and multi-step planning.
- Research and experimentation with a reasoning-focused model that can be run outside a single hosted API.
- Applications where a 32K-token context and potentially long reasoning response are useful.
Another option may be more appropriate when you need image, audio, or video understanding; web-grounded answers; built-in tool use; strict structured responses; broad multilingual coverage; or consistently short, low-latency replies. A larger hosted model may also be preferable for difficult tasks requiring capabilities not documented for Phi-4-reasoning, while a smaller conventional language model may be more efficient for simple text transformations.
Overall assessment
Phi-4-reasoning is a focused open-weight model for users who want Microsoft’s Phi-4 reasoning approach in a text-only deployment. Its 14-billion-parameter size, MIT license, 32K context, coding and mathematics orientation, and support for local serving make it particularly relevant to technical experimentation and resource-conscious reasoning applications.
Its trade-offs are equally important: no native multimodal support, no documented tool calling or structured-output mode, English-focused training and evaluation, potentially long and expensive reasoning traces, and no published official per-token price. It is best evaluated as a specialized reasoning component rather than as a complete assistant platform.

