Phi-4

Phi-4-reasoning-plus

by Microsoft Copilot · Current downloadable open-weight model

Microsoft's Phi-4-reasoning-plus is a 14B open-weight, MIT-licensed text model designed for extended mathematical, scientific, coding, and multi-step reasoning. It offers a 32,768-token context and can generate up to 32,768 new tokens for complex queries, but its longer reasoning traces increase latency and compute requirements. It has no identified official Microsoft per-token API price, native multimodal support, or built-in tool execution.

Text Reasoning Coding
Phi-4-reasoning-plus is a 14-billion-parameter dense language model released by Microsoft Research on April 30, 2025. It is designed to spend more tokens working through difficult problems before presenting an answer, with particular emphasis on mathematics, science, coding, and algorithmic reasoning. Unlike a metered cloud model, it is distributed as downloadable weights under the MIT license, so users generally provide their own hardware or select a third-party hosting service. That makes it attractive for private experimentation and customized deployments, but its longer reasoning traces increase latency, memory requirements, and operating cost.
Outputs

What Phi-4-reasoning-plus can produce

Text
Inputs

What it can understand

Text
Model profile

Performance characteristics

8/10 Reasoning
7/10 Coding
5/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Phi-4
Model type Reasoning
Context window 33K tokens
Maximum output 33K tokens
Knowledge cutoff March 2025
Release date 2025-04-30
Status Current downloadable open-weight model
Knowledge cutoff notes

The model card describes the model as static and trained on an offline dataset with cutoff dates of March 2025 and earlier for publicly available data. This is the underlying training-data cutoff and is not changed by retrieval or external context supplied during inference.

Model notes

Phi-4-reasoning-plus is a 14B dense decoder-only model fine-tuned from Phi-4 with supervised fine-tuning and outcome-based reinforcement learning. Microsoft reports that it generates approximately 50% more tokens on average than Phi-4-reasoning, which can improve accuracy but increase latency and compute requirements. The model card documents a 32K-token context length and recommends allowing up to 32,768 new tokens for complex queries. It is distributed as downloadable weights under the MIT license rather than with an official Microsoft per-token hosted API price. The model is primarily English-focused and was evaluated mainly for mathematical reasoning, with additional evaluations covering science, coding, instruction following, planning, and algorithmic tasks.

Model guide

Phi-4-reasoning-plus: Microsoft’s Open-Weight Model for Extended Mathematical Reasoning

Phi-4-reasoning-plus is Microsoft's 14-billion-parameter open-weight reasoning model for mathematical, scientific, coding, and other multi-step tasks. It extends Phi-4-reasoning with outcome-based reinforcement learning that improves reported reasoning accuracy while producing longer responses and requiring more computation. Its MIT license and downloadable weights make it suitable for local, private, and customized deployments, but it is text-only, English-focused, slower than smaller or non-reasoning models, and has no official Microsoft per-token hosted API price in the supplied research.

What is Phi-4-reasoning-plus?

Phi-4-reasoning-plus is Microsoft's open-weight reasoning model in the Phi-4 family. It is a dense decoder-only Transformer with approximately 14 billion parameters. In practical terms, it generates text one token at a time, but it has been tuned to work through multi-step problems before giving a final solution or summary.

Microsoft Research released the model on April 30, 2025. The model is fine-tuned from Phi-4 using supervised learning on reasoning demonstrations, followed by outcome-based reinforcement learning. The plus version is intended to produce longer reasoning traces than Phi-4-reasoning. Microsoft reports that this additional training improves accuracy on difficult reasoning tasks, although it also increases the number of generated tokens.

Phi-4-reasoning-plus is primarily an English-language text model. It is not a general-purpose multimodal assistant: the supplied model information identifies text as its input and output modality and does not list native image, audio, or video support.

Where it fits in Microsoft's model lineup

Phi-4-reasoning-plus is positioned as a smaller open-weight reasoning model rather than as a Microsoft-hosted consumer assistant or a general multimodal product. Its downloadable weights distinguish it from services accessed through Microsoft Copilot, and its MIT license permits local deployment, research, and commercial use subject to the license and applicable laws.

Within the Phi-4 reasoning work, it is the longer-trace “plus” variant. Compared with a conventional instruction model, its purpose is not simply to answer quickly; it is tuned to spend more generation capacity on problems where intermediate steps can improve the result. That positioning makes it relevant to researchers, developers, and organizations that want to control deployment rather than rely on a fixed hosted endpoint.

Reasoning and benchmark performance

The central trade-off is computation for reasoning performance. Microsoft reports that Phi-4-reasoning-plus generates approximately 50% more tokens on average than Phi-4-reasoning. Longer traces can give the model more room to work through mathematical or scientific problems, but they also make responses slower and more expensive to run on hardware with limited memory or throughput.

In Microsoft's reported evaluations, the model scored 81.3 on AIME 2024, 78.0 on AIME 2025, 81.9 on OmniMath, 68.9 on GPQA Diamond, and 53.1 on the reported LiveCodeBench interval. These are provider-reported benchmark results, not guarantees for an individual application. Benchmark conditions, prompting, sampling settings, and evaluation contamination can all affect results, so production users should test representative workloads independently.

The model produces a reasoning section followed by a solution or summary section. Those generated steps should not automatically be treated as a reliable proof of correctness. The final answer and any intermediate reasoning can still contain mistakes or fabricated claims, particularly outside the model's strongest evaluated areas.

Technical specifications and limits

SpecificationVerified detail
ProviderMicrosoft
Release dateApril 30, 2025
Model familyPhi-4
ParametersApproximately 14 billion
ArchitectureDense decoder-only Transformer
Context length32,768 tokens
Maximum documented new-token allowanceUp to 32,768 tokens for complex queries
Input and outputText input and text output
LicenseMIT
Knowledge cutoffMarch 2025 and earlier public training data, according to the model card

The 32,768-token context window limits the combined amount of prompt context and generated content that an implementation can accommodate. The model documentation also recommends allowing up to 32,768 newly generated tokens for complex queries. Actual usable limits can depend on the inference engine, memory available, serving configuration, and the prompt format.

The model card recommends using its ChatML template and provides sampling guidance of temperature 0.8, top-k 50, top-p 0.95, with sampling enabled. These are documented starting points rather than universal settings. Deterministic or lower-variance applications may need their own testing to decide whether those settings are appropriate.

Coding, tools, and supported modalities

Phi-4-reasoning-plus supports text-based coding assistance and is reported as particularly useful for Python-oriented tasks and common Python libraries. It can help explain algorithms, draft code, work through programming problems, and reason about implementation steps. However, the available information does not establish broad reliability across every programming language, framework, or less common package. Code should be executed and reviewed before use.

The supplied specifications do not list native function calling, tool use, web search, or a built-in code-execution environment. The model should therefore be treated as a text-generation model that can be connected to external systems by a surrounding application, not as an autonomous agent with verified built-in tools.

It accepts and generates text only. There is no verified native image understanding, image generation, audio processing, video processing, speech output, or other non-text output. Applications needing multimodal input or output should select a model with those capabilities or add separate specialized components.

Deployment and pricing

Microsoft distributes Phi-4-reasoning-plus as downloadable model weights through its Hugging Face repository. The model can be loaded with Transformers and served through compatible systems such as vLLM, SGLang, Ollama, and llama.cpp. Quantized and ONNX Runtime variants are also available in the supplied research, which can help reduce memory requirements compared with running the full-precision weights.

There is no official Microsoft per-token hosted API price identified in the supplied research. The model is therefore not directly comparable to a conventional cloud model with a fixed input and output rate. The practical cost depends on hardware purchase or rental, quantization, batch size, serving software, electricity, and the number and length of generated tokens. Third-party hosting may add a usage charge, but its price and availability are provider-specific.

Running a 14-billion-parameter model locally requires substantially more memory than running a small model, and the plus variant's longer responses increase the compute consumed per request. Quantization can lower the hardware barrier, but users should validate that the selected quantized version retains adequate quality for their task.

Main strengths and limitations

Strengths

  • Reasoning specialization: It is explicitly tuned for multi-step mathematical, scientific, coding, and algorithmic problems.
  • Strong reported results for its size: Microsoft's evaluations show competitive performance on several mathematics and science benchmarks for a 14-billion-parameter model.
  • Open deployment: Downloadable weights allow local, private, and customized deployments instead of requiring a Microsoft-hosted endpoint.
  • Permissive licensing: The MIT license supports research and commercial use subject to the license and applicable laws.
  • Extended reasoning traces: The plus training approach is intended to give difficult problems more reasoning capacity.

Limitations

  • Latency and compute: Longer traces increase response time, memory use, and operating cost.
  • Text-only operation: It does not natively handle images, audio, or video.
  • English emphasis: The model is primarily trained and evaluated for English-language use.
  • No verified built-in tools: The supplied specifications do not establish native web search, function calling, or code execution.
  • Limited current knowledge: Its training data is static with a stated cutoff of March 2025 and earlier; current facts require retrieval or another external source.
  • Uncertain production behavior: Reported benchmarks do not remove the need for validation, safety controls, and human review.
  • Coding scope: Performance is focused largely on Python and common libraries, so generated code for other languages or unusual dependencies needs additional checking.

When to choose Phi-4-reasoning-plus

Choose Phi-4-reasoning-plus when the priority is relatively strong mathematical or scientific reasoning from an open model that can be downloaded and controlled. It is a good candidate for local research, mathematical tutoring prototypes, algorithmic problem solving, coding assistance, synthetic-data experiments, private question answering, and evaluation of reasoning-model behavior.

Its open-weight design is especially useful when sending prompts to a hosted service is undesirable, when an organization needs to customize the serving stack, or when a team wants to test quantized or optimized deployments. The MIT license can also simplify commercial experimentation compared with models using more restrictive terms, although legal and compliance review remains the user's responsibility.

Another option may be more appropriate when low latency is more important than extended reasoning, when the workload requires current information without a retrieval layer, or when native multimodal features and built-in tools are essential. A smaller non-reasoning model may serve simple classification or short-answer tasks more economically. A hosted model may be easier for teams that do not want to manage hardware, model files, inference servers, and scaling. A multimodal or tool-enabled model is preferable for applications that must inspect images, call external functions, browse the web, or execute code.

Practical deployment guidance

For an initial evaluation, use the published ChatML formatting and begin with the model card's recommended sampling settings. Test both answer quality and operational behavior: measure time to first response, total generation time, memory consumption, average output length, and error rates on representative problems. Because the plus variant can generate substantially longer traces, a token budget that works for a standard instruction model may not provide a comparable cost or latency profile.

Applications should separate the model's explanation from the final answer where possible, validate mathematical and programming outputs with independent checks, and use retrieval when information beyond the March 2025 cutoff is required. High-stakes decisions should not rely on Phi-4-reasoning-plus without domain-specific review and safeguards.


Answers to Frequently Asked Questions

How can Phi-4-reasoning-plus be deployed?
The downloadable model weights can be obtained through Microsoft's Hugging Face repository and deployed locally or on private infrastructure using systems such as Transformers, vLLM, SGLang, Ollama, and llama.cpp. Quantized and ONNX Runtime variants may reduce memory requirements, while total costs depend on hardware, serving configuration, and generated token volume.
Can Phi-4-reasoning-plus process images, use tools, or browse the web?
No native support is verified for images, audio, video, web search, function calling, or code execution. Phi-4-reasoning-plus accepts and generates text, although external applications can connect it to tools or specialized multimodal components.
How does Phi-4-reasoning-plus compare with Phi-4-reasoning?
Phi-4-reasoning-plus is the longer-trace variant. Microsoft reports that it generates approximately 50% more tokens on average than Phi-4-reasoning, giving it more capacity for difficult problems but also increasing response time, memory usage, and operating cost.
What is Phi-4-reasoning-plus?
Phi-4-reasoning-plus is Microsoft's open-weight reasoning model in the Phi-4 family. It is a dense decoder-only Transformer with approximately 14 billion parameters, designed for extended multi-step reasoning in mathematical, scientific, coding, and algorithmic tasks.
What are the main specifications of Phi-4-reasoning-plus?
The model has approximately 14 billion parameters, a 32,768-token context length, text input and output, and an MIT license. Microsoft released it on April 30, 2025, and the model's stated knowledge cutoff is March 2025 and earlier public training data.


Sources 3
Provider

About Microsoft Copilot