Microsoft Foundry model-router

model-router

by Microsoft Copilot · Current; latest documented version is 2025-11-18

Microsoft model-router is a Foundry deployment abstraction that dynamically selects eligible underlying models for each request. Balanced, Quality, and Cost modes help teams trade off response quality, latency, and spending, while custom model subsets and automatic failover provide operational control. The router supports text and eligible image inputs, text output, agents, and tool calling, but context size, maximum output, compatibility, and behavior depend on the selected underlying model.

Text Reasoning Coding
Microsoft model-router is designed for teams that do not want to hard-code one language model for every request. Deployed through Microsoft Foundry, it evaluates incoming prompts and routes them to an eligible underlying model using Balanced, Quality, or Cost modes. This can simplify model-selection logic and reduce spending on routine requests, but it also means that context limits, tool compatibility, response behavior, and maximum output length can vary according to the model selected at runtime.
Outputs

What model-router can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Streaming Prompt caching
Model profile

Performance characteristics

8/10 Reasoning
7/10 Coding
7/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Microsoft Foundry model-router
Model type General Purpose
Context window 200K tokens
Release date 2025-05-19
Status Current; latest documented version is 2025-11-18
Knowledge cutoff notes

Microsoft does not publish a separate knowledge cutoff for the model-router routing model. Responses are generated by the underlying model selected for each request, so knowledge behavior depends on that model.

Model notes

Model-router is a routing model and deployment abstraction rather than a single conventional generative model. The latest documented version is 2025-11-18. It routes among eligible models from OpenAI, Anthropic, xAI, DeepSeek, and Meta, with regional and deployment-specific availability. Supported Claude models must be separately deployed before they can be selected. Balanced is the default routing mode; Quality and Cost modes are also available. Custom model subsets can constrain routing for compliance, capability, cost, or context requirements. The effective context window is limited by the smallest underlying model in the deployment, and maximum output tokens vary by the selected underlying model. Model-router accepts image input for vision-enabled chats but does not process audio input. Prompt caching is inherited from the selected underlying model and is not guaranteed across requests. Editorial scores reflect the usefulness of dynamic routing as a model-selection system, not a standalone benchmark of text-generation quality.

Cost

Model pricing

Input Not publicly listed as a numeric rate; Microsoft’s pricing page currently shows $- and directs users to Azure pricing for applicable deployment details.
Output Not separately listed; billing depends on model-router and the underlying routed model pricing configuration.
Model guide

Microsoft Model Router: One Foundry Deployment for Dynamic Model Selection

Microsoft model-router is a deployable Microsoft Foundry routing model that evaluates each request and selects an eligible underlying model according to the chosen balance between quality, cost, and latency. It gives applications one deployment while allowing requests to be handled by supported models from providers including OpenAI, Anthropic, xAI, DeepSeek, and Meta.

What is Microsoft model-router?

Microsoft model-router is a deployable routing model in Microsoft Foundry. Rather than generating every response through one fixed language model, it evaluates a request and selects an eligible model from its configured pool. The application can continue using one deployment while the router makes the model-selection decision for each request.

This makes model-router different from a conventional language model. Its main value is not a single, fixed benchmark profile or a permanently fixed set of capabilities. It is an optimization layer intended to match requests with models that are expected to provide an appropriate combination of quality, latency, cost, and capability.

For example, a routine classification or straightforward question may be sent to a lower-cost model, while a complex reasoning task may be routed to a higher-rated model. The exact choice depends on the routing mode, eligible model pool, region, deployment configuration, permissions, and the characteristics of the request.

Where it fits in Microsoft Foundry

Model-router sits within Microsoft Foundry's catalog of deployable models and model services. The documented model ID is model-router, and the latest documented version supplied for this page is 2025-11-18. Microsoft positions it as a managed alternative to building and maintaining custom application-side routing rules.

The supported pool can include models from OpenAI, Anthropic, xAI, DeepSeek, and Meta. This does not mean that every model from those providers is always available. Eligibility depends on the router version, selected region, deployment type, access permissions, and the model subset configured for the deployment.

Claude models require separate deployment in the Foundry resource before model-router can invoke them. Other supported models generally do not need to be separately deployed for use through the router, although availability and requirements remain deployment-specific.

How the routing modes work

Microsoft documents three routing modes:

  • Balanced: The default mode. It seeks a practical trade-off between response quality and cost.
  • Quality: Prioritizes the highest-rated eligible model for a request. This is intended for more demanding reasoning or higher-impact outputs where quality matters more than expense.
  • Cost: Favors lower-cost models when they are expected to provide adequate results. This is suited to high-volume workloads with many routine requests.

Applications can also restrict routing to a custom subset of models. A custom subset can help enforce regional, compliance, capability, cost, or context requirements. It can also make behavior more predictable than allowing the router to use every otherwise eligible model.

These modes are trade-off controls, not guarantees that every request will use one particular model or produce a particular quality level. Teams should evaluate the selected mode against representative prompts, including failure cases and tool-calling workflows.

Inputs, outputs, and supported modalities

Model-router accepts text input and can accept image input for supported vision-enabled chat scenarios. However, Microsoft documentation states that the routing decision is based on the text portion of the request. Image input therefore does not turn the router into a general vision system with independent image reasoning behavior; the selected underlying model must still support the relevant vision scenario.

The router does not process audio input and does not natively generate images, audio, or video. Its direct output is text. It can be used through Foundry Responses and Chat Completions APIs, and it supports eligible tool-calling workflows and Foundry agents.

Automatic failover can redirect a request to another eligible model when the selected model encounters a transient issue. This can improve service resilience, but failover may also change response characteristics because the replacement model can have different limits, tool behavior, or output style.

Context window and output limits

The supplied documentation lists a context length of up to 200,000 tokens for the model-router entry. In practice, the effective context window is constrained by the smallest underlying model available to the deployment. A request that fits within the nominal router limit may therefore fail or behave differently if the eligible model selected for that request supports a smaller context.

Custom model subsets can help control this issue. For example, a team handling long documents may restrict routing to models that support the required context size rather than leaving the router free to select a smaller-context option.

A fixed maximum output-token value is not documented for model-router. Maximum output length varies according to the underlying model selected for the request. Applications should therefore avoid assuming that all routed responses share one uniform output limit and should design request handling around the limits documented for the eligible model pool.

Reasoning, coding, and tool use

Model-router can route requests to models appropriate for different levels of reasoning, but it does not have one independent reasoning profile of its own. Complex reasoning performance depends on which eligible model handles the request and on whether the selected routing mode prioritizes quality.

The same distinction applies to coding. Model-router can support coding applications when the selected underlying model is suitable for code generation, debugging, or code-related reasoning. It should not be treated as a single coding model with fixed behavior across requests.

Eligible tool-calling workflows are supported, including use through Foundry agents. Tool compatibility must be tested across the configured model pool because different underlying models may vary in how they interpret tool schemas, handle arguments, support reasoning parameters, or respond to tool results. Some parameters unsupported by a selected reasoning model may be ignored.

The router can be useful for agents that perform several types of work. A classification step, retrieval decision, tool call, and long-form reasoning task do not necessarily need the same model. Model-router can provide a common deployment while selecting among eligible models for these different workloads.

Pricing and availability

Microsoft's pricing page supplied for this entry does not provide a public numeric token rate for model-router; it currently displays a placeholder and directs customers to Azure pricing information. Billing depends on the pricing configuration for the model-router deployment and the underlying routed-model usage. The applicable rate can vary by subscription, region, deployment type, agreement, and the model selected.

Because there is no verified numeric price in the supplied research, a reliable cost comparison should be made in the Azure pricing experience for the intended deployment rather than inferred from a general list price. The Cost routing mode may reduce spend for suitable routine requests, but it is not a promise of a specific percentage reduction.

Model-router supports Global Standard and Data Zone Standard deployment types in supported Microsoft Foundry regions. Individual underlying models can have different regional availability. The latest version and eligible model list should be checked before deployment, particularly when a workload depends on a specific provider or model family.

Main strengths

  • One deployment for mixed workloads: Applications can use a common endpoint and let the router select among eligible models.
  • Configurable trade-offs: Balanced, Quality, and Cost modes provide a direct way to prioritize different operational goals.
  • Reduced routing maintenance: Teams do not need to implement every model-selection rule themselves.
  • Multi-provider model pool: The supported catalog can include models from several providers, subject to access and regional conditions.
  • Operational resilience: Automatic failover can redirect transiently failed requests to another eligible model.
  • Agent and tool integration: The router can participate in eligible tool-calling workflows and Foundry agent deployments.
  • Configurable boundaries: Custom model subsets can constrain routing for compliance, capability, cost, or context requirements.

Important limitations

  • Less deterministic model selection: The same application does not necessarily use the same underlying model for every similar request.
  • Variable capabilities: Context size, maximum output, tool compatibility, supported parameters, latency, and response style depend on the selected model.
  • No audio processing or native media generation: Model-router accepts text and supported image inputs, but it does not process audio or directly generate images, audio, or video.
  • Pricing is deployment-dependent: A public numeric rate is not provided in the supplied pricing information.
  • Performance requires evaluation: A routing mode that works well for one workload may be unsuitable for another, especially when quality and cost requirements differ.
  • Caching is not uniform: Prompt-cache behavior is inherited from the selected underlying model and is not guaranteed across requests.

These limitations are especially important for systems that require reproducibility, fixed model behavior, strict latency targets, or model-specific prompt and tool semantics. A custom model subset can reduce variability, but it does not turn model-router into a fully fixed-model deployment.

When to choose Microsoft model-router

Model-router is a strong fit when an application receives a broad mixture of simple and difficult requests and the team wants quality-cost trade-offs to be managed centrally. Examples include enterprise assistants, classification pipelines, research agents, customer-support systems, and high-volume applications where routine requests do not need the most expensive available model.

It is also useful when a team wants to experiment with multiple supported model providers without rewriting application integration for every model. A custom subset can preserve some control while still allowing runtime selection within an approved group.

A fixed model may be more appropriate when deterministic behavior is a primary requirement, when prompts have been carefully optimized for one model, or when every request must have the same context, output, tool, and latency characteristics. A dedicated lower-cost model can also be preferable for a simple high-volume task if routing overhead and variability provide little benefit. Conversely, a dedicated high-quality reasoning model may be preferable for a critical workflow where predictable performance matters more than automatic cost optimization.

What to evaluate before deployment

Before adopting model-router, test it against a fixed-model baseline using representative production prompts. Measure answer quality, factual and task-specific accuracy, latency, token usage, tool-call success, failure recovery, and cost. Include short prompts, long-context requests, image inputs where relevant, difficult reasoning tasks, and prompts that require structured tool interactions.

Also test whether the application tolerates variation between underlying models. Check maximum response length, unsupported parameters, formatting assumptions, prompt caching, and behavior after automatic failover. If a request needs a known context size or a particular capability, configure a suitable custom model subset instead of relying on the entire eligible pool.

In short, Microsoft model-router is best understood as a managed model-selection system rather than a single general-purpose model. Its value comes from dynamically matching requests to eligible models and exposing practical quality, cost, and latency controls. That flexibility can simplify enterprise deployments, but it must be balanced against the variability that dynamic routing introduces.


Answers to Frequently Asked Questions

When should an application use Microsoft model-router instead of a fixed model?
Model-router is a good fit for applications with mixed workloads, such as enterprise assistants, classification pipelines, research agents, and customer-support systems, where different requests need different quality and cost levels. A fixed model may be better when deterministic behavior, consistent context and output limits, predictable latency, or model-specific prompt and tool behavior is essential.
Does Microsoft model-router support images, audio, tools, and agents?
Model-router accepts text and supported image inputs, but its routing decision is based on the text portion of the request. It does not process audio or natively generate images, audio, or video. It supports eligible tool-calling workflows and Foundry agents, although tool compatibility can vary between underlying models.
Which models and providers can Microsoft model-router use?
The eligible model pool can include models from OpenAI, Anthropic, xAI, DeepSeek, and Meta. Availability depends on the router version, region, deployment type, permissions, and configured model subset. Claude models require separate deployment in the Foundry resource before the router can invoke them.
What is Microsoft model-router in Microsoft Foundry?
Microsoft model-router is a deployable routing model in Microsoft Foundry that evaluates each request and selects an eligible underlying model from a configured pool. Applications can use one deployment while the router dynamically balances quality, cost, latency, and capability.
What routing modes does Microsoft model-router support?
Microsoft model-router supports Balanced, Quality, and Cost modes. Balanced is the default and seeks a trade-off between quality and cost, Quality prioritizes higher-rated eligible models, and Cost favors lower-cost models when they are expected to provide adequate results.


Sources 5
Provider

About Microsoft Copilot