Text-guided biomedical image segmentation, annotation assistance, organ and tumor delineation, pathology-cell analysis, and research-oriented medical imaging pipelines
Type
Other
Reasoning
1/10
Speed
5/10
Multimodal
Image input
Media output
Status
Current; open-weight research model and available for managed deployment through Microsoft Foundry
Input
Not publicly listed as a per-inference model price; Microsoft Foundry deployments incur managed Azure compute charges
Output
Not publicly listed as a per-inference model price; Microsoft Foundry deployments incur managed Azure compute charges
View model
→
First-pass chest X-ray report drafting, structured findings extraction, radiologist workflow assistance, research evaluation, and institution-specific fine-tuning
Type
Multimodal
Reasoning
2/10
Speed
9/10
Multimodal
Image input
Fine-tuning
Status
Limited preview; registration and eligibility approval required
Input
$0.00218 per image for standard inference
Output
Included in the per-image inference price; output is generated text findings
View model
→
Local image captioning, object detection, phrase grounding, region description, OCR and lightweight computer-vision pipelines
Type
Multimodal
Context
1K
Reasoning
3/10
Speed
8/10
Multimodal
Image input
Fine-tuning
Status
Current open-weight model; downloadable from Hugging Face and usable for local or self-hosted inference
View model
→
Local image captioning, object detection, visual grounding, OCR, region annotation and multi-task computer-vision pipelines.
Type
Multimodal
Context
1K
Reasoning
2/10
Speed
7/10
Multimodal
Image input
Fine-tuning
Status
Current open-weight model; publicly accessible through Hugging Face for local or self-hosted inference.
View model
→
Open-weight reasoning, mathematics, coding, research, and general text-generation applications requiring DeepSeek-R1-style reasoning with Microsoft post-training
Type
Reasoning
Context
164K
Reasoning
8/10
Speed
3/10
Fine-tuning
Status
Available as an open-weights model and through Microsoft Foundry hosted API
View model
→
High-quality text-to-image generation, photorealistic imagery, product and marketing visuals, presentation graphics, and precise image-to-image editing.
Type
Other
Context
131K
Reasoning
6/10
Speed
6/10
Multimodal
Image input
Media output
Input
$5 per 1 million text input tokens; $8 per 1 million image input tokens
Output
$47 per 1 million image output tokens
View model
→
High-quality text-to-image generation, controlled image editing, commercial imagery, product and branding visuals, photorealistic scenes, and multi-reference creative workflows
Type
Multimodal
Reasoning
2/10
Speed
7/10
Multimodal
Image input
Media output
View model
→
Fast, cost-conscious text-to-image generation, image editing, creative production, concept visualization, and high-volume image workflows
Type
Multimodal
Context
32K
Reasoning
5/10
Speed
9/10
Multimodal
Image input
Media output
Status
Public preview; scheduled for retirement on 2026-10-01
Input
$1.75 per 1M text input tokens; $1.75 per 1M image input tokens
Output
$19.50 per 1M image output tokens
View model
→
High-fidelity text-to-image generation, precise image editing, hero imagery, commercial and photorealistic creative work, accurate in-image typography, and visually dense scenes requiring consistent objects, characters, materials, and spatial relationship
Type
Multimodal
Context
131K
Reasoning
7/10
Speed
4/10
Multimodal
Image input
Media output
Input
$5 per 1 million text input tokens; $8 per 1 million image input tokens
Output
$106 per 1 million image output tokens
View model
→
Fast, high-volume text-to-image generation, image editing, marketing assets, product imagery and production design
Type
Multimodal
Context
32K
Reasoning
5/10
Speed
9/10
Multimodal
Image input
Media output
View model
→
Complex mathematical reasoning, software engineering, quantitative enterprise analysis, long-context document work, and application-managed agent workflows.
Type
Reasoning
Context
256K
Reasoning
9/10
Speed
7/10
Tool use
Streaming
View model
→
Multilingual speech-to-text, captions, meeting transcription, accessibility, call analysis, content workflows, voice-agent audio understanding, and domain-specific terminology
Type
Other
Reasoning
1/10
Speed
9/10
Audio input
Status
Preview; currently accessible through Microsoft Foundry and Azure Speech
Input
$0.36 per hour of audio
View model
→
Multilingual audio transcription, meeting and contact-center records, captions, clinical notes, accessibility, media search, voice-agent evaluation, and domain-specific transcription with speaker labels and timestamps
Audio input
Input
$0.10 per hour of audio; limited-time launch pricing through December 31, 2026
Output
Included; no separate text-output charge documented
View model
→
Expressive long-form narration, audiobooks, podcasts, educational content, voice-over, accessibility, and high-fidelity branded audio
Type
Other
Reasoning
1/10
Speed
7/10
Multimodal
Audio input
Media output
Input
$22 per 1 million characters
Output
$22 per 1 million characters
View model
→
Low-latency expressive speech for voice agents, assistants, call centers, IVR systems, and interactive multilingual applications
Media output
Streaming
Input
$15 per 1 million characters
Output
$15 per 1 million characters
View model
→
Medical image and text embeddings, similarity search, multimodal retrieval, downstream classification, outlier detection, dataset curation, and healthcare AI development workflows
Multimodal
Image input
Fine-tuning
View model
→
Text-prompted segmentation of complete CT or MRI volumes, organ and lesion delineation, volumetry, annotation assistance, and medical-imaging research
Type
Other
Reasoning
1/10
Speed
5/10
Multimodal
Image input
Media output
Status
Available through Microsoft Foundry classic managed compute; research and model-development use only
View model
→
Microsoft Foundry model-router
Enterprise applications with mixed-complexity workloads, model selection automation, agent workflows, cost optimization, and configurable quality-versus-latency trade-offs.
Type
General Purpose
Context
200K
Reasoning
8/10
Speed
7/10
Multimodal
Image input
Tool use
Status
Current; latest documented version is 2025-11-18
Input
Not publicly listed as a numeric rate; Microsoft’s pricing page currently shows $- and directs users to Azure pricing for applicable deployment details.
Output
Not separately listed; billing depends on model-router and the underlying routed model pricing configuration.
View model
→
Local and private text generation, coding assistance, mathematics, reasoning, summarization, and latency-sensitive applications
Type
General Purpose
Context
4K
Reasoning
7/10
Speed
7/10
Fine-tuning
Status
Retired from Microsoft Foundry; open-weight repository remains available for independent deployment
Input
$0.00017 per 1,000 tokens historically on Azure AI; hosted price no longer applicable after retirement
Output
$0.00068 per 1,000 tokens historically on Azure AI; hosted price no longer applicable after retirement
View model
→
Long-context chat, retrieval-augmented generation, document summarization, coding, mathematics, reasoning, and self-hosted text-generation applications
Type
General Purpose
Context
131K
Reasoning
7/10
Speed
7/10
Fine-tuning
Streaming
Status
Generally available in Microsoft Foundry; open-weight checkpoint available for download and self-hosting
Input
$0.17 per 1 million input tokens in Microsoft Foundry
Output
$0.68 per 1 million output tokens in Microsoft Foundry
View model
→
Local or low-latency text generation, lightweight assistants, mathematics, coding, summarization, and resource-constrained deployments
Type
Lightweight
Context
4K
Reasoning
6/10
Speed
9/10
Fine-tuning
Streaming
Status
Retired from Microsoft Foundry on August 30, 2025
Input
$0.00013 per 1,000 input tokens (historical Azure pricing)
Output
$0.00052 per 1,000 output tokens (historical Azure pricing)
View model
→
Long-document analysis, local assistants, code and math tasks, retrieval-augmented generation, private or offline inference, and resource-constrained deployments.
Type
Lightweight
Context
131K
Reasoning
6/10
Speed
8/10
Fine-tuning
Streaming
Status
Retired from Microsoft Foundry on August 30, 2025; open-weight checkpoint remains downloadable and usable for self-hosted inference.
Input
$0.00013 per 1,000 input tokens, historical Azure Models-as-a-Service pricing; hosted offering retired
Output
$0.00052 per 1,000 output tokens, historical Azure Models-as-a-Service pricing; hosted offering retired
View model
→
Local and private text generation, instruction following, code and mathematics assistance, summarization, extraction, and resource-constrained deployments
Type
Lightweight
Context
8K
Reasoning
7/10
Speed
7/10
Fine-tuning
Streaming
Status
Retired from Microsoft Foundry on 2025-08-30; downloadable open-weight checkpoint remains available
Input
$0.00015 per 1,000 input tokens (historical Azure Models as a Service pricing; hosted service retired)
Output
$0.0006 per 1,000 output tokens (historical Azure Models as a Service pricing; hosted service retired)
View model
→
Long-context text generation, document analysis, summarization, local assistants, coding support, mathematics, and cost-sensitive self-hosted applications
Type
Lightweight
Context
131K
Reasoning
6/10
Speed
7/10
Fine-tuning
Streaming
Status
Retired from Microsoft Foundry on 2025-08-30; open-weight checkpoint remains available for self-hosted deployment
Input
USD 0.00015 per 1,000 input tokens historically on Azure Foundry; hosted Azure access is retired
Output
USD 0.0006 per 1,000 output tokens historically on Azure Foundry; hosted Azure access is retired
View model
→
Local multilingual chat, summarization, document analysis, coding assistance, long-context retrieval, and resource-constrained deployments
Type
Lightweight
Context
131K
Reasoning
6/10
Speed
8/10
Fine-tuning
Streaming
Status
Available open-weight model; also supported in Microsoft and third-party local or hosted inference environments
View model
→
Long-context multilingual assistants, coding, mathematics, reasoning, retrieval-augmented generation, and controlled local deployment.
Type
General Purpose
Context
131K
Reasoning
7/10
Speed
6/10
Fine-tuning
Streaming
Status
Retired from Azure Foundry on August 30, 2025; downloadable open-weight checkpoint remains available for self-hosted deployment.
Input
$0.16 per 1M input tokens historically on Azure; current hosted pricing unavailable after retirement.
Output
$0.64 per 1M output tokens historically on Azure; current hosted pricing unavailable after retirement.
View model
→
Lightweight image and text reasoning, OCR, charts, tables, diagrams, documents, screenshots, and multi-image comparison
Type
Multimodal
Context
131K
Reasoning
7/10
Speed
8/10
Multimodal
Image input
Fine-tuning
Status
Current and accessible; open-weight model and available through Microsoft Foundry deployment options
View model
→
Efficient local or hosted text generation, STEM reasoning, mathematics, coding assistance, technical question answering, and research on small language models
Type
General Purpose
Context
16K
Reasoning
8/10
Speed
8/10
Status
Preview in Microsoft Foundry; open-weight model available under the MIT license
View model
→
Efficient mathematical reasoning, math tutoring, automated assessment, lightweight reasoning agents, edge deployment, mobile applications, and latency-sensitive local inference
Type
Reasoning
Context
66K
Reasoning
8/10
Speed
9/10
Fine-tuning
Streaming
Status
Current open-weight model; available through Microsoft’s Hugging Face repository and documented Azure AI Foundry availability
View model
→
Efficient local or cloud text generation, multilingual applications, mathematics, coding, reasoning, retrieval-augmented generation, and edge deployment
Type
Lightweight
Context
131K
Reasoning
7/10
Speed
9/10
Tool use
Fine-tuning
Streaming
Status
Generally available
Input
$0.00075 per 1,000 tokens in Microsoft's February 2025 Azure announcement; current Foundry pricing may vary
Output
$0.0003 per 1,000 tokens in Microsoft's February 2025 Azure announcement; current Foundry pricing may vary
View model
→
Mathematical reasoning, educational tutoring, formal proof assistance, symbolic computation, local inference, edge deployment, and latency-sensitive applications
Type
Reasoning
Context
128K
Reasoning
8/10
Speed
9/10
Fine-tuning
Status
Preview; currently available through Microsoft Foundry and open-weight distribution
View model
→
Compact multimodal assistants, OCR, document and chart analysis, image question answering, speech recognition, speech translation, audio summarization, and private or local deployment
Type
Multimodal
Context
131K
Reasoning
7/10
Speed
8/10
Multimodal
Image input
Audio input
Status
Available; open-weight; listed in Microsoft Foundry and Hugging Face
Input
Not publicly listed as a current numeric price; Azure pricing page displays $-
Output
Not publicly listed as a current numeric price; Azure pricing page displays $-
View model
→
Mathematical reasoning, coding, science, logic, algorithmic problem solving, research, and resource-constrained local deployments
Type
Reasoning
Context
33K
Reasoning
8/10
Speed
7/10
Streaming
Status
Preview; open-weight and currently listed in Microsoft Foundry
Input
No official per-token price published; open-weight model for self-hosted or separately priced managed deployment
Output
No official per-token price published; open-weight model for self-hosted or separately priced managed deployment
View model
→
Mathematical reasoning, scientific problem solving, coding assistance, algorithmic tasks, local deployment, and reasoning-model research
Type
Reasoning
Context
33K
Reasoning
8/10
Speed
5/10
Status
Current downloadable open-weight model
View model
→