What the Microsoft Foundry API is and when to use it
Microsoft Foundry is a developer platform for creating, deploying, evaluating, governing, and operating AI applications and agents on Azure. It brings together models available through Azure, Azure OpenAI-compatible APIs, Foundry project resources, Foundry Agent Service, tools, evaluations, fine-tuning, content safety, and Azure identity and security controls.
The main generative interface for new applications is the Responses API. It accepts text, image, and supported file inputs and can return generated text, structured JSON, tool calls, and streamed events. It also supports multi-turn interactions, background tasks, reasoning models, response retrieval, and continuation with previous_response_id.
There are two important ways to work with the platform:
| Use this route | Best suited to |
|---|---|
| Azure OpenAI-compatible endpoint | Applications that prioritize compatibility with OpenAI clients, embeddings, image generation, chat completions, and latency. |
| Foundry project endpoint | Applications that need Foundry-specific agents, evaluations, project resources, file search, code interpreter, web search, memory, MCP, SharePoint, WorkIQ, or Fabric IQ. |
The platform is aimed at application developers, backend engineers, data and automation teams, and organizations that need Azure governance, regional deployment, identity integration, quota management, and enterprise controls. It is less suitable when you want a simple consumer chatbot subscription, a single fixed model with uniform behavior everywhere, or a service with no Azure resource and deployment administration.
Getting access and obtaining credentials
Start by creating a Microsoft Foundry or Azure OpenAI resource in Azure and deploying a supported model. The exact model identifier depends on the deployment and Azure environment, so do not assume that a public model name is valid for your resource. Store the deployed model identifier, endpoint, and credentials as environment variables rather than putting them directly in source code.
Microsoft supports two main authentication methods for Azure OpenAI-compatible APIs:
- Microsoft Entra ID: Usually the better production choice because it works with Azure role-based access control and avoids distributing long-lived API keys.
- API key: A straightforward option for development and supported deployments. Send it in the
api-keyHTTP header and protect it like any other secret.
Authentication requirements and permissions can differ between Azure OpenAI-compatible resources and Foundry project endpoints. In production, use managed identity or another Entra-based approach where practical, restrict permissions through Azure RBAC, and keep credentials on the server rather than in browser or mobile application code.
Choosing the endpoint and model
Choose the endpoint based on the application features you need rather than treating Microsoft Foundry as one single API URL. The Azure OpenAI-compatible base URL follows this pattern:
https://YOUR-RESOURCE-NAME.openai.azure.com/openai/v1/The Foundry project endpoint follows this pattern:
https://RESOURCE-NAME.services.ai.azure.com/api/projects/PROJECT-NAMEFor a conventional application that sends prompts, receives responses, uses streaming, or needs OpenAI-compatible client behavior, begin with the Azure OpenAI-compatible Responses API. For project-scoped agents, evaluations, platform tools, or persistent agent resources, use the Foundry project endpoint and the relevant Foundry SDK.
The model request field must contain the exact model or deployment identifier configured in your resource. Availability varies by model, region, deployment type, API version, preview status, quota, and access permissions. Check those conditions before designing around a particular capability such as image input, reasoning, fine-tuning, or a tool.
Making your first request
The following example uses the current OpenAI Python SDK with the Azure OpenAI-compatible Responses API. It sends a simple text input and prints the returned text.
pip install openaiimport os
from openai import OpenAI
endpoint = os.environ["AZURE_OPENAI_ENDPOINT"].rstrip("/")
model = os.environ["AZURE_OPENAI_MODEL"]
client = OpenAI(
api_key=os.environ["AZURE_OPENAI_API_KEY"],
base_url=f"{endpoint}/openai/v1/",
)
response = client.responses.create(
model=model,
instructions="You are a concise technical assistant.",
input="Explain what an API rate limit is in two sentences.",
)
print(response.output_text)The equivalent HTTP operation is a POST request to /responses. The request includes the deployed model identifier and either a string input or a structured list of input messages and content items.
Understanding the response
A Responses API result can contain more than one output item. The SDK's output_text convenience property is useful when you only need the generated text. Applications that support tools, structured results, reasoning, or other output types should inspect the response items and handle the relevant event or output types explicitly.
Responses can be used for multi-turn interaction. One simple approach is to send the earlier response ID with a later request:
follow_up = client.responses.create(
model=model,
previous_response_id=response.id,
input="Now give one practical mitigation for this limit.",
)
print(follow_up.output_text)Stateful behavior has data-retention implications. The model inference itself is described as stateless, but stored responses, conversations, threads, files, vector stores, batch data, and agent resources may persist according to the feature and configuration you use.
How Microsoft Foundry API pricing works
Microsoft Foundry and Azure OpenAI use Azure consumption pricing rather than one universal monthly API plan. Costs vary by model, input and output tokens, image or audio usage, batch processing, deployment type, region, and service tier. Exact rates should be checked in the current Azure pricing calculator and Azure OpenAI pricing documentation.
Standard and Global Standard deployments use quota and shared-capacity controls. Provisioned Throughput Units provide dedicated capacity and are intended for workloads that need more predictable throughput and latency. This can change the cost model substantially, so compare expected request volume and latency requirements before selecting a deployment type.
API usage can also be affected by regional quota, tokens per minute, requests per minute, model selection, retries, and the amount of input context sent with each request. Monitor token usage and failed or retried requests rather than estimating cost only from the number of API calls.
Capabilities available through the API
Streaming responses
Set stream: true to receive server-sent events as the response is generated. Streaming lets an interface display partial text instead of waiting for the complete result. Your client must process the event stream and handle interruptions, errors, and incomplete output.
stream = client.responses.create(
model=model,
input="Give three short API reliability tips.",
stream=True,
)
for event in stream:
if getattr(event, "type", "") == "response.output_text.delta":
print(event.delta, end="", flush=True)
print()Tools and function calling
Tool calling lets the model request an operation that your application performs, such as looking up an account or querying an internal service. Your code remains responsible for validating arguments, authorizing the operation, executing it, and returning the result. Use the current tools convention; older functions and function_call parameter patterns should not be the basis of new integrations.
Foundry project workflows can also expose platform tools such as web search, file search, code interpreter, memory, MCP, SharePoint, WorkIQ, and Fabric IQ where supported. Tool availability depends on the endpoint, model, region, permissions, and current service status.
Files and image input
Vision-capable models can receive supported image content, including image URLs or other accepted image inputs. Files can be uploaded through the Files API and referenced by supported model or tool workflows. File search and code interpreter are associated particularly with Foundry project and agent scenarios.
Do not assume that every deployed model accepts every media type. Confirm support for the selected model and deployment, and review how uploaded files are stored, accessed, and deleted before sending sensitive information.
Structured outputs
Structured outputs use a JSON Schema supplied under text.format. With strict validation enabled, the application can request a predictable JSON shape instead of trying to extract JSON from ordinary prose.
structured = client.responses.create(
model=model,
input="Extract a person and an organization from: Ada works at Example Corp.",
text={
"format": {
"type": "json_schema",
"name": "entity_extraction",
"schema": {
"type": "object",
"properties": {
"person": {"type": "string"},
"organization": {"type": "string"}
},
"required": ["person", "organization"],
"additionalProperties": False
},
"strict": True
}
}
)
print(structured.output_text)Agents and fine-tuning
Foundry supports ephemeral agent patterns through the Responses API and persistent agent capabilities through Foundry Agent Service. Project-level agent features can connect models with tools, files, evaluations, and other managed resources. Fine-tuning is available for supported models and methods, subject to model, region, quota, and access restrictions.
SDKs, playground, and developer tools
For the Azure OpenAI-compatible Responses API, Microsoft provides current examples using the official OpenAI SDK for Python and JavaScript or TypeScript. The Python package is openai, and the JavaScript or TypeScript package is also openai. For project-level Foundry APIs, Microsoft provides Foundry SDKs, including the stable Python package azure-ai-projects, along with Foundry packages for JavaScript or TypeScript, C#, and Java.
The Microsoft Foundry portal at ai.azure.com provides model catalog discovery, deployments, agent and evaluation workflows, playground testing, configuration, and monitoring. The Azure portal remains important for resource creation, quota, identity, networking, deployment administration, and billing.
Use the SDK that matches the endpoint you selected. Avoid combining examples from older Azure SDK generations with the current OpenAI-compatible Responses API syntax.
Important limits and production considerations
Quota is assigned by region, model, subscription, and deployment type. Common controls include tokens per minute and requests per minute. Responses may include headers such as x-ratelimit-limit-requests, x-ratelimit-remaining-requests, x-ratelimit-limit-tokens, and retry-after-ms.
Production clients should:
- Implement exponential backoff for transient failures and HTTP 429 responses.
- Honor retry headers rather than retrying immediately in a tight loop.
- Avoid sudden traffic bursts and monitor remaining request and token capacity.
- Track latency, token usage, failures, content-filter results, and tool errors.
- Use Provisioned Throughput when consistent capacity and latency are more important than the flexibility of shared capacity.
- Verify regional model availability and API-version lifecycle information before deployment.
- Review retention and geographic-processing requirements before enabling stored responses, threads, files, vector stores, or agents.
Microsoft states that prompts and completions sent to models sold by Azure are not used to train, retrain, or improve the base models. This does not mean that every feature is storage-free: stateful features can store customer data in the Azure tenant. The Responses API documentation describes stored response data as retained for 30 days by default, while feature-specific, preview, regional, and zero-data-retention arrangements can differ.
When Microsoft Foundry is a good or poor choice
Microsoft Foundry is a strong fit when your application already runs on Azure or needs Azure RBAC, Entra ID, regional deployment, enterprise networking, quota controls, model governance, evaluations, and managed agent features. The Azure OpenAI-compatible endpoint also makes it practical to reuse OpenAI-oriented client code while retaining Azure deployment and administration.
It may be a poor fit when your project needs a minimal setup with no cloud-resource administration, a fixed and globally uniform model catalog, or pricing that is easy to express as one subscription fee. The platform has many endpoint, deployment, region, quota, identity, and preview-status choices. Capabilities can vary considerably across models and environments, so a successful prototype does not automatically guarantee the same behavior in production.
For a new application, a sensible starting path is to deploy a supported model, use the Azure OpenAI-compatible Responses API, authenticate with Entra ID where possible, and add streaming, structured outputs, files, or tools only after confirming support for the chosen model and deployment. Move to Foundry project endpoints and Agent Service when the application needs persistent agents, evaluations, or Foundry-specific project resources.
