What the Databricks API is and when to use it
The Databricks API is a collection of services rather than one endpoint. The general REST API manages platform resources such as jobs, compute, Unity Catalog, permissions, files, and serving endpoints. Model Serving exposes custom models, foundation models, external model providers, and agents through HTTPS endpoints. Foundation Model APIs provide access to supported hosted models without requiring you to manage the underlying deployment. Unity Gateway model services provide a governed interface for supported models.
For generative-AI applications, Databricks currently emphasizes OpenAI-compatible interfaces. Depending on the service, you can use Chat Completions, the Open Responses API, the Databricks OpenAI client, a standard OpenAI Python or JavaScript client, raw HTTP, or Databricks SDKs.
Databricks is a good fit when an application needs model access alongside enterprise data, Unity Catalog governance, deployment controls, monitoring, model routing, or cloud-based infrastructure. It is less suitable for someone who wants a simple consumer chatbot or a small standalone model endpoint with minimal administration.
Getting access and obtaining credentials
You need a Databricks workspace or another supported Databricks deployment, permission to use the relevant model service or serving endpoint, and credentials accepted by that workspace. Access can differ according to the cloud, region, workspace tier, model, endpoint configuration, and account permissions.
Databricks recommends OAuth for users and service principals, especially for automated production workloads. OAuth token federation can exchange an identity-provider JWT for a Databricks OAuth token and reduces the need to manage long-lived secrets. Personal access tokens remain supported in applicable environments, but they are a less-preferred or legacy-style option for many production scenarios.
Keep credentials outside source code. Use environment variables, a secret manager, workload identity, or OAuth token federation. The examples below use a token environment variable to keep the request easy to understand; production systems should use the authentication method appropriate for their deployment.
Choosing the appropriate API and model
Choose the API based on what you are deploying:
| Requirement | Typical Databricks path |
|---|---|
| Call a governed hosted or routed model | Unity Gateway model service through the OpenAI-compatible interface |
| Deploy a custom model, external provider, or agent | Model Serving endpoint |
| Use a supported hosted foundation model without managing deployment | Foundation Model APIs |
| Manage jobs, compute, permissions, catalogs, or workspace resources | Databricks REST API or an official SDK |
| Store or transfer documents in governed storage | Files API and Unity Catalog volumes |
Model names, supported features, regions, and lifecycle status are model-specific. Databricks can deprecate or retire models, so check current availability before hard-coding a model into a new application. For a new model-service integration, a workspace-specific base URL commonly looks like https://<workspace-url>/ai-gateway/mlflow/v1. A directly deployed serving endpoint generally uses https://<workspace-url>/serving-endpoints.
Making the first request
The following example uses the current OpenAI-compatible Chat Completions pattern. Replace the workspace URL, token, and model service with values available in your Databricks environment. The model identifier shown is an example from the documented pattern; model availability can vary.
curl -sS -X POST "https://${DATABRICKS_WORKSPACE_URL}/ai-gateway/mlflow/v1/chat/completions"
-H "Authorization: Bearer ${DATABRICKS_TOKEN}"
-H "Content-Type: application/json"
-d '{
"model": "system.ai.claude-sonnet-4-5",
"messages": [
{"role": "system", "content": "You are a concise technical assistant."},
{"role": "user", "content": "Explain Databricks Model Serving in one paragraph."}
],
"max_tokens": 256,
"temperature": 0.2
}'A directly deployed Model Serving endpoint uses the serving endpoint name as the model value and normally uses the serving-endpoints base URL. The traditional invocation route remains available at /serving-endpoints/{name}/invocations. For Responses-style requests, Databricks documents the /serving-endpoints/open-responses path and uses an input field instead of Chat Completions' messages field.
Understanding the response
Chat Completions responses follow the familiar OpenAI-compatible structure. The generated text is normally available at choices[0].message.content. A response can also contain tool-call information rather than a final answer when the model decides that an application function should be invoked.
Responses-style calls use an output array and can expose provider-independent features such as reasoning, tool calls, structured output, and image input when the selected model supports them. Do not assume that every model or endpoint exposes the same fields. The serving mode and selected provider determine the available feature set.
How Databricks pricing generally works
Databricks pricing is usage-based rather than a single universal API subscription price. Total cost depends on the cloud, region, workspace tier, product, model, deployment mode, and workload.
- Foundation Model APIs may charge according to input tokens, output tokens, and query-based quotas.
- Model Serving can use pay-per-token, priority pay-per-token, provisioned throughput, or compute-based serving.
- General platform capabilities are billed through Databricks Units and related cloud infrastructure charges.
- Custom models, agents, external model endpoints, and supporting data services can have different cost behavior.
Use the current cloud-specific Databricks price list for estimates. Avoid treating a price from one model, cloud, or serving mode as a universal API price. In production, track token usage, endpoint usage, infrastructure charges, and any enabled inference or usage tables separately.
Important capabilities
Streaming responses
Compatible chat, completion, agent, and Responses-style workloads can support streaming. Instead of waiting for the complete answer, the client receives incremental events or chunks and can display text as it arrives. The exact event structure depends on the API being used.
curl -N -sS -X POST "https://${DATABRICKS_WORKSPACE_URL}/ai-gateway/mlflow/v1/chat/completions"
-H "Authorization: Bearer ${DATABRICKS_TOKEN}"
-H "Content-Type: application/json"
-d '{
"model": "system.ai.claude-sonnet-4-5",
"messages": [{"role": "user", "content": "Give three practical Databricks API use cases."}],
"stream": true,
"max_tokens": 256
}'Function calling and tools
Function calling lets a model return a structured request to call an application-defined function. Your application, not the model, executes the function, checks its arguments, and sends the result back in a subsequent model request. This is useful for connecting a model to internal search, ticketing, databases, or business actions.
Function calling is OpenAI-compatible for Foundation Model APIs and serving endpoints that support external models. Availability and tool behavior remain model-specific. Treat tool arguments as untrusted input: validate types, permissions, allowed operations, and authorization before executing anything.
Structured outputs
Supported chat models can return JSON objects or JSON Schema-based responses. Structured output is useful when application code needs predictable fields rather than free-form prose. Model-specific restrictions apply, and documented configurations generally cannot combine structured output with streaming. Claude structured output also has additional restrictions involving tools or tool choice.
const structured = await client.chat.completions.create({
model: "system.ai.claude-sonnet-4-5",
messages: [{
role: "user",
content: "Return a short summary of Databricks Model Serving and two benefits."
}],
response_format: {
type: "json_schema",
json_schema: {
name: "serving_summary",
strict: true,
schema: {
type: "object",
properties: {
summary: { type: "string" },
benefits: { type: "array", items: { type: "string" } }
},
required: ["summary", "benefits"],
additionalProperties: false
}
}
},
max_tokens: 256
});Files and documents
The Databricks Files API supports uploading, downloading, listing, and deleting files in Unity Catalog volumes and other supported locations. It is a governed file-management API, not a universal model-file-upload endpoint.
A typical document workflow stores files in governed volumes and then implements the required parsing, chunking, indexing, retrieval, or agent integration. Some specialized product experiences have separate upload behavior. For example, certain Genie Agent file uploads are UI-only and do not support API-based uploads through the general Files API.
Agents and model deployment
Databricks supports deployed agents through Databricks Apps and Model Serving. Databricks recommends its Databricks OpenAI client for querying deployed agents because it supports native agent features and streaming. Custom models, hosted models, external model providers, and agents can be exposed through serving endpoints, subject to supported deployment modes and permissions.
Using an SDK
Because the model-service interface is OpenAI-compatible, the standard OpenAI Python client can be configured with a Databricks base URL. This example uses one current SDK generation consistently.
# Install: pip install -U openai
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["DATABRICKS_TOKEN"],
base_url=f"https://{os.environ['DATABRICKS_WORKSPACE_URL']}/ai-gateway/mlflow/v1",
)
response = client.chat.completions.create(
model="system.ai.claude-sonnet-4-5",
messages=[
{"role": "system", "content": "You are a helpful technical assistant."},
{"role": "user", "content": "What is Databricks Model Serving?"},
],
max_tokens=256,
)
print(response.choices[0].message.content)
stream = client.chat.completions.create(
model="system.ai.claude-sonnet-4-5",
messages=[{"role": "user", "content": "List three API integration tips."}],
stream=True,
max_tokens=256,
)
for chunk in stream:
text = chunk.choices[0].delta.content
if text:
print(text, end="", flush=True)
print()Databricks also provides official SDKs and tooling for platform operations, MLflow Deployments APIs, the Databricks CLI, language-specific tools for supported operations, and raw REST access for any language. Use the Databricks-specific SDKs when managing workspaces, jobs, catalogs, permissions, or other platform resources rather than forcing those tasks through a model client.
Playground and developer tools
The AI Playground provides an interactive way to test supported models, compare responses, and prototype tool-calling agents. It is available only in supported workspaces and regions. The REST API reference, model availability documentation, SDKs, CLI, and MLflow deployment tooling are useful for moving from an experiment to a controlled deployment.
Databricks development typically requires more configuration than a standalone model API. You may need to understand workspace URLs, Unity Catalog names, endpoint permissions, cloud regions, serving modes, model lifecycle status, authentication, and usage-based billing.
Limits and production considerations
Databricks applies endpoint, workspace, query, input-token, and output-token limits. Foundation Model API limits vary by workspace tier, model, deployment mode, and region. Model Serving and REST APIs also enforce endpoint and workspace limits. Requests that exceed applicable limits can receive HTTP 429 responses.
- Use bounded exponential backoff for retryable rate-limit responses.
- Set connection and request timeouts rather than allowing calls to wait indefinitely.
- Track request IDs, latency, token usage, errors, and endpoint status where available.
- Limit concurrency to the capacity of the selected endpoint and workspace.
- Check the model lifecycle policy before deployment and plan for model replacement.
- Review data residency, retention, logging, inference-table, external-provider, and contract settings before sending regulated data.
- Use OAuth, least-privilege permissions, endpoint ACLs, network controls, and secret management for production access.
Databricks states that designated AI services protect customer content under platform and contractual controls and do not use it to train generative foundation models made available to third parties. However, retention and handling depend on the service, cloud, workspace configuration, external model provider, contract, and enabled logging. Do not generalize one service's retention behavior to every endpoint.
The former Foundation Model Fine-tuning feature and the legacy databricks_genai package are end-of-life. Current training and fine-tuning work should use the supported AI Runtime and Model Training capabilities. A trained or customized model can then be registered in Unity Catalog and deployed with Model Serving when its architecture, region, quotas, and product availability are supported.
Advantages and limitations
Why use the Databricks API
- Model access can be combined with enterprise data, Unity Catalog governance, permissions, and audit controls.
- OpenAI-compatible interfaces reduce the amount of application code needed for common model calls.
- Model Serving supports hosted models, custom models, external providers, and agents.
- Developers can choose pay-per-token, provisioned-throughput, priority, or compute-based serving where available.
- REST APIs, SDKs, files, jobs, data services, and model endpoints can be managed within one broader platform.
Where it may be a poor fit
- A casual user seeking a simple chatbot may find workspace and cloud configuration unnecessary.
- Costs can be difficult to estimate because they depend on usage, cloud infrastructure, model, region, and serving mode.
- Feature availability is not uniform across models and endpoints.
- Free or trial-oriented environments and workspace tiers can have important quotas and administrative restrictions.
- Developers must understand Databricks authentication, permissions, Unity Catalog, endpoint management, and model lifecycle changes.
Databricks is strongest when the application needs governed AI integrated with data engineering, analytics, machine learning, deployment, and enterprise operations. If the main requirement is an isolated text-generation call with minimal infrastructure, a simpler standalone API may be easier to operate.
