Developer platform

Yandex AI Studio API

Developer overview for Yandex AI, including API access, pricing, SDK support, endpoints, capabilities and platform policies.

API key Required
Primary API Yandex Cloud AI Studio
SDK support Official Yandex Cloud SDKs support Node.js, Go, Python, Java, and .NET. A dedicated Yandex AI Studio SDK is available, and compatible Python and JavaScript OpenAI clients can be used with the OpenAI-compatible endpoint. No PHP-specific SDK is required for
Rate limits Documented AI Studio quotas include up to 10 concurrent synchronous text generations, 10 asynchronous submission requests per second, 50 asynchronous result requests per second, 5,000 asynchronous submissions per hour, 50 tokenization requests per second,
Platform

API overview

Endpoints

API access

Base URL https://ai.api.cloud.yandex.net/v1
Primary API Yandex Cloud AI Studio
Pricing

API pricing

Pricing model Usage-based cloud billing

Prices vary by model and resource consumption, including input/output tokens, image or audio processing, batch jobs, dedicated instances, agents, and fine-tuning. Consult the current official price list for exact rates.

Developer experience

SDKs & usability

SDKs Official Yandex Cloud SDKs support Node.js, Go, Python, Java, and .NET. A dedicated Yandex AI Studio SDK is available, and compatible Python and JavaScript OpenAI clients can be used with the OpenAI-compatible endpoint. No PHP-specific SDK is required for
Ease of use Moderate. The OpenAI-compatible endpoint lowers migration effort, while native Yandex Cloud APIs require familiarity with folders, service accounts, IAM permissions, model URIs, and cloud billing.
Documentation Good and actively maintained. Documentation covers REST, gRPC, OpenAI-compatible access, Responses API, SDKs, agents, files, fine-tuning, quotas, pricing, and tutorials. Capability details can vary by model and feature status.
Latency No single platform-wide latency guarantee is published. Synchronous requests return after model processing, streaming can provide intermediate output, asynchronous mode is available for longer operations, and dedicated instances are available for workload
Features

API capabilities

✓ Streaming
✓ Function calling
✓ Assistants API
✓ File uploads
✓ Fine-tuning
✓ Web search
✓ Image input
✓ Structured outputs
✓ Playground
Feature notes

Yandex AI Studio currently provides native REST and gRPC APIs, an OpenAI-compatible API, AI Playground, official cloud SDKs, a dedicated AI Studio SDK, streaming, synchronous, asynchronous and batch modes, structured JSON responses, function calling, hosted tools including web search and MCP-related integrations, AI agents, file management, multimodal model access, embeddings, image generation, and selected fine-tuning workflows. Capability support is model-specific. Some models and features may be preview, region-restricted, permission-restricted, or unavailable through every API surface. The OpenAI-compatible base URL is https://ai.api.cloud.yandex.net/v1. Native Yandex Cloud API-key authentication uses the Api-Key scheme, while OpenAI-compatible client examples commonly pass the same credential through the client's API-key configuration. The platform's current service documentation distinguishes current Responses API functionality from older or endpoint-specific interfaces; developers should follow the API and model documentation for the selected integration.

Examples

API examples

curl --fail-with-body --silent --show-error https://ai.api.cloud.yandex.net/v1/chat/completions \
  -H "Authorization: Bearer ${YANDEX_API_KEY}" \
  -H "Content-Type: application/json" \
  -d @- <<'JSON'
{
  "model": "gpt://FOLDER_ID/MODEL_ID/latest",
  "messages": [
    {
      "role": "system",
      "content": "You are a concise technical assistant."
    },
    {
      "role": "user",
      "content": "Explain retry with exponential backoff in three sentences."
    }
  ],
  "temperature": 0.2,
  "stream": false
}
JSON

curl --fail-with-body --silent --show-error https://ai.api.cloud.yandex.net/v1/chat/completions \
  -H "Authorization: Bearer ${YANDEX_API_KEY}" \
  -H "Content-Type: application/json" \
  -d @- <<'JSON'
{
  "model": "gpt://FOLDER_ID/MODEL_ID/latest",
  "messages": [
    {
      "role": "system",
      "content": "Return only valid JSON."
    },
    {
      "role": "user",
      "content": "Return a JSON object with the fields name and priority for a database migration task."
    }
  ],
  "response_format": {
    "type": "json_object"
  }
}
JSON
# Install: pip install openai
import json
import os
from openai import OpenAI

api_key = os.environ["YANDEX_API_KEY"]
model = os.environ["YANDEX_MODEL"]

client = OpenAI(
    api_key=api_key,
    base_url="https://ai.api.cloud.yandex.net/v1",
)

try:
    response = client.chat.completions.create(
        model=model,
        messages=[
            {"role": "system", "content": "You are a concise technical assistant."},
            {"role": "user", "content": "Explain retry with exponential backoff in three sentences."},
        ],
        temperature=0.2,
    )
    print(response.choices[0].message.content)

    stream = client.chat.completions.create(
        model=model,
        messages=[
            {"role": "system", "content": "Answer clearly and briefly."},
            {"role": "user", "content": "Give a checklist for deploying an API."},
        ],
        stream=True,
    )
    for chunk in stream:
        text = chunk.choices[0].delta.content
        if text:
            print(text, end="", flush=True)
    print()
except Exception as exc:
    print(f"Yandex AI Studio request failed: {exc}")
    raise
// Install: npm install openai
import OpenAI from "openai";

const apiKey = process.env.YANDEX_API_KEY;
const model = process.env.YANDEX_MODEL;

if (!apiKey || !model) {
  throw new Error("Set YANDEX_API_KEY and YANDEX_MODEL");
}

const client = new OpenAI({
  apiKey,
  baseURL: "https://ai.api.cloud.yandex.net/v1",
});

try {
  const response = await client.chat.completions.create({
    model,
    messages: [
      { role: "system", content: "You are a concise technical assistant." },
      { role: "user", content: "Explain retry with exponential backoff in three sentences." },
    ],
    temperature: 0.2,
  });
  console.log(response.choices[0]?.message?.content ?? "");

  const stream = await client.chat.completions.create({
    model,
    messages: [
      { role: "system", content: "Answer clearly and briefly." },
      { role: "user", content: "Give a checklist for deploying an API." },
    ],
    stream: true,
  });

  for await (const chunk of stream) {
    const text = chunk.choices[0]?.delta?.content;
    if (text) process.stdout.write(text);
  }
  process.stdout.write("\n");
} catch (error) {
  console.error("Yandex AI Studio request failed:", error);
  process.exitCode = 1;
}
<?php
$apiKey = getenv('YANDEX_API_KEY');
$model = getenv('YANDEX_MODEL');

if (!$apiKey || !$model) {
    fwrite(STDERR, "Set YANDEX_API_KEY and YANDEX_MODEL\n");
    exit(1);
}

$url = 'https://ai.api.cloud.yandex.net/v1/chat/completions';
$payload = [
    'model' => $model,
    'messages' => [
        [
            'role' => 'system',
            'content' => 'You are a concise technical assistant.'
        ],
        [
            'role' => 'user',
            'content' => 'Explain retry with exponential backoff in three sentences.'
        ]
    ],
    'temperature' => 0.2,
    'stream' => false
];

$ch = curl_init($url);
curl_setopt_array($ch, [
    CURLOPT_POST => true,
    CURLOPT_HTTPHEADER => [
        'Authorization: Bearer ' . $apiKey,
        'Content-Type: application/json',
        'Accept: application/json'
    ],
    CURLOPT_POSTFIELDS => json_encode($payload, JSON_THROW_ON_ERROR),
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_TIMEOUT => 60,
]);

$body = curl_exec($ch);
$curlError = curl_error($ch);
$status = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);

if ($body === false || $curlError !== '') {
    throw new RuntimeException('Network error: ' . $curlError);
}

$data = json_decode($body, true, 512, JSON_THROW_ON_ERROR);
if ($status < 200 || $status >= 300) {
    $message = $data['error']['message'] ?? $body;
    throw new RuntimeException("HTTP $status: $message");
}

$result = $data['choices'][0]['message']['content'] ?? null;
if ($result === null) {
    throw new RuntimeException('The response did not contain message content.');
}

echo $result . PHP_EOL;
Policies

Data & usage

Data training

No universal statement that all API-submitted data is used for model training was verified in the consulted current first-party documentation. Treat API data-handling and training use as governed by the applicable Yandex Cloud terms, service documentation, region, contract, and account configuration. Avoid submitting secrets or unnecessary personal data and verify the applicable enterprise privacy terms before production use.

Data retention

Retention is service- and resource-specific. AI Studio file objects expose expiration information, and documented asynchronous text-generation results are stored on the server for up to 3 days. Uploaded files, assistant resources, batch outputs, datasets, and agent resources can have separate expiration or lifecycle settings. Applications should configure deletion or expiration where available and confirm current service-specific retention terms.

Rate limits

Documented AI Studio quotas include up to 10 concurrent synchronous text generations, 10 asynchronous submission requests per second, 50 asynchronous result requests per second, 5,000 asynchronous submissions per hour, 50 tokenization requests per second,

Developer guide

Yandex AI Studio API: Beginner-Friendly Developer Guide

Yandex Cloud AI Studio is a usage-based developer platform for building generative AI applications with Yandex and selected third-party models. It provides native REST and gRPC APIs, an OpenAI-compatible endpoint, official SDKs, AI Playground, streaming, structured JSON responses, function calling, agents, file processing, multimodal workflows, embeddings, image generation, and selected fine-tuning capabilities.
Yandex AI Studio is intended for developers who want to add text generation, multimodal processing, agents, search-related tools, file workflows, and other AI features to applications or business systems. Access requires a Yandex Cloud account, billing configuration, permissions, and authentication. Beginners can start with the OpenAI-compatible API, while applications needing Yandex Cloud-specific services can use the native REST, gRPC, or SDK interfaces.
Yandex AI Studio is Yandex Cloud's usage-based developer platform for generative AI. It offers native REST and gRPC APIs, an OpenAI-compatible endpoint, official SDKs, AI Playground, streaming, structured JSON, function calling, agents, file processing, multimodal models, embeddings, image generation, and selected fine-tuning.

What is the Yandex AI Studio API?

Yandex AI Studio is Yandex Cloud's developer platform for integrating generative AI into applications. It provides access to Yandex models such as YandexGPT and YandexART, along with selected open-source and third-party models available in the model catalog.

The platform is separate from consumer Alice subscriptions. It is designed for production applications, internal tools, automation, agents, batch workloads, and multimodal systems. Developers can use native Yandex Cloud REST and gRPC APIs, an OpenAI-compatible endpoint, official SDKs, or the AI Playground in the Yandex Cloud console.

Who should use it?

AI Studio is suitable for developers and teams that need cloud-hosted model access with Yandex Cloud identity management, usage-based billing, service accounts, quotas, and production-oriented APIs. It may be especially relevant when applications operate in Yandex Cloud or need Yandex-specific models and services.

It is less suitable for users seeking a simple consumer chatbot subscription, a universally available global service, or a single provider-wide model with identical capabilities across every endpoint. Model support, regional availability, permissions, and feature status can vary.

Getting access and obtaining credentials

Before sending requests, create or select a Yandex Cloud account and folder, configure billing, and create a service account with the permissions required by your workload. You then authenticate with an API key or IAM token.

For application-to-application access, a scoped service-account API key is generally the practical starting point. Native Yandex Cloud APIs use the Api-Key authorization scheme. OpenAI-compatible clients commonly receive the same credential through their normal API-key configuration.

Keep credentials in environment variables or a secret manager rather than placing them directly in source code. Use the narrowest permission scope available for the operation, such as yc.ai.languageModels.execute where appropriate.

Choosing an API and model

There are two main integration approaches:

  • OpenAI-compatible API: useful when an application already uses a compatible Python or JavaScript client and needs a relatively familiar chat or responses-style interface.
  • Native Yandex Cloud APIs: appropriate when you need Yandex Cloud-specific resources, REST or gRPC services, management operations, or capabilities that are not exposed identically through the compatibility layer.

The OpenAI-compatible base URL is https://ai.api.cloud.yandex.net/v1. Model identifiers commonly use Yandex Cloud URI formats such as gpt://<folder_ID>/<model_ID>/latest. Select the model from the current AI Studio catalog rather than assuming that every model supports the same context size, modalities, tools, or output limits.

Making a first request

The following example uses the OpenAI-compatible chat-completions interface with the current OpenAI Python client package. Set the API key and model identifier in environment variables first.

pip install openai
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["YANDEX_API_KEY"],
    base_url="https://ai.api.cloud.yandex.net/v1",
)

response = client.chat.completions.create(
    model=os.environ["YANDEX_MODEL"],
    messages=[
        {"role": "system", "content": "You are a concise technical assistant."},
        {"role": "user", "content": "Explain retry with exponential backoff in three sentences."},
    ],
    temperature=0.2,
)

print(response.choices[0].message.content)

The model value is not a generic name in this example. It should be replaced with a currently available model URI from your Yandex Cloud folder and account configuration.

Understanding the response

A chat-completions response contains a collection of choices. For a standard single-answer request, application code usually reads the generated text from response.choices[0].message.content. Production code should also handle empty content, refusals or errors, usage information where returned, timeouts, and non-success HTTP responses.

How pricing works

AI Studio uses usage-based cloud billing rather than a general consumer subscription. Charges depend on the selected model and resource consumption. Relevant factors can include input and output tokens, image or audio processing, batch jobs, dedicated instances, agents, and fine-tuning.

There is no single platform-wide price that represents every AI Studio request. Check the current official Yandex Cloud pricing documentation before deployment, because rates and billing structures can change independently of the API syntax. For cost control, select models deliberately, limit unnecessary context, monitor usage, and test representative workloads before estimating production spending.

What developers can build

Streaming responses

Streaming sends generated output incrementally instead of waiting for the complete answer. This is useful for chat interfaces and long responses because the user can see progress sooner. Enable streaming when the selected endpoint and model support it.

stream = client.chat.completions.create(
    model=os.environ["YANDEX_MODEL"],
    messages=[
        {"role": "user", "content": "Give a short deployment checklist."}
    ],
    stream=True,
)

for chunk in stream:
    text = chunk.choices[0].delta.content
    if text:
        print(text, end="", flush=True)
print()

Structured JSON output

AI Studio documentation describes plain text, JSON-object output, and JSON Schema output for supported Responses API configurations. Structured output is useful when an application needs fields that it can process programmatically rather than free-form prose.

For a simple JSON object, a compatible request can use a JSON response format:

response = client.chat.completions.create(
    model=os.environ["YANDEX_MODEL"],
    messages=[
        {"role": "system", "content": "Return only valid JSON."},
        {"role": "user", "content": "Return name and priority for a database migration task."},
    ],
    response_format={"type": "json_object"},
)

print(response.choices[0].message.content)

Always parse and validate the result in application code. JSON mode or schema support can vary by model and API surface, and valid JSON does not guarantee that the values satisfy your business rules.

Function calling and hosted tools

Function calling allows a model to request an application-defined operation using a name and JSON Schema parameters. Your application remains responsible for validating the arguments, executing the function, enforcing permissions, and returning the result to the model.

AI Studio also documents hosted tools such as web search, file search, and MCP-related integrations for supported configurations. Tool availability is model- and configuration-specific, so confirm support before designing a workflow around a particular tool.

Agents

AI Studio provides first-party agent capabilities, including agent tools, file search, code execution, MCP integrations, and related resources. Developers can also connect AI Studio models to external agent frameworks and serverless components such as Yandex Cloud Functions.

Files and multimodal input

The Files API supports uploading and managing resources for use cases including assistants, batch processing, fine-tuning, vision, and user data. Supported image, video, and audio workflows depend on the selected model and API configuration. File objects can have expiration information or other lifecycle rules, so applications should not assume that uploaded resources remain available indefinitely.

Fine-tuning, embeddings, and image generation

AI Studio supports embeddings and image-generation workflows, as well as LoRA-based fine-tuning for selected language, classification, and embedding models. Fine-tuning depends on model availability, dataset requirements, quotas, preview status, and additional resource costs. Fine-tuned models are usable only through the supported generation APIs and interfaces documented for that model.

Playground and SDK options

AI Playground in the Yandex Cloud management console provides an interactive way to test models and prompts before writing application code. It is useful for comparing prompts, checking whether a model supports a desired modality, and exploring basic behavior.

Official Yandex Cloud SDKs support Node.js, Go, Python, Java, and .NET. A dedicated Yandex AI Studio SDK is also available. Applications using the OpenAI-compatible endpoint can use compatible Python or JavaScript clients, while PHP applications can call the HTTPS API directly with a standard HTTP client or cURL.

Limits and production considerations

AI Studio publishes quotas for different resources rather than one universal rate limit. Documented examples include up to 10 concurrent synchronous text generations, 10 asynchronous submission requests per second, 50 asynchronous result requests per second, 5,000 asynchronous submissions per hour, and 50 tokenization requests per second. Fine-tuning documentation includes limits of up to 10 runs per day and 3 per hour. Actual limits can vary by resource and may be adjustable.

Asynchronous text-generation results are documented as being stored on the server for up to three days. Uploaded files, batch outputs, datasets, assistants, and agent resources can have different expiration or lifecycle behavior.

For production systems:

  • Set connection and request timeouts.
  • Retry transient failures with exponential backoff and jitter.
  • Respect quotas and avoid unbounded concurrency.
  • Validate structured output and tool arguments.
  • Monitor token and resource consumption.
  • Store keys in a secret manager and use scoped service accounts.
  • Define deletion or expiration policies for uploaded data.
  • Check whether a required feature is generally available, preview, region-restricted, or model-specific.
  • Do not send secrets, unnecessary personal data, or regulated information unless the applicable terms and security configuration permit it.

Advantages and limitations

Advantages

  • Native REST and gRPC APIs plus an OpenAI-compatible endpoint.
  • Official SDK coverage for several major programming languages.
  • AI Playground for testing models and prompts.
  • Support for streaming, asynchronous and batch workloads.
  • Access to structured responses, tools, agents, files, multimodal workflows, embeddings, image generation, and selected fine-tuning.
  • Yandex Cloud identity, permissions, quotas, and service-account controls.

Limitations

  • Model capabilities and API support are not uniform across the catalog.
  • Pricing is usage-based and can involve several resource categories.
  • Quotas, preview status, permissions, and regional restrictions may affect deployment.
  • Retention differs between asynchronous results, files, datasets, and agent resources.
  • The OpenAI-compatible interface reduces migration effort but does not guarantee complete compatibility with every OpenAI feature or client behavior.
  • English-language documentation and international availability may be less comprehensive than those of some major US-based competitors.

When is Yandex AI Studio a good or poor choice?

AI Studio is a good choice when you need Yandex Cloud-hosted model APIs, Yandex models, cloud IAM, an OpenAI-compatible migration path, or a combination of text, multimodal, agent, file, and fine-tuning workflows. It is also a practical option for teams already operating applications and billing inside Yandex Cloud.

It may be a poor choice when your application requires guaranteed global availability, identical model behavior across regions, fully transparent model specifications, a simple fixed subscription, or a broad ecosystem of consumer custom agents. Before committing, verify the target model, region, quota, tool support, retention behavior, and current pricing for your exact workload.

A practical starting checklist

  1. Create a Yandex Cloud folder and configure billing.
  2. Create a service account with the minimum required AI Studio permissions.
  3. Create an API key and store it securely.
  4. Choose a current model identifier from the AI Studio catalog.
  5. Test a basic request in AI Playground or through the OpenAI-compatible endpoint.
  6. Add streaming, structured output, files, tools, or agents only after confirming support for the selected model.
  7. Implement validation, timeouts, retries, quota handling, monitoring, and data-retention controls before production use.
Sources 13