Developer platform

Gemini API

Developer overview for Google DeepMind, including API access, pricing, SDK support, endpoints, capabilities and platform policies.

API key Required
Primary API Interactions API
SDK support Official Google GenAI SDKs for Python, JavaScript/TypeScript, Go, and Java. REST is also supported. Google recommends migrating from legacy Gemini SDKs to the Google GenAI SDK.
Rate limits Project-level limits commonly include requests per minute, input tokens per minute, and requests per day. Limits vary by model and usage tier and are viewable in Google AI Studio. Batch requests have separate quotas, including up to 100 concurrent batch r
Platform

API overview

Endpoints

API access

Base URL https://generativelanguage.googleapis.com
Primary API Interactions API
Pricing

API pricing

Pricing model Free tier plus prepaid and pay-as-you-go paid tiers; usage is primarily metered by model, input tokens, output tokens, cached tokens, storage duration, service tier, batch mode, and tool usage.

Free access is available for eligible models and limits. Paid access provides higher limits, advanced models, context caching, batch processing, and production features. Current prices vary by model and are listed on the official pricing page.

Developer experience

SDKs & usability

SDKs Official Google GenAI SDKs for Python, JavaScript/TypeScript, Go, and Java. REST is also supported. Google recommends migrating from legacy Gemini SDKs to the Google GenAI SDK.
Ease of use High. Google provides AI Studio, API-key creation, quickstarts, REST examples, and official production-ready SDKs. The recommended Interactions API offers a unified model and agent interface.
Documentation Strong and extensive, with current quickstarts, API reference pages, migration guidance, SDK documentation, capability guides, pricing, rate limits, files, agents, and privacy documentation.
Latency No universal latency guarantee for the direct Gemini API. Latency varies by model, prompt and output size, region, traffic, tools, and service tier. The API exposes standard, priority, and other model-dependent service options where supported.
Features

API capabilities

✓ Streaming
✓ Function calling
✓ Assistants API
✓ File uploads
✓ Web search
✓ Image input
✓ Structured outputs
✓ Playground
Feature notes

As of September 26, 2026, Google documents the Interactions API as the recommended interface for new Gemini API projects and states that it is generally available. The original generateContent API remains supported but is considered legacy for new projects. The Interactions API supports model calls, multimodal content, streaming, server-side conversation state through previous_interaction_id, structured output through response_format, function calling, built-in tools, background execution, and managed agents. Managed agents, including Antigravity-based agents and Deep Research agents, are currently preview capabilities. The Files API supports text, images, audio, video, and documents, with model- and endpoint-specific restrictions. Direct Gemini API fine-tuning is currently unavailable, although tuning remains available through Gemini Enterprise Agent Platform for supported configurations. Image input and structured output are supported by the developer platform, but exact support depends on the selected model and endpoint.

Examples

API examples

# Basic Interactions API request
curl -sS -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
  -H "x-goog-api-key: ${GEMINI_API_KEY}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-3.8-flash",
    "system_instruction": "You are a concise technical assistant.",
    "input": "Explain server-side streaming in one paragraph."
  }' | jq -r '.steps[]?.content[]?.text // empty'

# Structured JSON output
curl -sS -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
  -H "x-goog-api-key: ${GEMINI_API_KEY}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-3.8-flash",
    "input": "Extract the product and year from: Google released Gemini in 2023.",
    "response_format": {
      "type": "text",
      "mime_type": "application/json",
      "schema": {
        "type": "object",
        "properties": {
          "product": {"type": "string"},
          "year": {"type": "integer"}
        },
        "required": ["product", "year"]
      }
    }
  }' | jq -r '.steps[]?.content[]?.text // empty'

# Streaming request
curl -N -sS -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
  -H "x-goog-api-key: ${GEMINI_API_KEY}" \
  -H "Content-Type: application/json" \
  -d '{"model":"gemini-3.8-flash","input":"Write a short welcome message.","stream":true}'
# Install: pip install -U google-genai pydantic
import os
from google import genai
from pydantic import BaseModel

client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])

# Basic request with system instruction
interaction = client.interactions.create(
    model="gemini-3.8-flash",
    system_instruction="You are a helpful technical assistant.",
    input="Explain APIs in two sentences.",
)
print(interaction.output_text)

# Multi-turn conversation using server-side interaction state
follow_up = client.interactions.create(
    model="gemini-3.8-flash",
    previous_interaction_id=interaction.id,
    input="Now give one practical example.",
)
print(follow_up.output_text)

# Streaming
stream = client.interactions.create(
    model="gemini-3.8-flash",
    input="List three benefits of structured output.",
    stream=True,
)
for event in stream:
    if getattr(event, "type", None) == "content_delta":
        print(getattr(event, "delta", ""), end="", flush=True)
print()

# Structured JSON output
class Contact(BaseModel):
    name: str
    email: str

structured = client.interactions.create(
    model="gemini-3.8-flash",
    input="Extract the contact from: Ada Lovelace, ada@example.com",
    response_format={
        "type": "text",
        "mime_type": "application/json",
        "schema": Contact.model_json_schema(),
    },
)
contact = Contact.model_validate_json(structured.output_text)
print(contact.name, contact.email)

# File input
uploaded = client.files.upload(file="document.pdf")
file_interaction = client.interactions.create(
    model="gemini-3.8-flash",
    input=[
        {"type": "text", "text": "Summarize this document."},
        {"type": "document", "uri": uploaded.uri, "mime_type": uploaded.mime_type},
    ],
)
print(file_interaction.output_text)
// Install: npm install @google/genai
import { GoogleGenAI } from "@google/genai";

const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });

async function main() {
  try {
    const interaction = await ai.interactions.create({
      model: "gemini-3.8-flash",
      system_instruction: "You are a helpful technical assistant.",
      input: "Explain APIs in two sentences.",
    });
    console.log(interaction.output_text);

    const followUp = await ai.interactions.create({
      model: "gemini-3.8-flash",
      previous_interaction_id: interaction.id,
      input: "Now give one practical example.",
    });
    console.log(followUp.output_text);

    const stream = await ai.interactions.create({
      model: "gemini-3.8-flash",
      input: "Write a short welcome message.",
      stream: true,
    });
    for await (const event of stream) {
      if (event.type === "content_delta") process.stdout.write(event.delta ?? "");
    }
    process.stdout.write("\n");

    const structured = await ai.interactions.create({
      model: "gemini-3.8-flash",
      input: "Return the status as JSON for a completed job.",
      response_format: {
        type: "text",
        mime_type: "application/json",
        schema: {
          type: "object",
          properties: {
            status: { type: "string" },
            completed: { type: "boolean" }
          },
          required: ["status", "completed"]
        }
      }
    });
    console.log(JSON.parse(structured.output_text));
  } catch (error) {
    console.error(error);
    process.exitCode = 1;
  }
}

main();
<?php
$apiKey = getenv('GEMINI_API_KEY');
if (!$apiKey) {
    throw new RuntimeException('GEMINI_API_KEY is not set');
}

$url = 'https://generativelanguage.googleapis.com/v1beta/interactions';
$payload = [
    'model' => 'gemini-3.8-flash',
    'system_instruction' => 'You are a helpful technical assistant.',
    'input' => 'Explain APIs in two sentences.',
];

$ch = curl_init($url);
curl_setopt_array($ch, [
    CURLOPT_POST => true,
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_HTTPHEADER => [
        'x-goog-api-key: ' . $apiKey,
        'Content-Type: application/json',
    ],
    CURLOPT_POSTFIELDS => json_encode($payload, JSON_THROW_ON_ERROR),
    CURLOPT_TIMEOUT => 60,
]);

$responseBody = curl_exec($ch);
if ($responseBody === false) {
    throw new RuntimeException('cURL error: ' . curl_error($ch));
}
$statusCode = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);

$data = json_decode($responseBody, true, 512, JSON_THROW_ON_ERROR);
if ($statusCode < 200 || $statusCode >= 300) {
    $message = $data['error']['message'] ?? 'Gemini API request failed';
    throw new RuntimeException($message, $statusCode);
}

$text = '';
foreach (($data['steps'] ?? []) as $step) {
    foreach (($step['content'] ?? []) as $content) {
        if (($content['type'] ?? null) === 'text') {
            $text .= $content['text'] ?? '';
        }
    }
}

echo $text . PHP_EOL;
?>
Policies

Data & usage

Data training

For the Gemini API free tier, Google states that content may be used to improve Google's products. Paid Gemini API usage is described as not being used to improve Google's products. Enterprise Gemini deployments on Gemini Enterprise Agent Platform have separate contractual controls, and Google states that customer data is not used to train or fine-tune managed models without prior permission or instruction, subject to the applicable service terms.

Data retention

Gemini Files API uploads are stored for 48 hours and cannot generally be downloaded by users after upload, although generated files may have separate download behavior. The Interactions API stores requests by default to support previous_interaction_id and server-side state. Managed-agent environments may persist between interactions and are permanently deleted after seven days of inactivity. Enterprise Gemini deployments provide different retention, residency, and zero-data-retention controls subject to configuration and applicable Google Cloud terms.

Rate limits

Project-level limits commonly include requests per minute, input tokens per minute, and requests per day. Limits vary by model and usage tier and are viewable in Google AI Studio. Batch requests have separate quotas, including up to 100 concurrent batch r

Developer guide

Google Gemini API Guide: Getting Started with the Interactions API

The Google Gemini API is a developer platform for building applications with Gemini models. The current recommended interface for new projects is the Interactions API, which supports multimodal input, streaming, structured JSON responses, function calling, built-in tools, file inputs, conversation state, and managed agents. This guide explains access, authentication, SDKs, requests, pricing, capabilities, limits, and production considerations.
Google's Gemini API gives developers direct access to Gemini models through Google AI Studio, REST, and official Google GenAI SDKs. New applications should generally use the Interactions API because it brings model calls, multimodal content, streaming, tools, structured output, multi-turn state, and agent workflows into one interface. The original generateContent API remains supported for existing applications and direct stateless generation, but it is not the preferred starting point for new projects.
The Gemini API provides direct developer access to Gemini models through Google AI Studio, REST, and official Google GenAI SDKs. The recommended Interactions API supports multimodal inputs, streaming, structured JSON, function calling, built-in tools, file analysis, conversation state, and preview managed agents. The guide covers authentication, first requests, pricing, limits, privacy, and production planning.

What the Gemini API is and when to use it

The Gemini API is Google's direct developer platform for integrating Gemini models into applications. It lets software send prompts and files to a selected model, receive generated responses, and connect model reasoning to application code through tools and function calls.

Unlike the consumer Gemini applications, the API is intended for developers building their own products, services, automations, research tools, and internal workflows. It supports text generation as well as image, audio, video, and document understanding when the selected model and endpoint support those inputs.

For new projects, Google recommends the Interactions API. It provides a common request format for model calls and managed agent workflows, including streaming, structured responses, built-in tools, function calling, and server-side multi-turn state through previous_interaction_id. The older generateContent interface remains supported and can still be useful for direct stateless generation or compatibility with existing code.

How to get access and create an API key

Gemini API requests require an API key. Google AI Studio is the simplest starting point for most developers: it provides a visual playground, project and key management, model testing, usage monitoring, and agent prototyping.

  1. Open Google AI Studio and create or select a project.
  2. Create an API key for the project.
  3. Store the key in an environment variable such as GEMINI_API_KEY.
  4. Use the key from a server-side application, SDK, or REST request.

Do not place an API key in browser code, a mobile application bundle, a public repository, or client-side JavaScript that users can inspect. Use a backend service or a secret manager to keep the key private. REST requests normally send the key in the x-goog-api-key header.

Choosing the interface and model

Start by selecting the model that matches the application's task, input types, latency needs, and budget. Model availability and exact feature support can differ, so check the current model documentation before deploying.

The Interactions API is the recommended general interface for new applications. It is especially useful when an application needs conversation state, streaming events, tools, structured responses, or agent execution. Use the original generateContent API only when its stateless request format is a better fit or when maintaining an existing integration.

The model name is supplied in each request. The examples below use gemini-3.8-flash, matching the supplied current API research. Replace it with a model available to the project and suitable for the required capabilities.

Making the first request

Install the current Google GenAI SDK rather than an older legacy Gemini client library. Python projects use google-genai, while JavaScript and TypeScript projects use @google/genai. Official implementations are also available for Go and Java.

Python example

pip install -U google-genai

import os
from google import genai

client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])

interaction = client.interactions.create(
    model="gemini-3.8-flash",
    system_instruction="You are a helpful technical assistant.",
    input="Explain APIs in two sentences.",
)

print(interaction.output_text)

JavaScript example

npm install @google/genai

import { GoogleGenAI } from "@google/genai";

const ai = new GoogleGenAI({
  apiKey: process.env.GEMINI_API_KEY
});

const interaction = await ai.interactions.create({
  model: "gemini-3.8-flash",
  system_instruction: "You are a helpful technical assistant.",
  input: "Explain APIs in two sentences."
});

console.log(interaction.output_text);

The SDK examples use the same current client and method family throughout: GoogleGenAI or genai.Client with interactions.create. REST is also available when an application needs direct HTTP requests or does not use an official SDK.

Understanding the response

An interaction can contain output text and execution steps. For a simple request, the SDK's output_text property is the convenient way to retrieve the generated text. Applications using tools or managed agents may inspect the individual steps to understand tool calls, intermediate events, and returned content.

Do not treat generated text as automatically reliable or safe to execute. Validate data, apply authorization checks around tools, and handle errors before using a response in a business process. Structured output can make parsing more predictable, but it does not remove the need for application-level validation.

How Gemini API pricing works

The Gemini API has a free tier for eligible models and usage levels, followed by paid access for production workloads. Paid usage is generally metered according to factors such as:

  • Input tokens sent to the model.
  • Output tokens generated by the model.
  • Cached tokens and cached-token storage.
  • The selected model.
  • Service tier.
  • Batch processing.
  • Tool usage where applicable.

Google also provides batch processing, context caching, priority service tiers, and enterprise deployments with additional controls. Prices vary by model and can change, so use the current official pricing documentation rather than embedding assumed prices in an application or budget.

Free-tier terms are also important for data governance: Google states that content from the Gemini API free tier may be used to improve products. Paid Gemini API usage is described as not being used to improve Google's products. Enterprise Gemini deployments have separate contractual and governance terms.

What developers can build with the API

Multimodal input

Gemini can process text, images, audio, video, and documents, subject to model and endpoint restrictions. This supports applications such as document analysis, image interpretation, audio or video understanding, and workflows that combine several input types.

Small inputs can be included directly in a request. For larger files or files reused across requests, use the Files API. Uploaded files can be used by supported models and are stored for 48 hours. The documented limits are up to 2 GB per file and 20 GB of file storage per project.

Streaming responses

Streaming sends incremental events while the model is generating a response instead of waiting for the complete result. It is useful for chat interfaces, progress displays, and applications where users should see output as it becomes available.

const stream = await ai.interactions.create({
  model: "gemini-3.8-flash",
  input: "Write a short welcome message.",
  stream: true
});

for await (const event of stream) {
  if (event.type === "content_delta") {
    process.stdout.write(event.delta ?? "");
  }
}
process.stdout.write("n");

The Live API provides bidirectional real-time streaming for suitable voice and multimodal experiences. Its availability and supported input or output combinations depend on the relevant model and feature documentation.

Function calling and built-in tools

Function calling lets the model request an action defined by the application. For example, a model might request a weather lookup or database search, after which the application executes the function and sends the result back. The model does not receive permission to perform arbitrary actions automatically; the application controls which functions exist and whether each call is allowed.

Gemini also supports built-in tools such as Google Search, URL context, code execution, and file search. Tool availability depends on the model and whether a capability is in preview. Production systems should validate tool arguments, enforce authorization, set timeouts, and record tool errors separately from model output.

Structured JSON output

Structured output allows an application to provide a JSON Schema and request a response that conforms to it. This is useful when generated information must be passed to another program rather than displayed only as prose.

from pydantic import BaseModel
from google import genai
import os

class Contact(BaseModel):
    name: str
    email: str

client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])

result = client.interactions.create(
    model="gemini-3.8-flash",
    input="Extract the contact from: Ada Lovelace, ada@example.com",
    response_format={
        "type": "text",
        "mime_type": "application/json",
        "schema": Contact.model_json_schema(),
    },
)

contact = Contact.model_validate_json(result.output_text)
print(contact.name, contact.email)

The Interactions API uses the top-level response_format field with a text format, the application/json MIME type, and a schema. The official SDKs support schema definitions through standard dictionaries, Pydantic models in Python, and JSON-Schema-compatible approaches in JavaScript. Validate the result before storing it or using it in downstream operations.

Multi-turn state and agents

Applications can continue an interaction by passing the previous interaction's identifier:

follow_up = client.interactions.create(
    model="gemini-3.8-flash",
    previous_interaction_id=interaction.id,
    input="Now give one practical example.",
)

print(follow_up.output_text)

This server-side state is convenient for multi-turn workflows, but it also affects retention and data-governance decisions. The Interactions API stores requests by default to support this stateful workflow.

Managed agents are available in public preview. They can run in isolated Linux sandbox environments, execute code, manage files, browse the web, and use configured tools. Developers can invoke predefined agents or create agents with persistent instructions, skills, files, and network rules. Preview schemas and behavior may change, so avoid making an irreversible production dependency on undocumented agent behavior.

Uploading and analyzing files

The Files API supports reusable text, image, audio, video, and document inputs, subject to model-specific restrictions. A typical workflow uploads a file and then passes its URI to an interaction:

uploaded = client.files.upload(file="document.pdf")

file_interaction = client.interactions.create(
    model="gemini-3.8-flash",
    input=[
        {"type": "text", "text": "Summarize this document."},
        {
            "type": "document",
            "uri": uploaded.uri,
            "mime_type": uploaded.mime_type
        },
    ],
)

print(file_interaction.output_text)

Files uploaded through the Gemini Files API are stored for 48 hours. Plan workflows around that retention period and upload again when a file must be processed later.

Limits and production considerations

Limits are applied at the project level and commonly include requests per minute, input tokens per minute, and requests per day. The exact values depend on the model and billing or usage tier and can be viewed in Google AI Studio.

Batch processing has separate documented limits, including up to 100 concurrent batch requests, a 2 GB input-file limit, and a 20 GB project file-storage limit. These figures should be checked against the current documentation before capacity planning.

  • Protect credentials: keep API keys on the server and rotate or revoke exposed keys.
  • Handle transient failures: implement bounded retries with backoff for rate-limit and temporary service errors.
  • Control spending: use project quotas, billing controls, and appropriate model or service tiers.
  • Validate outputs: parse and validate structured responses before using them in application logic.
  • Minimize sensitive logging: request identifiers and operational metadata are often more useful than storing complete prompts and responses.
  • Plan for retention: account for Files API retention, Interactions API state, and managed-agent environment lifetimes.
  • Check regional availability: model access, limits, tools, and enterprise options can vary by location and account configuration.

There is no single latency figure for the Gemini API. Response time depends on the model, prompt size, generated output, region, traffic, tools, and service tier. Applications with strict latency or throughput requirements should test their own workload and evaluate supported service tiers, quotas, regional availability, and enterprise deployment options.

Fine-tuning and data policy

Direct fine-tuning is not currently available through the Gemini API or Google AI Studio. The last supported Gemini API tuning model, Gemini 1.5 Flash-001, was deprecated in May 2025. Fine-tuning remains available through Google's enterprise platform for supported models and configurations.

Retention and training treatment depend on the API tier and feature. Free-tier content may be used to improve Google's products, while paid Gemini API usage is described as not being used for that purpose. Enterprise deployments have separate terms and governance controls. Review the applicable current service terms before sending confidential, regulated, or proprietary information.

When the Gemini API is a good or poor choice

The Gemini API is a good choice when an application needs a Google-supported multimodal platform with official SDKs, Google AI Studio, streaming, structured output, function calling, file analysis, Google Search or other built-in tools, and a unified path toward stateful interactions and managed agents. It is also a practical option for teams already using Google services or evaluating Gemini models through Google's developer tooling.

It may be a poor fit when the project requires a fixed, unchanging model interface, guaranteed deterministic answers, unrestricted access to every capability, or complete independence from Google's account, service, and data-governance ecosystem. Free-tier data-use terms, feature and model availability, preview APIs, project quotas, file retention, and changing model pricing should all be considered during evaluation.

For a new application, a sensible starting path is to create a protected API key in Google AI Studio, install the current Google GenAI SDK, make a small Interactions API request, add structured output or streaming only when needed, and then test the selected model and workload against the project's cost, latency, privacy, and reliability requirements.

Sources 16