What the Gemini API is and when to use it
The Gemini API is Google's direct developer platform for integrating Gemini models into applications. It lets software send prompts and files to a selected model, receive generated responses, and connect model reasoning to application code through tools and function calls.
Unlike the consumer Gemini applications, the API is intended for developers building their own products, services, automations, research tools, and internal workflows. It supports text generation as well as image, audio, video, and document understanding when the selected model and endpoint support those inputs.
For new projects, Google recommends the Interactions API. It provides a common request format for model calls and managed agent workflows, including streaming, structured responses, built-in tools, function calling, and server-side multi-turn state through previous_interaction_id. The older generateContent interface remains supported and can still be useful for direct stateless generation or compatibility with existing code.
How to get access and create an API key
Gemini API requests require an API key. Google AI Studio is the simplest starting point for most developers: it provides a visual playground, project and key management, model testing, usage monitoring, and agent prototyping.
- Open Google AI Studio and create or select a project.
- Create an API key for the project.
- Store the key in an environment variable such as
GEMINI_API_KEY. - Use the key from a server-side application, SDK, or REST request.
Do not place an API key in browser code, a mobile application bundle, a public repository, or client-side JavaScript that users can inspect. Use a backend service or a secret manager to keep the key private. REST requests normally send the key in the x-goog-api-key header.
Choosing the interface and model
Start by selecting the model that matches the application's task, input types, latency needs, and budget. Model availability and exact feature support can differ, so check the current model documentation before deploying.
The Interactions API is the recommended general interface for new applications. It is especially useful when an application needs conversation state, streaming events, tools, structured responses, or agent execution. Use the original generateContent API only when its stateless request format is a better fit or when maintaining an existing integration.
The model name is supplied in each request. The examples below use gemini-3.8-flash, matching the supplied current API research. Replace it with a model available to the project and suitable for the required capabilities.
Making the first request
Install the current Google GenAI SDK rather than an older legacy Gemini client library. Python projects use google-genai, while JavaScript and TypeScript projects use @google/genai. Official implementations are also available for Go and Java.
Python example
pip install -U google-genai
import os
from google import genai
client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
interaction = client.interactions.create(
model="gemini-3.8-flash",
system_instruction="You are a helpful technical assistant.",
input="Explain APIs in two sentences.",
)
print(interaction.output_text)JavaScript example
npm install @google/genai
import { GoogleGenAI } from "@google/genai";
const ai = new GoogleGenAI({
apiKey: process.env.GEMINI_API_KEY
});
const interaction = await ai.interactions.create({
model: "gemini-3.8-flash",
system_instruction: "You are a helpful technical assistant.",
input: "Explain APIs in two sentences."
});
console.log(interaction.output_text);The SDK examples use the same current client and method family throughout: GoogleGenAI or genai.Client with interactions.create. REST is also available when an application needs direct HTTP requests or does not use an official SDK.
Understanding the response
An interaction can contain output text and execution steps. For a simple request, the SDK's output_text property is the convenient way to retrieve the generated text. Applications using tools or managed agents may inspect the individual steps to understand tool calls, intermediate events, and returned content.
Do not treat generated text as automatically reliable or safe to execute. Validate data, apply authorization checks around tools, and handle errors before using a response in a business process. Structured output can make parsing more predictable, but it does not remove the need for application-level validation.
How Gemini API pricing works
The Gemini API has a free tier for eligible models and usage levels, followed by paid access for production workloads. Paid usage is generally metered according to factors such as:
- Input tokens sent to the model.
- Output tokens generated by the model.
- Cached tokens and cached-token storage.
- The selected model.
- Service tier.
- Batch processing.
- Tool usage where applicable.
Google also provides batch processing, context caching, priority service tiers, and enterprise deployments with additional controls. Prices vary by model and can change, so use the current official pricing documentation rather than embedding assumed prices in an application or budget.
Free-tier terms are also important for data governance: Google states that content from the Gemini API free tier may be used to improve products. Paid Gemini API usage is described as not being used to improve Google's products. Enterprise Gemini deployments have separate contractual and governance terms.
What developers can build with the API
Multimodal input
Gemini can process text, images, audio, video, and documents, subject to model and endpoint restrictions. This supports applications such as document analysis, image interpretation, audio or video understanding, and workflows that combine several input types.
Small inputs can be included directly in a request. For larger files or files reused across requests, use the Files API. Uploaded files can be used by supported models and are stored for 48 hours. The documented limits are up to 2 GB per file and 20 GB of file storage per project.
Streaming responses
Streaming sends incremental events while the model is generating a response instead of waiting for the complete result. It is useful for chat interfaces, progress displays, and applications where users should see output as it becomes available.
const stream = await ai.interactions.create({
model: "gemini-3.8-flash",
input: "Write a short welcome message.",
stream: true
});
for await (const event of stream) {
if (event.type === "content_delta") {
process.stdout.write(event.delta ?? "");
}
}
process.stdout.write("n");The Live API provides bidirectional real-time streaming for suitable voice and multimodal experiences. Its availability and supported input or output combinations depend on the relevant model and feature documentation.
Function calling and built-in tools
Function calling lets the model request an action defined by the application. For example, a model might request a weather lookup or database search, after which the application executes the function and sends the result back. The model does not receive permission to perform arbitrary actions automatically; the application controls which functions exist and whether each call is allowed.
Gemini also supports built-in tools such as Google Search, URL context, code execution, and file search. Tool availability depends on the model and whether a capability is in preview. Production systems should validate tool arguments, enforce authorization, set timeouts, and record tool errors separately from model output.
Structured JSON output
Structured output allows an application to provide a JSON Schema and request a response that conforms to it. This is useful when generated information must be passed to another program rather than displayed only as prose.
from pydantic import BaseModel
from google import genai
import os
class Contact(BaseModel):
name: str
email: str
client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
result = client.interactions.create(
model="gemini-3.8-flash",
input="Extract the contact from: Ada Lovelace, ada@example.com",
response_format={
"type": "text",
"mime_type": "application/json",
"schema": Contact.model_json_schema(),
},
)
contact = Contact.model_validate_json(result.output_text)
print(contact.name, contact.email)The Interactions API uses the top-level response_format field with a text format, the application/json MIME type, and a schema. The official SDKs support schema definitions through standard dictionaries, Pydantic models in Python, and JSON-Schema-compatible approaches in JavaScript. Validate the result before storing it or using it in downstream operations.
Multi-turn state and agents
Applications can continue an interaction by passing the previous interaction's identifier:
follow_up = client.interactions.create(
model="gemini-3.8-flash",
previous_interaction_id=interaction.id,
input="Now give one practical example.",
)
print(follow_up.output_text)This server-side state is convenient for multi-turn workflows, but it also affects retention and data-governance decisions. The Interactions API stores requests by default to support this stateful workflow.
Managed agents are available in public preview. They can run in isolated Linux sandbox environments, execute code, manage files, browse the web, and use configured tools. Developers can invoke predefined agents or create agents with persistent instructions, skills, files, and network rules. Preview schemas and behavior may change, so avoid making an irreversible production dependency on undocumented agent behavior.
Uploading and analyzing files
The Files API supports reusable text, image, audio, video, and document inputs, subject to model-specific restrictions. A typical workflow uploads a file and then passes its URI to an interaction:
uploaded = client.files.upload(file="document.pdf")
file_interaction = client.interactions.create(
model="gemini-3.8-flash",
input=[
{"type": "text", "text": "Summarize this document."},
{
"type": "document",
"uri": uploaded.uri,
"mime_type": uploaded.mime_type
},
],
)
print(file_interaction.output_text)Files uploaded through the Gemini Files API are stored for 48 hours. Plan workflows around that retention period and upload again when a file must be processed later.
Limits and production considerations
Limits are applied at the project level and commonly include requests per minute, input tokens per minute, and requests per day. The exact values depend on the model and billing or usage tier and can be viewed in Google AI Studio.
Batch processing has separate documented limits, including up to 100 concurrent batch requests, a 2 GB input-file limit, and a 20 GB project file-storage limit. These figures should be checked against the current documentation before capacity planning.
- Protect credentials: keep API keys on the server and rotate or revoke exposed keys.
- Handle transient failures: implement bounded retries with backoff for rate-limit and temporary service errors.
- Control spending: use project quotas, billing controls, and appropriate model or service tiers.
- Validate outputs: parse and validate structured responses before using them in application logic.
- Minimize sensitive logging: request identifiers and operational metadata are often more useful than storing complete prompts and responses.
- Plan for retention: account for Files API retention, Interactions API state, and managed-agent environment lifetimes.
- Check regional availability: model access, limits, tools, and enterprise options can vary by location and account configuration.
There is no single latency figure for the Gemini API. Response time depends on the model, prompt size, generated output, region, traffic, tools, and service tier. Applications with strict latency or throughput requirements should test their own workload and evaluate supported service tiers, quotas, regional availability, and enterprise deployment options.
Fine-tuning and data policy
Direct fine-tuning is not currently available through the Gemini API or Google AI Studio. The last supported Gemini API tuning model, Gemini 1.5 Flash-001, was deprecated in May 2025. Fine-tuning remains available through Google's enterprise platform for supported models and configurations.
Retention and training treatment depend on the API tier and feature. Free-tier content may be used to improve Google's products, while paid Gemini API usage is described as not being used for that purpose. Enterprise deployments have separate terms and governance controls. Review the applicable current service terms before sending confidential, regulated, or proprietary information.
When the Gemini API is a good or poor choice
The Gemini API is a good choice when an application needs a Google-supported multimodal platform with official SDKs, Google AI Studio, streaming, structured output, function calling, file analysis, Google Search or other built-in tools, and a unified path toward stateful interactions and managed agents. It is also a practical option for teams already using Google services or evaluating Gemini models through Google's developer tooling.
It may be a poor fit when the project requires a fixed, unchanging model interface, guaranteed deterministic answers, unrestricted access to every capability, or complete independence from Google's account, service, and data-governance ecosystem. Free-tier data-use terms, feature and model availability, preview APIs, project quotas, file retention, and changing model pricing should all be considered during evaluation.
For a new application, a sensible starting path is to create a protected API key in Google AI Studio, install the current Google GenAI SDK, make a small Interactions API request, add structured output or streaming only when needed, and then test the selected model and workload against the project's cost, latency, privacy, and reliability requirements.
