Gemini 2.5

Gemini 2.5 Computer Use

by Google DeepMind · Legacy preview; currently documented and accessible, with newer Gemini 3.x models recommended for new computer-use applications

Gemini 2.5 Computer Use is Google’s legacy preview model for browser automation and computer-control agents. It accepts text and screenshots, returns UI-action function calls, supports a 128,000-token input limit and 64,000-token maximum output, and uses standard token-based Gemini pricing.

Text Actions Reasoning Coding
Gemini 2.5 Computer Use is a specialized Google model for agents that need to understand and operate graphical user interfaces. Rather than directly controlling a person’s computer, it examines screenshots and returns structured actions that an application can execute in a browser or other controlled environment. The model is useful for browser automation, data entry, form filling, repetitive web workflows, and interface testing. It remains documented and accessible as a preview endpoint, but Google recommends evaluating newer Gemini 3.x models for new computer-use projects.
Outputs

What Gemini 2.5 Computer Use can produce

Text Actions
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Streaming Structured output
Model profile

Performance characteristics

7/10 Reasoning
5/10 Coding
6/10 Speed
5/10 Cost efficiency
Specifications

Technical details

Model family Gemini 2.5
Model type Multimodal
Context window 128K tokens
Maximum output 64K tokens
Release date 2025-10-07
Status Legacy preview; currently documented and accessible, with newer Gemini 3.x models recommended for new computer-use applications
Knowledge cutoff notes

Google's model documentation specifies the model limits and supported data types but does not publish a separate knowledge-cutoff date for this exact Computer Use preview model.

Model notes

Canonical API identifier: gemini-2.5-computer-use-preview-10-2025. The model accepts text and screenshots and returns text plus function calls representing UI actions. Client applications must execute those actions and provide screenshots or results in a controlled environment. Google now labels it a legacy preview model because Gemini 3.x models support computer use without a separate specialized model. Action output is marked positive because the model natively returns executable UI-control function calls intended to operate an external browser or interface. No separate official knowledge-cutoff date was identified.

Cost

Model pricing

Input $1.25 per 1M tokens for prompts up to 200K tokens; $2.50 per 1M tokens for prompts over 200K tokens
Output $10.00 per 1M tokens for prompts up to 200K tokens; $15.00 per 1M tokens for prompts over 200K tokens
Model guide

Gemini 2.5 Computer Use: Google’s Legacy Browser-Automation Model

Gemini 2.5 Computer Use is Google’s specialized preview model for browser automation and computer-control agents. It interprets text instructions and screenshots, then returns text and executable UI-action function calls for activities such as clicking, typing, navigation, scrolling, hovering, and waiting. It offers a 128,000-token input limit and 64,000-token maximum output, but Google now classifies it as a legacy preview model because newer Gemini 3.x models support computer use.

What is Gemini 2.5 Computer Use?

Gemini 2.5 Computer Use is a specialized vision-language model from Google for computer-control agents. A vision-language model can process both written instructions and visual input. In this case, the visual input is primarily a screenshot of a browser or software interface.

The model does not operate a computer by itself. Instead, it reasons about the visible interface and returns a response containing text and structured function calls that represent actions. A client application then executes those actions, captures the updated screen or an action result, and sends that information back to the model for the next step.

For example, an automation agent could be instructed to open a website, locate a form, click a field, type information, scroll to another section, and wait for a page to load. Gemini 2.5 Computer Use can propose the individual interface actions, while the surrounding application remains responsible for execution, permissions, error handling, and safety checks.

Model identity and current status

The canonical API identifier is gemini-2.5-computer-use-preview-10-2025. Google launched the model as a preview on October 7, 2025. Its documented status is now legacy preview: the endpoint remains documented and accessible, but newer Gemini 3.x models support computer use without requiring a separate specialized model.

This status matters when choosing a model for a new deployment. Gemini 2.5 Computer Use can still be relevant when an existing application depends on its endpoint or when a team specifically wants to evaluate its behavior. However, a new project should compare it with Google’s newer computer-use-capable models before committing to an older preview model.

Gemini 2.5 Computer Use should also be distinguished from Google’s Computer Use tool and implementation documentation. The model supplies the reasoning and action calls; the developer’s application supplies the browser, virtual machine, container, or other environment that carries out those actions.

What the model can do

The model accepts text instructions and screenshot images. Its output can include ordinary text as well as structured calls for interface actions. Supported browser-oriented actions include:

  • Opening or navigating to pages.
  • Clicking at interface coordinates.
  • Typing text into fields.
  • Scrolling through a page.
  • Using keyboard combinations.
  • Hovering over interface elements.
  • Waiting for a page or interface state to change.

These capabilities make the model suitable for multi-step workflows in which an agent must respond to what is visibly present rather than follow a fixed sequence of selectors. This can help with websites or internal applications that lack a convenient conventional API, although visual automation can still be sensitive to layout changes and unexpected screens.

The model can be integrated through the Gemini API and Google’s Vertex AI preview tooling. Function calls are central to its computer-use design: the application must interpret the returned action, decide whether it is permitted, execute it, and report the result. A model response should therefore be treated as a proposed operation rather than an automatically trusted command.

Technical specifications and supported modalities

SpecificationVerified value
Model familyGemini 2.5
API identifiergemini-2.5-computer-use-preview-10-2025
Input modalitiesText and image, including screenshots
Output typesText and UI-action function calls
Input token limit128,000 tokens
Maximum output64,000 tokens
AvailabilityLegacy preview
Release dateOctober 7, 2025

The model does not natively generate images, audio, video, speech, or embeddings. Its “computer-use” output is action-oriented function calling, not direct control of a user’s device. The application must provide the execution loop and decide how screenshots, action results, and confirmation requests are handled.

Pricing and cost considerations

Google’s published pricing charges Gemini 2.5 Computer Use by tokens. For prompts up to 200,000 tokens, input costs $1.25 per 1 million tokens and text output, including response and reasoning tokens, costs $10 per 1 million tokens. For prompts above 200,000 tokens, the listed prices are $2.50 per 1 million input tokens and $15 per 1 million output tokens.

Prompt sizeInputOutput, including reasoning tokens
Up to 200,000 tokens$1.25 per 1 million tokens$10 per 1 million tokens
More than 200,000 tokens$2.50 per 1 million tokens$15 per 1 million tokens

These are model-token charges only. A real computer-use system may also incur costs for browser infrastructure, virtual machines, containers, storage, monitoring, and application-side orchestration. Repeated screenshot-and-action cycles can increase token usage, so the total workflow cost depends on how many model turns are needed to complete a task.

Reasoning, coding, and tool support

The model’s primary strength is visual task reasoning: it can connect a written objective with the visible state of an interface and propose the next UI operation. It is not positioned as a general-purpose coding model, although developers can use it within software that contains application logic and tool handlers.

The supplied evaluation records an editorial reasoning score of 7 out of 10, coding score of 5 out of 10, speed score of 6 out of 10, and cost score of 5 out of 10. These are catalog evaluations rather than Google-published benchmark results and should not be treated as official performance measurements.

Tool support is built into the model’s intended interaction pattern through function calls. The model can return actions such as click, type, scroll, key combination, hover, navigation, and wait operations. It does not independently verify that an action is safe or appropriate. Developers should validate coordinates and targets, restrict access to sensitive systems, and require confirmation for consequential operations such as purchases, account changes, data deletion, or message submission.

Implementation and safety requirements

Computer-use actions should run inside a sandboxed browser, virtual machine, or container instead of directly on an unrestricted host system. The application should maintain a clear loop: send the task and current screenshot, receive the model’s proposed action, validate it, execute it, capture the result, and continue only when appropriate.

Safety controls are particularly important because a visual agent may encounter login pages, misleading buttons, unexpected pop-ups, private information, or irreversible actions. The model’s returned function call is not a substitute for application authorization. Teams should limit available websites and tools, isolate credentials where possible, log actions, and provide a human confirmation step for high-impact operations.

Best use cases and limitations

Gemini 2.5 Computer Use is best suited to:

  • Browser-based automation when a conventional API is unavailable.
  • Repetitive website workflows and data-entry tasks.
  • Form filling and navigation across multi-step pages.
  • User-interface testing and visual regression workflows.
  • Agents that need to interpret screenshots before selecting the next action.

It is less suitable when a stable, documented API can perform the same operation more reliably. Direct API calls are usually easier to validate and maintain than coordinate-based screen interaction. The model is also a weaker fit for applications requiring native image, audio, video, speech, or embedding generation, because those are not supported outputs of this endpoint.

Other limitations follow from its preview and legacy status. Interface changes can affect visual actions, the model may require several turns to complete a task, and the developer must build the execution environment rather than receiving a complete hosted computer-control service. Google’s documentation also identifies newer Gemini 3.x models as current alternatives with computer-use support.

When to choose Gemini 2.5 Computer Use

Choose this model when you specifically need Google’s Gemini 2.5 Computer Use endpoint for screenshot-based browser automation, UI testing, or an existing computer-control workflow. Its 128,000-token input capacity and structured action output provide enough room for multi-step interactions, while the dedicated design makes its purpose clear.

Choose another option when you are starting a new project and want the current computer-use direction from Google, when the task can be completed more reliably through a conventional API, or when you need direct generation of media or speech. Newer Gemini 3.x computer-use-capable models deserve evaluation for new deployments because Gemini 2.5 Computer Use is now classified as a legacy preview model.

Overall, Gemini 2.5 Computer Use is a focused bridge between visual language-model reasoning and browser automation. Its value comes from returning executable interface actions, not from acting as a general multimedia model or a replacement for secure application logic. The most dependable deployments will pair it with a restricted execution environment, explicit validation, and human approval for sensitive actions.


Answers to Frequently Asked Questions

What safety measures are required when using Gemini 2.5 Computer Use?
Computer-use actions should run in a sandboxed browser, virtual machine, or container. Developers should validate returned actions, restrict websites and tools, isolate credentials, log activity, and require human confirmation for sensitive operations such as purchases, account changes, data deletion, or message submission.
How much does Gemini 2.5 Computer Use cost?
For prompts up to 200,000 tokens, pricing is $1.25 per 1 million input tokens and $10 per 1 million output tokens, including reasoning tokens. For prompts above 200,000 tokens, pricing is $2.50 per 1 million input tokens and $15 per 1 million output tokens. Browser infrastructure, virtual machines, containers, monitoring, and orchestration may create additional costs.
How does Gemini 2.5 Computer Use work with browser automation?
The application sends the model a task and a current screenshot. Gemini 2.5 Computer Use proposes an interface action through a structured function call, and the application validates and executes it. The application then captures the updated screen or action result and sends it back for the next step.
What is Gemini 2.5 Computer Use?
Gemini 2.5 Computer Use is a Google vision-language model for computer-control agents. It interprets text instructions and screenshots, then returns text and structured function calls for actions such as clicking, typing, scrolling, navigating, hovering, and waiting. A separate application must execute those actions.
What is the API identifier and current status of Gemini 2.5 Computer Use?
The canonical API identifier is gemini-2.5-computer-use-preview-10-2025. Google released it as a preview on October 7, 2025, and it is now classified as a legacy preview model. Newer Gemini 3.x models provide computer-use capabilities for current deployments.


Sources 7
Provider

About Google DeepMind