What is Gemini 2.5 Computer Use?
Gemini 2.5 Computer Use is a specialized vision-language model from Google for computer-control agents. A vision-language model can process both written instructions and visual input. In this case, the visual input is primarily a screenshot of a browser or software interface.
The model does not operate a computer by itself. Instead, it reasons about the visible interface and returns a response containing text and structured function calls that represent actions. A client application then executes those actions, captures the updated screen or an action result, and sends that information back to the model for the next step.
For example, an automation agent could be instructed to open a website, locate a form, click a field, type information, scroll to another section, and wait for a page to load. Gemini 2.5 Computer Use can propose the individual interface actions, while the surrounding application remains responsible for execution, permissions, error handling, and safety checks.
Model identity and current status
The canonical API identifier is gemini-2.5-computer-use-preview-10-2025. Google launched the model as a preview on October 7, 2025. Its documented status is now legacy preview: the endpoint remains documented and accessible, but newer Gemini 3.x models support computer use without requiring a separate specialized model.
This status matters when choosing a model for a new deployment. Gemini 2.5 Computer Use can still be relevant when an existing application depends on its endpoint or when a team specifically wants to evaluate its behavior. However, a new project should compare it with Google’s newer computer-use-capable models before committing to an older preview model.
Gemini 2.5 Computer Use should also be distinguished from Google’s Computer Use tool and implementation documentation. The model supplies the reasoning and action calls; the developer’s application supplies the browser, virtual machine, container, or other environment that carries out those actions.
What the model can do
The model accepts text instructions and screenshot images. Its output can include ordinary text as well as structured calls for interface actions. Supported browser-oriented actions include:
- Opening or navigating to pages.
- Clicking at interface coordinates.
- Typing text into fields.
- Scrolling through a page.
- Using keyboard combinations.
- Hovering over interface elements.
- Waiting for a page or interface state to change.
These capabilities make the model suitable for multi-step workflows in which an agent must respond to what is visibly present rather than follow a fixed sequence of selectors. This can help with websites or internal applications that lack a convenient conventional API, although visual automation can still be sensitive to layout changes and unexpected screens.
The model can be integrated through the Gemini API and Google’s Vertex AI preview tooling. Function calls are central to its computer-use design: the application must interpret the returned action, decide whether it is permitted, execute it, and report the result. A model response should therefore be treated as a proposed operation rather than an automatically trusted command.
Technical specifications and supported modalities
| Specification | Verified value |
|---|---|
| Model family | Gemini 2.5 |
| API identifier | gemini-2.5-computer-use-preview-10-2025 |
| Input modalities | Text and image, including screenshots |
| Output types | Text and UI-action function calls |
| Input token limit | 128,000 tokens |
| Maximum output | 64,000 tokens |
| Availability | Legacy preview |
| Release date | October 7, 2025 |
The model does not natively generate images, audio, video, speech, or embeddings. Its “computer-use” output is action-oriented function calling, not direct control of a user’s device. The application must provide the execution loop and decide how screenshots, action results, and confirmation requests are handled.
Pricing and cost considerations
Google’s published pricing charges Gemini 2.5 Computer Use by tokens. For prompts up to 200,000 tokens, input costs $1.25 per 1 million tokens and text output, including response and reasoning tokens, costs $10 per 1 million tokens. For prompts above 200,000 tokens, the listed prices are $2.50 per 1 million input tokens and $15 per 1 million output tokens.
| Prompt size | Input | Output, including reasoning tokens |
|---|---|---|
| Up to 200,000 tokens | $1.25 per 1 million tokens | $10 per 1 million tokens |
| More than 200,000 tokens | $2.50 per 1 million tokens | $15 per 1 million tokens |
These are model-token charges only. A real computer-use system may also incur costs for browser infrastructure, virtual machines, containers, storage, monitoring, and application-side orchestration. Repeated screenshot-and-action cycles can increase token usage, so the total workflow cost depends on how many model turns are needed to complete a task.
Reasoning, coding, and tool support
The model’s primary strength is visual task reasoning: it can connect a written objective with the visible state of an interface and propose the next UI operation. It is not positioned as a general-purpose coding model, although developers can use it within software that contains application logic and tool handlers.
The supplied evaluation records an editorial reasoning score of 7 out of 10, coding score of 5 out of 10, speed score of 6 out of 10, and cost score of 5 out of 10. These are catalog evaluations rather than Google-published benchmark results and should not be treated as official performance measurements.
Tool support is built into the model’s intended interaction pattern through function calls. The model can return actions such as click, type, scroll, key combination, hover, navigation, and wait operations. It does not independently verify that an action is safe or appropriate. Developers should validate coordinates and targets, restrict access to sensitive systems, and require confirmation for consequential operations such as purchases, account changes, data deletion, or message submission.
Implementation and safety requirements
Computer-use actions should run inside a sandboxed browser, virtual machine, or container instead of directly on an unrestricted host system. The application should maintain a clear loop: send the task and current screenshot, receive the model’s proposed action, validate it, execute it, capture the result, and continue only when appropriate.
Safety controls are particularly important because a visual agent may encounter login pages, misleading buttons, unexpected pop-ups, private information, or irreversible actions. The model’s returned function call is not a substitute for application authorization. Teams should limit available websites and tools, isolate credentials where possible, log actions, and provide a human confirmation step for high-impact operations.
Best use cases and limitations
Gemini 2.5 Computer Use is best suited to:
- Browser-based automation when a conventional API is unavailable.
- Repetitive website workflows and data-entry tasks.
- Form filling and navigation across multi-step pages.
- User-interface testing and visual regression workflows.
- Agents that need to interpret screenshots before selecting the next action.
It is less suitable when a stable, documented API can perform the same operation more reliably. Direct API calls are usually easier to validate and maintain than coordinate-based screen interaction. The model is also a weaker fit for applications requiring native image, audio, video, speech, or embedding generation, because those are not supported outputs of this endpoint.
Other limitations follow from its preview and legacy status. Interface changes can affect visual actions, the model may require several turns to complete a task, and the developer must build the execution environment rather than receiving a complete hosted computer-control service. Google’s documentation also identifies newer Gemini 3.x models as current alternatives with computer-use support.
When to choose Gemini 2.5 Computer Use
Choose this model when you specifically need Google’s Gemini 2.5 Computer Use endpoint for screenshot-based browser automation, UI testing, or an existing computer-control workflow. Its 128,000-token input capacity and structured action output provide enough room for multi-step interactions, while the dedicated design makes its purpose clear.
Choose another option when you are starting a new project and want the current computer-use direction from Google, when the task can be completed more reliably through a conventional API, or when you need direct generation of media or speech. Newer Gemini 3.x computer-use-capable models deserve evaluation for new deployments because Gemini 2.5 Computer Use is now classified as a legacy preview model.
Overall, Gemini 2.5 Computer Use is a focused bridge between visual language-model reasoning and browser automation. Its value comes from returning executable interface actions, not from acting as a general multimedia model or a replacement for secure application logic. The most dependable deployments will pair it with a restricted execution environment, explicit validation, and human approval for sensitive actions.

