Computer-Using Agent

computer-use-preview

by OpenAI · Deprecated

OpenAI’s computer-use-preview is a deprecated research-preview model for screenshot-based browser and desktop automation. It accepts text and images, returns structured mouse and keyboard actions through the Responses API, and supports function calling. Its 8,192-token context window, 1,024-token maximum output, limited modalities, launch-era reliability results, and lifecycle status make it most suitable for controlled experiments and existing integrations rather than new high-stakes production systems.

Text Actions Reasoning Coding
computer-use-preview was designed for applications that need an AI model to look at a computer interface and decide what to do next. It can inspect screenshots, choose actions such as clicking, typing, scrolling, or moving the pointer, and return those actions for an application to execute. OpenAI introduced it as a research preview on March 11, 2025, but now labels the model deprecated, so its main relevance is to existing integrations, experimentation, and evaluation of computer-using agents.
Outputs

What computer-use-preview can produce

Text Actions
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Batch API
Model profile

Performance characteristics

6/10 Reasoning
4/10 Coding
7/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family Computer-Using Agent
Model type Other
Context window 8K tokens
Maximum output 1K tokens
Knowledge cutoff 2023-10-01
Release date 2025-03-11
Status Deprecated
Knowledge cutoff notes

The official model page lists October 1, 2023 as the knowledge cutoff. This cutoff describes the model’s underlying training knowledge and is not changed by screenshots, external tool results, or browsing integrations.

Model notes

OpenAI introduced the model as a research preview on March 11, 2025. It is specialized for the computer-use tool and is documented for the Responses API. The model accepts text and screenshots and can return structured mouse and keyboard actions for an application runtime to execute. OpenAI currently labels the canonical model deprecated. The dated snapshot computer-use-preview-2025-03-11 was shut down on July 23, 2026, but the canonical alias has no separately documented shutdown date in the current deprecation material. The model has an October 1, 2023 knowledge cutoff. OpenAI reported a 38.1% OSWorld success rate in launch-era system-card testing and recommends strong isolation, confirmation, permission controls, and outcome verification.

Cost

Model pricing

Input $3.00 per 1 million input tokens; tool-specific computer-use calls may incur separate fees
Output $12.00 per 1 million output tokens
Model guide

computer-use-preview: OpenAI’s Specialized Model for Browser and GUI Automation

OpenAI’s computer-use-preview is a deprecated research-preview model built to interpret screenshots and return computer-interaction actions through the Responses API. It accepts text and image inputs, supports function calling and structured mouse and keyboard actions, and is intended for controlled browser automation, interface testing, and repetitive computer workflows rather than general conversation or new high-reliability production deployments.

What is computer-use-preview?

computer-use-preview is a specialized OpenAI model for computer-use automation. Instead of primarily generating a conversational answer, it analyzes a visual computer interface and produces an action that an application can carry out. A typical cycle looks like this: the application sends instructions and a screenshot, the model returns a computer action, the application executes that action in a browser or desktop environment, and a new screenshot is sent back for the next decision.

Its action vocabulary is intended for ordinary graphical-interface operations, including mouse movement, clicking, typing, scrolling, and related keyboard or pointer interactions. The model does not itself take control of a computer. The surrounding runtime must execute its response, manage permissions, collect updated observations, and decide whether an action is safe to perform.

OpenAI introduced computer-use-preview on March 11, 2025, as a research-preview model associated with its computer-use tool and Operator-related work. OpenAI currently lists the canonical model as deprecated. That status is important: the model can still be useful for understanding or maintaining an existing integration, but it is not the natural default for a new production system when a supported successor or general-purpose computer-use option is available.

How the model works in an application

computer-use-preview is documented for the Responses API. Developers provide a task in text and can provide screenshots or other image observations. The model then returns text-based API content containing structured computer-use actions. A client application translates those actions into browser or operating-system events and sends the resulting state back to the model.

This design makes the model an action planner inside a control loop rather than a standalone automation script. For example, a workflow might ask it to find a particular setting in a web application. The model can inspect the visible page, click a likely control, observe the changed screen, and continue. The application remains responsible for implementing the loop and for stopping when the task is complete or an action requires human approval.

Function calling is supported, and the model can return structured computer actions for the integration to process. This is useful when an application needs machine-readable actions instead of prose instructions. It does not mean that every returned action is guaranteed to be correct, safe, or suitable for unattended execution.

Inputs, outputs, and modalities

The verified input modalities are text and images. Screenshots are the central visual input for computer-use workflows, allowing the model to reason about buttons, fields, menus, page layouts, and other visible interface elements. The supplied research does not document native audio or video input.

The model’s output is primarily text-based API content containing computer actions. It does not directly generate images, audio, or video, so its multimodal behavior is mainly on the input side. Its distinctive output is structured action information such as mouse and keyboard operations, which the host application can execute.

CapabilityVerified status
Text inputSupported
Image input and screenshotsSupported
Audio input or outputNot documented
Video input or outputNot documented
Computer-use actionsSupported
Function callingSupported
StreamingNot supported
Structured OutputsNot supported
Fine-tuningNot supported

Context limit, output limit, and pricing

OpenAI’s listed context window is 8,192 tokens, with a maximum output of 1,024 tokens. The context window covers the text and other information supplied as part of the model interaction. In a screenshot-driven loop, developers should account for the task instructions, prior interaction history, tool results, and current visual observations when designing prompts and state management.

The listed API prices are $3.00 per 1 million input tokens and $12.00 per 1 million output tokens. OpenAI also notes that computer-use tool calls may incur separate tool-specific fees. The effective cost of an automated task therefore depends not only on token volume but also on how many observation-and-action cycles the application performs and whether those cycles generate additional tool charges.

SpecificationValue
ProviderOpenAI
Release dateMarch 11, 2025
StatusDeprecated
Context window8,192 tokens
Maximum output1,024 tokens
Input price$3.00 per 1 million input tokens
Output price$12.00 per 1 million output tokens
Primary APIResponses API
Knowledge cutoffOctober 1, 2023

Strengths and practical trade-offs

The model’s main strength is specialization. It is designed around the difficult step between understanding a visual interface and expressing the next interaction in a form that software can execute. That makes it more directly applicable to browser workflows and interface testing than a text-only model that can describe what a user should click but cannot produce the corresponding computer-use action format.

It also combines screenshot interpretation with iterative action selection. This is useful for interfaces whose layout, state, or available controls change during a task. Rather than relying entirely on fixed coordinates or brittle scripts, an application can ask the model to reassess the current screen after each action.

Those benefits come with meaningful trade-offs. The model is a research preview, has a relatively small 8,192-token context window, and limits each response to 1,024 tokens. It does not support streaming, Structured Outputs, or fine-tuning. Its deprecated status adds lifecycle risk, particularly for a new system that needs a long support horizon.

The supplied research includes editorial ratings of reasoning, coding, speed, and cost, but those scores are comparative evaluations rather than OpenAI-published guarantees. The rating profile describes computer-use-preview as relatively quick and moderately priced for its category, while rating its reasoning and coding usefulness lower than a general-purpose model would typically be expected to perform. These scores should be treated as directional assessments, not formal benchmark specifications.

Reliability and safety limitations

Computer interaction can create real side effects: a mistaken click may submit a form, alter data, download a file, or trigger an external transaction. OpenAI released computer-use-preview as a research preview partly because this class of task requires caution. Launch-era system-card testing reported a 38.1% success rate on OSWorld, and OpenAI highlighted weaker reliability in non-browser environments. That result is a provider-reported evaluation claim from the launch context, not a guarantee for a particular application or workflow.

Deployments should use isolated environments and restrict the websites, accounts, and actions the model can access. Screen content should be treated as untrusted input because visible text can contain misleading instructions or content that conflicts with the user’s task. Consequential actions should require confirmation, especially purchases, account changes, communications, deletion, or data submission.

Verification is essential. An application should check the resulting page or system state instead of assuming that a returned action succeeded. Timeouts, recovery paths, permission boundaries, and human escalation are especially important for tasks that run outside a controlled browser or that can affect sensitive data.

Best use cases

  • Controlled browser automation: navigating a known set of websites or completing repetitive internal workflows in an isolated session.
  • Interface testing: exercising end-to-end user flows and checking how an application behaves when a user interacts through its visible interface.
  • Repetitive computer tasks: handling low-consequence steps that are easy for the surrounding application to verify.
  • Agent research and prototyping: evaluating how a computer-using agent can combine screenshots, decisions, actions, and feedback.
  • Action-execution pipelines: applications that can translate model-generated mouse and keyboard actions into events and independently verify the outcome.

When to choose this model

Choose computer-use-preview mainly when you are maintaining an existing OpenAI computer-use integration, testing the original research-preview behavior, or building a tightly controlled prototype that specifically needs screenshot-based computer actions. Its specialization can be more relevant than a conventional conversational model when the central problem is interacting with a graphical interface rather than producing an explanation or code sample.

For a new production deployment, the deprecated status should weigh heavily against it. A currently supported computer-use-capable model is generally a better starting point when lifecycle stability, updated capabilities, or documented migration support matters. A regular general-purpose model may be more appropriate for planning, text generation, coding, or knowledge work where direct GUI actions are unnecessary. A deterministic browser automation framework is preferable when the page structure is stable and reliability matters more than flexibility.

It is not a good fit for general chat, current-information tasks, image generation, audio or video processing, or high-reliability unattended desktop control. Its October 1, 2023 knowledge cutoff also means that it should not be treated as a source of current facts; external observations and tools may provide fresh state, but they do not change the model’s underlying training cutoff.

Bottom line

computer-use-preview is best understood as an early, specialized computer-interaction model rather than a general-purpose assistant. It accepts instructions and screenshots, returns structured mouse and keyboard actions through the Responses API, and can support controlled automation and interface testing. Its limited context and output sizes, lack of streaming and fine-tuning, safety requirements, research-preview reliability, and deprecated status make careful containment and verification necessary. For existing experiments it remains technically distinctive; for new production work, a supported alternative should normally be evaluated first.


Answers to Frequently Asked Questions

Is computer-use-preview suitable for production automation?
It can be useful for controlled browser automation, interface testing, prototypes, and low-consequence repetitive tasks. However, its deprecated status, reported reliability limitations, and potential for harmful side effects make a supported successor preferable for new production systems, with isolation, permission controls, verification, and human approval for consequential actions.
How much does computer-use-preview cost?
The listed API price is $3.00 per 1 million input tokens and $12.00 per 1 million output tokens. Computer-use tool calls may involve separate fees, so total task costs also depend on the number of screenshot-and-action cycles and tool charges.
What are the main limitations of computer-use-preview?
The model has an 8,192-token context window and a 1,024-token maximum output. It does not support streaming, Structured Outputs, or fine-tuning, and native audio and video modalities are not documented. It also has research-preview reliability limitations and is currently listed as deprecated.
What is computer-use-preview?
computer-use-preview is a specialized OpenAI model for browser and GUI automation. It analyzes text instructions and screenshots, then returns structured computer actions such as mouse movement, clicking, typing, and scrolling for a host application to execute.
How does computer-use-preview work with the Responses API?
An application sends the model a task and visual observations such as screenshots. The model returns computer-use actions, the application executes them in a browser or desktop environment, and updated screenshots are sent back so the model can choose the next action.


Sources 5
Provider

About OpenAI