What is computer-use-preview?
computer-use-preview is a specialized OpenAI model for computer-use automation. Instead of primarily generating a conversational answer, it analyzes a visual computer interface and produces an action that an application can carry out. A typical cycle looks like this: the application sends instructions and a screenshot, the model returns a computer action, the application executes that action in a browser or desktop environment, and a new screenshot is sent back for the next decision.
Its action vocabulary is intended for ordinary graphical-interface operations, including mouse movement, clicking, typing, scrolling, and related keyboard or pointer interactions. The model does not itself take control of a computer. The surrounding runtime must execute its response, manage permissions, collect updated observations, and decide whether an action is safe to perform.
OpenAI introduced computer-use-preview on March 11, 2025, as a research-preview model associated with its computer-use tool and Operator-related work. OpenAI currently lists the canonical model as deprecated. That status is important: the model can still be useful for understanding or maintaining an existing integration, but it is not the natural default for a new production system when a supported successor or general-purpose computer-use option is available.
How the model works in an application
computer-use-preview is documented for the Responses API. Developers provide a task in text and can provide screenshots or other image observations. The model then returns text-based API content containing structured computer-use actions. A client application translates those actions into browser or operating-system events and sends the resulting state back to the model.
This design makes the model an action planner inside a control loop rather than a standalone automation script. For example, a workflow might ask it to find a particular setting in a web application. The model can inspect the visible page, click a likely control, observe the changed screen, and continue. The application remains responsible for implementing the loop and for stopping when the task is complete or an action requires human approval.
Function calling is supported, and the model can return structured computer actions for the integration to process. This is useful when an application needs machine-readable actions instead of prose instructions. It does not mean that every returned action is guaranteed to be correct, safe, or suitable for unattended execution.
Inputs, outputs, and modalities
The verified input modalities are text and images. Screenshots are the central visual input for computer-use workflows, allowing the model to reason about buttons, fields, menus, page layouts, and other visible interface elements. The supplied research does not document native audio or video input.
The model’s output is primarily text-based API content containing computer actions. It does not directly generate images, audio, or video, so its multimodal behavior is mainly on the input side. Its distinctive output is structured action information such as mouse and keyboard operations, which the host application can execute.
| Capability | Verified status |
|---|---|
| Text input | Supported |
| Image input and screenshots | Supported |
| Audio input or output | Not documented |
| Video input or output | Not documented |
| Computer-use actions | Supported |
| Function calling | Supported |
| Streaming | Not supported |
| Structured Outputs | Not supported |
| Fine-tuning | Not supported |
Context limit, output limit, and pricing
OpenAI’s listed context window is 8,192 tokens, with a maximum output of 1,024 tokens. The context window covers the text and other information supplied as part of the model interaction. In a screenshot-driven loop, developers should account for the task instructions, prior interaction history, tool results, and current visual observations when designing prompts and state management.
The listed API prices are $3.00 per 1 million input tokens and $12.00 per 1 million output tokens. OpenAI also notes that computer-use tool calls may incur separate tool-specific fees. The effective cost of an automated task therefore depends not only on token volume but also on how many observation-and-action cycles the application performs and whether those cycles generate additional tool charges.
| Specification | Value |
|---|---|
| Provider | OpenAI |
| Release date | March 11, 2025 |
| Status | Deprecated |
| Context window | 8,192 tokens |
| Maximum output | 1,024 tokens |
| Input price | $3.00 per 1 million input tokens |
| Output price | $12.00 per 1 million output tokens |
| Primary API | Responses API |
| Knowledge cutoff | October 1, 2023 |
Strengths and practical trade-offs
The model’s main strength is specialization. It is designed around the difficult step between understanding a visual interface and expressing the next interaction in a form that software can execute. That makes it more directly applicable to browser workflows and interface testing than a text-only model that can describe what a user should click but cannot produce the corresponding computer-use action format.
It also combines screenshot interpretation with iterative action selection. This is useful for interfaces whose layout, state, or available controls change during a task. Rather than relying entirely on fixed coordinates or brittle scripts, an application can ask the model to reassess the current screen after each action.
Those benefits come with meaningful trade-offs. The model is a research preview, has a relatively small 8,192-token context window, and limits each response to 1,024 tokens. It does not support streaming, Structured Outputs, or fine-tuning. Its deprecated status adds lifecycle risk, particularly for a new system that needs a long support horizon.
The supplied research includes editorial ratings of reasoning, coding, speed, and cost, but those scores are comparative evaluations rather than OpenAI-published guarantees. The rating profile describes computer-use-preview as relatively quick and moderately priced for its category, while rating its reasoning and coding usefulness lower than a general-purpose model would typically be expected to perform. These scores should be treated as directional assessments, not formal benchmark specifications.
Reliability and safety limitations
Computer interaction can create real side effects: a mistaken click may submit a form, alter data, download a file, or trigger an external transaction. OpenAI released computer-use-preview as a research preview partly because this class of task requires caution. Launch-era system-card testing reported a 38.1% success rate on OSWorld, and OpenAI highlighted weaker reliability in non-browser environments. That result is a provider-reported evaluation claim from the launch context, not a guarantee for a particular application or workflow.
Deployments should use isolated environments and restrict the websites, accounts, and actions the model can access. Screen content should be treated as untrusted input because visible text can contain misleading instructions or content that conflicts with the user’s task. Consequential actions should require confirmation, especially purchases, account changes, communications, deletion, or data submission.
Verification is essential. An application should check the resulting page or system state instead of assuming that a returned action succeeded. Timeouts, recovery paths, permission boundaries, and human escalation are especially important for tasks that run outside a controlled browser or that can affect sensitive data.
Best use cases
- Controlled browser automation: navigating a known set of websites or completing repetitive internal workflows in an isolated session.
- Interface testing: exercising end-to-end user flows and checking how an application behaves when a user interacts through its visible interface.
- Repetitive computer tasks: handling low-consequence steps that are easy for the surrounding application to verify.
- Agent research and prototyping: evaluating how a computer-using agent can combine screenshots, decisions, actions, and feedback.
- Action-execution pipelines: applications that can translate model-generated mouse and keyboard actions into events and independently verify the outcome.
When to choose this model
Choose computer-use-preview mainly when you are maintaining an existing OpenAI computer-use integration, testing the original research-preview behavior, or building a tightly controlled prototype that specifically needs screenshot-based computer actions. Its specialization can be more relevant than a conventional conversational model when the central problem is interacting with a graphical interface rather than producing an explanation or code sample.
For a new production deployment, the deprecated status should weigh heavily against it. A currently supported computer-use-capable model is generally a better starting point when lifecycle stability, updated capabilities, or documented migration support matters. A regular general-purpose model may be more appropriate for planning, text generation, coding, or knowledge work where direct GUI actions are unnecessary. A deterministic browser automation framework is preferable when the page structure is stable and reliability matters more than flexibility.
It is not a good fit for general chat, current-information tasks, image generation, audio or video processing, or high-reliability unattended desktop control. Its October 1, 2023 knowledge cutoff also means that it should not be treated as a source of current facts; external observations and tools may provide fresh state, but they do not change the model’s underlying training cutoff.
Bottom line
computer-use-preview is best understood as an early, specialized computer-interaction model rather than a general-purpose assistant. It accepts instructions and screenshots, returns structured mouse and keyboard actions through the Responses API, and can support controlled automation and interface testing. Its limited context and output sizes, lack of streaming and fine-tuning, safety requirements, research-preview reliability, and deprecated status make careful containment and verification necessary. For existing experiments it remains technically distinctive; for new production work, a supported alternative should normally be evaluated first.

