What is UI-TARS-1.5-7B?
UI-TARS-1.5-7B is an open-weight multimodal agent model from ByteDance Seed. Its specialty is computer use: interpreting a screenshot or other visual interface, connecting that interface to a natural-language instruction, and selecting an action such as clicking a coordinate, entering text, pressing a key, or scrolling.
Unlike a general-purpose chat model, UI-TARS-1.5-7B is designed around the interaction loop between perception and action. A typical workflow might provide the model with a screenshot and the instruction “open the settings menu and enable dark mode.” The model analyzes the visible interface, identifies a likely target, reasons about the next step, and returns an action representation. A separate runtime then performs the action and supplies a new screenshot for the next turn.
The model was released on April 17, 2025, as the publicly downloadable 7B-class checkpoint in the UI-TARS-1.5 family. The model card describes the checkpoint as approximately 8 billion parameters even though its product name uses the 7B designation. The larger UI-TARS-1.5 system is the broader research model; the 7B release is positioned primarily for general computer-use tasks and is not presented as an equivalent replacement for the larger system in game-oriented scenarios.
How the model understands and acts on interfaces
UI-TARS-1.5-7B combines visual understanding with action-oriented reasoning. ByteDance Seed describes the UI-TARS-1.5 family as bringing visual perception, reasoning, memory, and action modeling into a unified agent architecture. For the released 7B checkpoint, the practical capabilities supported by the supplied documentation include:
- Understanding screenshots and graphical user interfaces.
- Following instructions that combine text and images.
- Locating or grounding interface elements and predicting coordinates.
- Planning multi-step computer-use tasks.
- Generating mouse and keyboard actions in a structured textual representation.
- Supporting browser and desktop interaction when connected to a suitable execution environment.
The model uses a think-then-act style: it can produce intermediate reasoning before selecting an action. Its primary output is still text. It does not natively return an image, audio, video, or other direct non-text output. “Action output” means that its text can describe or encode GUI operations; it does not mean that the model independently sends operating-system events.
Technical specifications and modalities
| Specification | Verified information |
|---|---|
| Provider | ByteDance Seed |
| Release date | April 17, 2025 |
| Model family | UI-TARS-1.5 |
| Checkpoint | ByteDance-Seed/UI-TARS-1.5-7B |
| Model size | Approximately 8B parameters according to the model card; the model name uses 7B |
| Context length | 128,000 positions in the published configuration |
| Input | Text and images |
| Output | Text containing reasoning and GUI action representations |
| License | Apache 2.0 |
| Hosted API pricing | No official per-token price verified for this checkpoint |
| Maximum output tokens | Not verified |
The 128,000-position context is useful for workflows that need to retain a long sequence of instructions, screenshots, observations, and previous actions. It should not be interpreted as a guarantee that every deployment will support the same effective visual history: image resolution, serving configuration, precision, and the surrounding agent framework can affect memory use and practical limits.
The available research does not verify an official hosted API, maximum output-token limit, knowledge-cutoff date, fine-tuning service, caching feature, batch API, or separate JSON-mode capability for this exact checkpoint. Its structured action representations should therefore not be confused with a separately documented provider JSON mode.
Reported performance and reasoning behavior
ByteDance Seed reports results for both the larger UI-TARS-1.5 system and the UI-TARS-1.5-7B comparison. The published comparison lists 42.5 on OSWorld and 61.6 on ScreenSpotPro for UI-TARS-1.5, while the 7B comparison reports 27.5 on OSWorld and 49.6 on ScreenSpotPro. These are provider-reported benchmark figures, and results in a real deployment can vary with screenshot capture, action parsing, coordinate conversion, environment state, and task design.
The model's reasoning is practical rather than a general claim of broad intellectual ability. It is intended to connect visual observations to the next computer action. That makes intermediate reasoning useful for diagnosing why an agent selected a particular button or coordinate, but it can also expose unreliable assumptions when the interface is ambiguous or changes between steps.
The supplied research rates the model's reasoning capability highly for its intended agent role, while coding capability is more limited and not the model's primary purpose. UI-TARS-1.5-7B can be used in automation workflows that involve text entry or developer tools, but it should not automatically be selected as a dedicated coding model.
Deployment and the missing execution layer
UI-TARS-1.5-7B is distributed through Hugging Face and can be deployed with compatible Transformers tooling or inference servers such as vLLM and SGLang. The model documentation also describes OpenAI-compatible local serving endpoints. These references describe ways to serve the model; they do not establish a ByteDance-hosted commercial API with published token pricing.
Actual computer control requires additional components. At minimum, an application needs to capture the current screen or browser state, send the relevant observation and instruction to the model, parse the returned action, convert coordinates when necessary, execute the mouse or keyboard event, and collect the next observation. Tools such as a browser driver, desktop automation library, or custom agent runtime may fill these roles, but no specific native tool or function-calling capability is verified for this checkpoint.
Hardware requirements are deployment-dependent. Precision, quantization, batch size, image resolution, and the chosen serving framework all affect memory consumption and throughput. Because the weights are self-hosted, cost is determined by the operator's hardware and infrastructure rather than by a documented per-input or per-output token rate. This can be attractive for experimentation and controlled workloads, but it transfers setup, scaling, monitoring, and security responsibilities to the user.
Strengths and limitations
UI-TARS-1.5-7B's clearest strength is specialization. It is built to map visual interface state to concrete actions, making it more directly relevant to GUI-agent research than a text-only language model. Open weights and the Apache 2.0 license also make it suitable for independent deployment, benchmarking, and experimentation without depending on an official hosted endpoint.
The model's 128,000-position configuration provides room for long multi-step interaction histories, and its text-and-image input supports screenshot-based workflows. Its smaller release can also be a practical entry point for users who want to test the UI-TARS approach without deploying the larger UI-TARS-1.5 research model.
There are important limitations. The checkpoint does not itself operate a computer, so a reliable action-execution layer is mandatory. It may misidentify controls, misunderstand an interface, hallucinate visual details, or choose an inefficient action, particularly when the screen is cluttered, ambiguous, unfamiliar, or changing quickly. Any autonomous workflow should restrict permissions and add confirmation steps around authentication, financial actions, sensitive files, destructive operations, and protected websites.
The 7B release is also not the strongest option in the UI-TARS family for game-oriented scenarios. ByteDance Seed reports that the larger UI-TARS-1.5 system has a significant advantage in game-related tasks. In addition, the lack of verified hosted pricing, service-level guarantees, maximum output limits, and provider-managed execution may make this model less convenient for production teams seeking a turnkey API.
When to choose UI-TARS-1.5-7B
Choose UI-TARS-1.5-7B when the main problem is understanding and acting on visual interfaces and you are prepared to manage the surrounding runtime. It is a reasonable candidate for:
- Research into computer-use and GUI agents.
- Browser or desktop automation prototypes.
- Screenshot-based interface grounding.
- Open-weight experimentation with visual action models.
- Benchmarking models that predict coordinates and interface operations.
Its open-weight deployment is especially relevant when an organization wants local control over inference or needs to adapt the execution environment rather than rely on a provider-hosted computer-use product.
Another option may be more appropriate when the priority is general conversation, dedicated code generation, a managed API with predictable per-token billing, or native end-to-end computer control. The larger UI-TARS-1.5 system is the better-supported comparison when game-oriented performance or the broader research model is the deciding factor, although the supplied information does not provide a direct commercial price comparison. A conventional automation system may also be preferable for stable, deterministic workflows where fixed selectors and explicit rules are safer than model-generated actions.
Pricing, availability, and practical trade-offs
No official hosted per-token price was verified for UI-TARS-1.5-7B. The model is available as an open-weight Hugging Face checkpoint under Apache 2.0, so users generally evaluate cost through local or independently managed inference infrastructure. That does not make the model cost-free: hardware, storage, engineering, monitoring, and execution infrastructure still contribute to the total cost.
This creates a clear trade-off. Compared with a hosted model, self-hosting can provide greater control and avoid dependence on an unpublished provider API price, but it requires more technical work and may deliver lower convenience or variable throughput. Compared with a small deterministic automation script, UI-TARS-1.5-7B can handle more visually variable interfaces, but model inference introduces latency and the possibility of incorrect actions.
Overall, UI-TARS-1.5-7B is best understood as a foundation for computer-use agents rather than a finished desktop assistant. Its value lies in open deployment, visual grounding, and action-oriented reasoning; its main costs are the need for an execution wrapper, careful safety controls, and user-managed infrastructure.

