UI-TARS-1.5

UI-TARS-1.5-7B

by ByteDance Seed · Available open-weight research release; superseded by UI-TARS-2 for the provider's newer UI-TARS development line

An open-weight 7B-class vision-language agent from ByteDance Seed for screenshot understanding, GUI grounding, computer-use research, and browser or desktop automation with an external execution layer.

Text Actions Reasoning Coding
UI-TARS-1.5-7B is the publicly available 7-billion-parameter variant of ByteDance Seed's UI-TARS-1.5 model family. Released on April 17, 2025, it accepts text and images, interprets graphical interfaces, and generates text-based reasoning and action instructions for tasks such as clicking, typing, scrolling, browser navigation, and desktop interaction. The model is distributed as open weights under the Apache 2.0 license, but it does not control a computer by itself: an external parser and execution environment are required to turn its responses into actual interface events.
Outputs

What UI-TARS-1.5-7B can produce

Text Actions
Inputs

What it can understand

Text Images Multimodal input
Model profile

Performance characteristics

8/10 Reasoning
4/10 Coding
5/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family UI-TARS-1.5
Model type Multimodal
Context window 128K tokens
Release date 2025-04-17
Status Available open-weight research release; superseded by UI-TARS-2 for the provider's newer UI-TARS development line
Knowledge cutoff notes

No authoritative knowledge-cutoff date was published for this exact open-weight checkpoint.

Model notes

The canonical downloadable checkpoint is ByteDance-Seed/UI-TARS-1.5-7B. ByteDance Seed describes UI-TARS-1.5 as a multimodal agent that reasons before acting and generates GUI actions. The public 7B release is primarily optimized for general computer use and is not specifically optimized for game scenarios. The Hugging Face model card lists an approximately 8B-parameter checkpoint despite the 7B model name. The model is distributed under Apache 2.0 and can be deployed with Transformers, vLLM, SGLang, Docker, and compatible inference infrastructure. Actual computer interaction requires an external parser and execution environment. No official hosted ByteDance API price, maximum output-token limit, knowledge-cutoff date, fine-tuning service, caching feature, batch API, or separate JSON-mode capability was verified for this exact checkpoint.

Model guide

UI-TARS-1.5-7B: An Open-Weight Model for Computer Use and GUI Automation

UI-TARS-1.5-7B is ByteDance Seed's open-weight vision-language agent for understanding screenshots, grounding interface elements, reasoning through computer-use tasks, and producing structured mouse and keyboard actions. It is designed for self-hosted GUI-agent research and automation prototypes rather than general chat or a hosted API service.

What is UI-TARS-1.5-7B?

UI-TARS-1.5-7B is an open-weight multimodal agent model from ByteDance Seed. Its specialty is computer use: interpreting a screenshot or other visual interface, connecting that interface to a natural-language instruction, and selecting an action such as clicking a coordinate, entering text, pressing a key, or scrolling.

Unlike a general-purpose chat model, UI-TARS-1.5-7B is designed around the interaction loop between perception and action. A typical workflow might provide the model with a screenshot and the instruction “open the settings menu and enable dark mode.” The model analyzes the visible interface, identifies a likely target, reasons about the next step, and returns an action representation. A separate runtime then performs the action and supplies a new screenshot for the next turn.

The model was released on April 17, 2025, as the publicly downloadable 7B-class checkpoint in the UI-TARS-1.5 family. The model card describes the checkpoint as approximately 8 billion parameters even though its product name uses the 7B designation. The larger UI-TARS-1.5 system is the broader research model; the 7B release is positioned primarily for general computer-use tasks and is not presented as an equivalent replacement for the larger system in game-oriented scenarios.

How the model understands and acts on interfaces

UI-TARS-1.5-7B combines visual understanding with action-oriented reasoning. ByteDance Seed describes the UI-TARS-1.5 family as bringing visual perception, reasoning, memory, and action modeling into a unified agent architecture. For the released 7B checkpoint, the practical capabilities supported by the supplied documentation include:

  • Understanding screenshots and graphical user interfaces.
  • Following instructions that combine text and images.
  • Locating or grounding interface elements and predicting coordinates.
  • Planning multi-step computer-use tasks.
  • Generating mouse and keyboard actions in a structured textual representation.
  • Supporting browser and desktop interaction when connected to a suitable execution environment.

The model uses a think-then-act style: it can produce intermediate reasoning before selecting an action. Its primary output is still text. It does not natively return an image, audio, video, or other direct non-text output. “Action output” means that its text can describe or encode GUI operations; it does not mean that the model independently sends operating-system events.

Technical specifications and modalities

SpecificationVerified information
ProviderByteDance Seed
Release dateApril 17, 2025
Model familyUI-TARS-1.5
CheckpointByteDance-Seed/UI-TARS-1.5-7B
Model sizeApproximately 8B parameters according to the model card; the model name uses 7B
Context length128,000 positions in the published configuration
InputText and images
OutputText containing reasoning and GUI action representations
LicenseApache 2.0
Hosted API pricingNo official per-token price verified for this checkpoint
Maximum output tokensNot verified

The 128,000-position context is useful for workflows that need to retain a long sequence of instructions, screenshots, observations, and previous actions. It should not be interpreted as a guarantee that every deployment will support the same effective visual history: image resolution, serving configuration, precision, and the surrounding agent framework can affect memory use and practical limits.

The available research does not verify an official hosted API, maximum output-token limit, knowledge-cutoff date, fine-tuning service, caching feature, batch API, or separate JSON-mode capability for this exact checkpoint. Its structured action representations should therefore not be confused with a separately documented provider JSON mode.

Reported performance and reasoning behavior

ByteDance Seed reports results for both the larger UI-TARS-1.5 system and the UI-TARS-1.5-7B comparison. The published comparison lists 42.5 on OSWorld and 61.6 on ScreenSpotPro for UI-TARS-1.5, while the 7B comparison reports 27.5 on OSWorld and 49.6 on ScreenSpotPro. These are provider-reported benchmark figures, and results in a real deployment can vary with screenshot capture, action parsing, coordinate conversion, environment state, and task design.

The model's reasoning is practical rather than a general claim of broad intellectual ability. It is intended to connect visual observations to the next computer action. That makes intermediate reasoning useful for diagnosing why an agent selected a particular button or coordinate, but it can also expose unreliable assumptions when the interface is ambiguous or changes between steps.

The supplied research rates the model's reasoning capability highly for its intended agent role, while coding capability is more limited and not the model's primary purpose. UI-TARS-1.5-7B can be used in automation workflows that involve text entry or developer tools, but it should not automatically be selected as a dedicated coding model.

Deployment and the missing execution layer

UI-TARS-1.5-7B is distributed through Hugging Face and can be deployed with compatible Transformers tooling or inference servers such as vLLM and SGLang. The model documentation also describes OpenAI-compatible local serving endpoints. These references describe ways to serve the model; they do not establish a ByteDance-hosted commercial API with published token pricing.

Actual computer control requires additional components. At minimum, an application needs to capture the current screen or browser state, send the relevant observation and instruction to the model, parse the returned action, convert coordinates when necessary, execute the mouse or keyboard event, and collect the next observation. Tools such as a browser driver, desktop automation library, or custom agent runtime may fill these roles, but no specific native tool or function-calling capability is verified for this checkpoint.

Hardware requirements are deployment-dependent. Precision, quantization, batch size, image resolution, and the chosen serving framework all affect memory consumption and throughput. Because the weights are self-hosted, cost is determined by the operator's hardware and infrastructure rather than by a documented per-input or per-output token rate. This can be attractive for experimentation and controlled workloads, but it transfers setup, scaling, monitoring, and security responsibilities to the user.

Strengths and limitations

UI-TARS-1.5-7B's clearest strength is specialization. It is built to map visual interface state to concrete actions, making it more directly relevant to GUI-agent research than a text-only language model. Open weights and the Apache 2.0 license also make it suitable for independent deployment, benchmarking, and experimentation without depending on an official hosted endpoint.

The model's 128,000-position configuration provides room for long multi-step interaction histories, and its text-and-image input supports screenshot-based workflows. Its smaller release can also be a practical entry point for users who want to test the UI-TARS approach without deploying the larger UI-TARS-1.5 research model.

There are important limitations. The checkpoint does not itself operate a computer, so a reliable action-execution layer is mandatory. It may misidentify controls, misunderstand an interface, hallucinate visual details, or choose an inefficient action, particularly when the screen is cluttered, ambiguous, unfamiliar, or changing quickly. Any autonomous workflow should restrict permissions and add confirmation steps around authentication, financial actions, sensitive files, destructive operations, and protected websites.

The 7B release is also not the strongest option in the UI-TARS family for game-oriented scenarios. ByteDance Seed reports that the larger UI-TARS-1.5 system has a significant advantage in game-related tasks. In addition, the lack of verified hosted pricing, service-level guarantees, maximum output limits, and provider-managed execution may make this model less convenient for production teams seeking a turnkey API.

When to choose UI-TARS-1.5-7B

Choose UI-TARS-1.5-7B when the main problem is understanding and acting on visual interfaces and you are prepared to manage the surrounding runtime. It is a reasonable candidate for:

  • Research into computer-use and GUI agents.
  • Browser or desktop automation prototypes.
  • Screenshot-based interface grounding.
  • Open-weight experimentation with visual action models.
  • Benchmarking models that predict coordinates and interface operations.

Its open-weight deployment is especially relevant when an organization wants local control over inference or needs to adapt the execution environment rather than rely on a provider-hosted computer-use product.

Another option may be more appropriate when the priority is general conversation, dedicated code generation, a managed API with predictable per-token billing, or native end-to-end computer control. The larger UI-TARS-1.5 system is the better-supported comparison when game-oriented performance or the broader research model is the deciding factor, although the supplied information does not provide a direct commercial price comparison. A conventional automation system may also be preferable for stable, deterministic workflows where fixed selectors and explicit rules are safer than model-generated actions.

Pricing, availability, and practical trade-offs

No official hosted per-token price was verified for UI-TARS-1.5-7B. The model is available as an open-weight Hugging Face checkpoint under Apache 2.0, so users generally evaluate cost through local or independently managed inference infrastructure. That does not make the model cost-free: hardware, storage, engineering, monitoring, and execution infrastructure still contribute to the total cost.

This creates a clear trade-off. Compared with a hosted model, self-hosting can provide greater control and avoid dependence on an unpublished provider API price, but it requires more technical work and may deliver lower convenience or variable throughput. Compared with a small deterministic automation script, UI-TARS-1.5-7B can handle more visually variable interfaces, but model inference introduces latency and the possibility of incorrect actions.

Overall, UI-TARS-1.5-7B is best understood as a foundation for computer-use agents rather than a finished desktop assistant. Its value lies in open deployment, visual grounding, and action-oriented reasoning; its main costs are the need for an execution wrapper, careful safety controls, and user-managed infrastructure.


Answers to Frequently Asked Questions

What are the main limitations of UI-TARS-1.5-7B?
UI-TARS-1.5-7B can misidentify controls, misunderstand changing or cluttered interfaces, hallucinate visual details, or select inefficient actions. It requires a separate computer-control runtime, and autonomous workflows should use restricted permissions and confirmation steps for authentication, financial actions, sensitive files, destructive operations, and protected websites. It is also not the strongest UI-TARS option for game-oriented tasks.
How much does UI-TARS-1.5-7B cost, and where can it be deployed?
No official hosted per-token price has been verified for this checkpoint. UI-TARS-1.5-7B is available as an open-weight Hugging Face model and can be self-hosted with compatible Transformers tooling or inference servers such as vLLM and SGLang. Users must account for hardware, infrastructure, engineering, monitoring, and execution costs.
What are the main technical specifications of UI-TARS-1.5-7B?
UI-TARS-1.5-7B is the ByteDance-Seed/UI-TARS-1.5-7B checkpoint, released on April 17, 2025. It accepts text and images, produces text containing reasoning and GUI actions, uses a published 128,000-position context configuration, and is distributed under the Apache 2.0 license. The model card describes it as approximately 8B parameters despite the 7B product name.
What is UI-TARS-1.5-7B designed to do?
UI-TARS-1.5-7B is an open-weight multimodal model from ByteDance Seed designed for computer use and GUI automation. It interprets screenshots and natural-language instructions, identifies interface elements, and generates text-based representations of actions such as clicking, typing, pressing keys, and scrolling.
Does UI-TARS-1.5-7B control a computer by itself?
No. The model generates GUI action representations but does not independently send operating-system events. A separate execution layer must capture screenshots, parse the model's output, convert coordinates when needed, perform mouse or keyboard actions, and provide the next screen observation.


Sources 4
Provider

About ByteDance Seed