Seed GR

Seed GR-3

by ByteDance Seed · Current officially documented research model; no public hosted API, commercial pricing, or downloadable weights verified

Seed GR-3 is ByteDance Seed’s robotics model for turning language instructions, camera observations, and robot state into physical actions. Its reported capabilities include unseen-object manipulation, long-horizon table cleaning, mobile bimanual control, and deformable-cloth handling with ByteMini. Public API access, pricing, weights, and standard language-model limits are not verified.

Actions Reasoning Coding
Seed GR-3 is a robotics-focused model from ByteDance Seed that turns visual observations and natural-language instructions into robot actions. Its reported demonstrations cover unseen-object pick-and-place, extended table-cleaning sequences, and bimanual clothing manipulation using the ByteMini mobile robot. GR-3 is technically significant for its attempt to combine broad visual and language understanding with direct embodied control, but public information does not verify a hosted API, commercial pricing, downloadable weights, or a conventional text-generation context window.
Outputs

What Seed GR-3 can produce

Actions
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Fine-tuning Multimodal output
Model profile

Performance characteristics

6/10 Reasoning
1/10 Coding
6/10 Speed
Specifications

Technical details

Model family Seed GR
Model type Other
Release date 2025-07-22
Status Current officially documented research model; no public hosted API, commercial pricing, or downloadable weights verified
Knowledge cutoff notes

No explicit knowledge cutoff for the exact GR-3 model was found in the official model page, technical report page, or release material. The model is trained for embodied control using vision-language, human-trajectory, and robot-trajectory data rather than being documented as a general knowledge chatbot.

Model notes

Seed GR-3 is a vision-language-action robotics model rather than a conventional text-generation model. ByteDance Seed describes an approximately 4B-parameter mixture-of-transformers system combining a vision-language module with a diffusion-transformer action generator. Inputs include natural-language instructions, RGB observations from multiple robot cameras, and robot state information. The model natively produces robot actions and task-status signals for embodied control. Its demonstrations use the ByteMini bimanual mobile robot. The official material documents research results and a technical report, but does not verify public model weights, a general developer API, token pricing, a standard context window, or a maximum text-output token limit. Editorial scores reflect robotics usefulness rather than general-purpose language-model performance.

Model guide

Seed GR-3: ByteDance Seed’s Vision-Language-Action Model for Dexterous Robots

Seed GR-3 is ByteDance Seed’s approximately 4-billion-parameter vision-language-action model for general-purpose robotic manipulation. It combines language instructions, camera observations, and robot state information to generate physical robot actions. The model is designed for long-horizon, bimanual, mobile, and dexterous tasks, including manipulation of unfamiliar objects and deformable items such as clothing. It is documented as a research model rather than a general-purpose chatbot or commercial API product.

What is Seed GR-3?

Seed GR-3 is a vision-language-action model developed by ByteDance Seed for general-purpose robot manipulation. A vision-language-action, or VLA, model connects perception and instruction following to physical control: instead of replying with text, it interprets what a robot sees, understands what it has been asked to do, and generates actions for the robot’s joints or actuators.

The model accepts natural-language instructions, visual observations from robot-mounted cameras, and robot state information. It then produces action sequences and task-status signals for an embodied robot system. This makes GR-3 fundamentally different from a conventional language model. Its primary output is not an explanation, image, or code sample, but a control policy that can be executed by compatible robotic hardware.

ByteDance Seed reports that GR-3 contains approximately 4 billion parameters and uses a mixture-of-transformers architecture. The model is presented in the provider’s robotics catalog alongside the broader Seed research portfolio, but it is not documented as a consumer assistant or general-purpose developer model.

How GR-3 combines vision, language, and action

GR-3 separates the problem into closely connected but specialized components. Its vision-language module processes the user’s instruction and the visual scene, while an action-generation module produces robot movements. ByteDance Seed describes the action component as a diffusion transformer that uses flow matching to generate actions.

In practical terms, the model can use information such as the location of an object, its relationship to another object, its approximate size, and the requested task. It also receives robot state information, which helps relate the intended movement to the robot’s current configuration. The result is a policy intended to connect semantic understanding with continuous physical control rather than merely identifying objects or describing a plan.

The training approach combines robot trajectory data, human trajectories recorded through virtual-reality devices, and publicly accessible vision-language data. Robot demonstrations provide examples of physical manipulation, while human trajectories can help the system learn new tasks and objects more efficiently. Vision-language training contributes broader semantic and visual knowledge, including concepts such as spatial relationships and object categories.

What can Seed GR-3 do?

According to ByteDance Seed’s official research material, GR-3 was evaluated on generalizable pick-and-place, long-horizon table-bussing, and dexterous cloth-manipulation tasks. The provider reports that the model can operate with objects, environments, and instructions that were not present in the robot training data.

Generalization to unfamiliar objects and scenes

Pick-and-place tasks are used to test whether a robot can move beyond memorized demonstrations. GR-3 was tested with unseen objects and environments and with instructions involving object categories and spatial relationships. This capability is important because a practical household or workplace robot cannot rely on every object being identical to the examples used during training.

The reported generalization should be understood as a provider claim tied to the evaluated setups, not as evidence that the model can reliably manipulate every unfamiliar object. Real-world performance still depends on camera quality, robot calibration, the action interface, the physical properties of the object, and the similarity between deployment conditions and training conditions.

Long-horizon table cleaning

GR-3 was demonstrated with the ByteMini mobile robot on a long-horizon table-cleaning task. The system performed multiple related subtasks, including packing food, organizing utensils, and moving containers. These tasks require maintaining progress across a sequence rather than completing one isolated grasp.

Long-horizon control is difficult because small mistakes can accumulate. A system must keep track of the task state, choose useful intermediate actions, and continue operating when the scene changes. GR-3’s table-bussing demonstration is therefore relevant to research into robots that perform connected household or service tasks rather than single-step laboratory actions.

Bimanual manipulation of clothing

The model was also tested on hanging clothing, a deformable-object task that requires coordinated use of both arms. Cloth is more difficult than a rigid object because its shape changes as it is lifted, folded, or pulled. Successful handling requires the robot to respond to visual changes instead of following a fixed movement sequence.

This demonstration highlights the intended scope of GR-3: it is designed not only for simple object relocation but also for manipulation that involves mobile platforms, two arms, and changing object geometry.

How GR-3 relates to the ByteMini robot

ByteMini is the physical robot platform used in the reported GR-3 demonstrations. It is described as a bimanual mobile robot with multiple cameras, two seven-degree-of-freedom arms, and spherical wrist designs intended to support dexterous movement in confined spaces.

GR-3 supplies the perception and action policy, while ByteMini provides the sensors, mobility, arms, and mechanical interface needed to carry out the policy. This distinction matters when evaluating the model. A demonstration attributed to GR-3 is also dependent on the robot’s hardware, camera placement, calibration, low-level controllers, and safety systems. GR-3 should not be treated as a self-contained robot that works independently of a compatible embodiment.

Inputs, outputs, and modalities

GR-3 supports multimodal embodied control rather than ordinary text generation. The documented inputs include natural-language instructions, RGB observations from multiple robot cameras, and robot state information. Its output is native robot action data, including joint or control trajectories and task-status signals.

For database purposes, this means the model has multimodal input and action output. It is not documented as producing images, video, audio, or ordinary text responses. The model’s action output should not be confused with image or video generation: it represents commands for a robot control system.

CapabilityVerified or reported status
Language inputNatural-language task instructions are supported.
Visual inputRGB observations from multiple robot cameras are used.
Robot state inputRobot state information is part of the model input.
Action outputRobot actions, including joint or control trajectories, are generated.
Text, image, audio, or video outputNot documented as the model’s primary output.

Availability, pricing, and API status

ByteDance Seed lists Seed GR-3 on its official robotics model catalog and provides a model page, research announcement, and technical report. The available material presents GR-3 primarily as a research model for embodied robotics.

No public hosted API, commercial token pricing, downloadable model weights, general-purpose chat interface, standard language-model context limit, or maximum text-output limit was verified for the exact model. Consequently, there is no supported price comparison with text or multimodal API models. The absence of published pricing does not mean that no future access route will exist; it means that a public commercial offering was not established in the supplied sources.

GR-3 also should not be evaluated through typical API features such as streaming text, JSON mode, or function calling. The documented control interface is an embodied action interface, and the supplied research does not verify a general developer API with those conventions.

Strengths and limitations

Where GR-3 is strong

  • Embodied control: It directly connects language and visual understanding to robot actions instead of stopping at recognition or planning.
  • Task generalization: ByteDance Seed reports tests involving unseen objects, environments, and instructions.
  • Long-horizon behavior: The table-cleaning demonstration addresses multiple dependent subtasks rather than one isolated movement.
  • Dexterous and bimanual work: The ByteMini demonstrations include coordinated manipulation of deformable clothing.
  • Mixed training data: Robot trajectories, VR-recorded human trajectories, and vision-language data are combined to support both physical control and broader semantic understanding.

Where GR-3 is limited

  • Hardware dependence: Performance depends on a compatible robot, cameras, calibration, state feedback, and action interface.
  • Research availability: Public commercial access, weights, hosted inference, and pricing are not verified.
  • Generalization is not unlimited: ByteDance Seed notes difficulties with unfamiliar object shapes, novel concepts, and unexpected physical events.
  • Not a conversational model: It is not intended for chat, text generation, software coding, image generation, or general knowledge tasks.
  • No conventional context specification: A standard token context window and maximum text output are not provided for this robotics model.

These limitations are especially important for deployment. A strong result in a controlled demonstration does not by itself establish safe operation in homes, factories, or public environments. Physical robots must also handle collisions, fragile objects, people, and unexpected changes that may fall outside the training distribution.

Reasoning, coding, and tool support

GR-3 demonstrates task-level reasoning in the practical sense that it can interpret instructions, relate objects spatially, sequence subtasks, and generate actions over an extended interaction. However, the supplied research does not define a separate reasoning mode, chain-of-thought interface, or general reasoning benchmark. Any editorial reasoning assessment should therefore be understood as an evaluation of embodied task usefulness, not a provider-published language reasoning score.

The model is not documented as a coding model and does not provide ordinary software-generation capabilities. Likewise, no conventional tool-calling or function-calling interface is verified. Its relationship with tools is physical: the robot platform and its sensors act as the model’s embodiment and control channel.

When to choose Seed GR-3

Choose Seed GR-3 when the main problem is embodied robotics research and the project can provide compatible hardware and sensor/action interfaces. It is particularly relevant for experiments involving generalist manipulation, long sequences of tasks, mobile bimanual robots, unfamiliar objects, or deformable materials.

GR-3 is a better fit than a text-only model when the desired result is a direct robot policy rather than a written plan. It may also be preferable to a narrow single-task controller when the research goal is to study transfer across objects, environments, and instructions.

Another type of model may be more appropriate when the project needs a public API, predictable usage pricing, downloadable weights, conventional text or code generation, or a mature production integration. A conventional vision-language model can be suitable for image understanding or high-level planning without directly controlling a robot. A specialized low-level controller may be preferable where strict latency, deterministic behavior, or safety certification matters more than broad task generalization.

Within ByteDance Seed’s wider portfolio, models such as Seed2.1, Seedance, and Seedream serve different purposes, including productivity, coding, and creative image or video generation. They are not substitutes for GR-3’s embodied action role. GR-3 should be selected specifically for robotics research, not simply because it belongs to the same provider.

Bottom line

Seed GR-3 is a research-oriented vision-language-action model that aims to give general-purpose robots a stronger connection between perception, language, and physical manipulation. Its reported demonstrations cover unseen-object tasks, long-horizon table cleaning, mobile bimanual control, and deformable clothing. The model’s most important distinction is its action output: it is designed to control an embodied robot, not to act as a chatbot or conventional multimodal API.

Its practical value depends on access to suitable hardware and on whether the deployment environment resembles the conditions used during training and evaluation. Because public API access, pricing, model weights, and standard language-model limits have not been verified, GR-3 is currently best understood as an important documented robotics research system rather than a readily deployable commercial model.


Answers to Frequently Asked Questions

Is Seed GR-3 available through a public API or commercial pricing?
A public hosted API, commercial token pricing, downloadable model weights, and a general-purpose chat interface were not verified for Seed GR-3. The available information presents it primarily as a research model for embodied robotics rather than as a conventional commercial language-model service.
What are the inputs and outputs of Seed GR-3?
The documented inputs are natural-language instructions, RGB observations from multiple robot cameras, and robot-state information. Its outputs are robot actions, including joint or control trajectories and task-status signals, rather than ordinary text, images, audio, or video.
How does Seed GR-3 work with the ByteMini robot?
GR-3 provides the perception and action policy, while ByteMini supplies the cameras, mobility, two seven-degree-of-freedom arms, mechanical interfaces, and low-level control systems. Performance therefore depends on the robot hardware, calibration, sensors, and safety systems as well as the model.
What is Seed GR-3?
Seed GR-3 is a vision-language-action model developed by ByteDance Seed for general-purpose robot manipulation. It combines natural-language instructions, camera observations, and robot-state information to generate physical actions and task-status signals for compatible robots.
What can Seed GR-3 do?
ByteDance Seed reports that GR-3 was evaluated on generalizable pick-and-place, long-horizon table cleaning, and bimanual cloth-manipulation tasks. The reported demonstrations include unfamiliar objects and environments, multi-step household actions, and coordinated handling of deformable clothing.


Sources 5
Provider

About ByteDance Seed