Gemini Robotics ER

Gemini Robotics ER 2

by Google DeepMind · Public preview

Google DeepMind’s Gemini Robotics ER 2 is a public-preview vision-language model that provides high-level embodied reasoning for robots. It accepts text, images, video, and audio, supports spatial and temporal understanding, orchestrates tools and lower-level robot models, tracks task progress, and coordinates multiple robots. Its 131,072-token input limit, 65,536-token maximum output, function calling, code execution, structured outputs, and streaming variant make it suited to complex robotics workflows, while physical safety and motor control remain the responsibility of surrounding systems.

Text Reasoning Coding
Gemini Robotics ER 2 is designed to act as a high-level intelligence layer for physical robots. It can reason about environments, objects, actions, and task progress, then use function calls, code, custom robot APIs, or other models to help complete a workflow. Its main distinction is that it focuses on embodied reasoning and orchestration instead of serving as a standalone low-level motion controller.
Outputs

What Gemini Robotics ER 2 can produce

Text
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching Batch API
Model profile

Performance characteristics

8/10 Reasoning
5/10 Coding
4/10 Speed
5/10 Cost efficiency
Specifications

Technical details

Model family Gemini Robotics ER
Model type Reasoning
Context window 131K tokens
Maximum output 66K tokens
Release date 2026-07-30
Status Public preview
Knowledge cutoff notes

Google's official ER 2 model card and Gemini API documentation do not publish an exact knowledge cutoff. The model is described as being based on Gemini 3.5 Flash, but the parent model's cutoff is not used as a substitute for an exact ER 2 cutoff.

Model notes

The canonical model is Gemini Robotics ER 2, with the standard API identifier gemini-robotics-er-2-preview. Google also offers gemini-robotics-er-2-streaming-preview for low-latency bidirectional Live API interactions; it is treated as a deployment variant rather than a separate model entity here. ER 2 is based on Gemini 3.5 Flash and is intended to coordinate with lower-level vision-language-action models or robot controllers. The model card states that robotics models should not be used for safety-critical applications without appropriate safeguards. No exact knowledge cutoff is published for ER 2.

Cost

Model pricing

Input $1.00 per 1 million tokens for the standard preview endpoint through December 31, 2026; $2.00 per 1 million tokens starting January 1, 2027
Output $5.00 per 1 million tokens, including reasoning tokens, through December 31, 2026; $10.00 per 1 million tokens starting January 1, 2027
Model guide

Gemini Robotics ER 2: Google’s High-Level Reasoning Model for Robots

Gemini Robotics ER 2 is Google DeepMind’s embodied reasoning vision-language model for robots. It interprets text, images, video, and audio to plan multi-step physical tasks, monitor progress, use tools, coordinate robots, and provide high-level instructions to lower-level controllers rather than directly generating motor commands.

What is Gemini Robotics ER 2?

Gemini Robotics ER 2 is Google DeepMind’s embodied reasoning vision-language model for robotics. It accepts text, images, video, and audio, and produces text-based reasoning, plans, and instructions that can be consumed by robot software, tools, or other models.

The model is intended to function as a robot’s high-level planning layer. In practical terms, it can help determine what is happening in a physical environment, decide which steps are needed to reach a goal, monitor whether those steps worked, and request another action when a step fails. A separate robot controller, vision-language-action model, or hardware system generally remains responsible for converting those instructions into safe motor movements.

This distinction is important. ER 2 is not presented as a general replacement for motion planning, hardware integration, perception pipelines, or safety systems. Its role is closer to an intelligent coordinator that connects perception, reasoning, tools, and robot actions over longer workflows.

Where ER 2 fits in Google DeepMind’s robotics lineup

Gemini Robotics ER 2 belongs to the Gemini Robotics 2 family and is specifically positioned as an embodied reasoning model. It can work alongside lower-level vision-language-action systems rather than replacing them. A vision-language-action model is designed to map visual and language information more directly to physical actions, while ER 2 concentrates on planning, task decomposition, progress monitoring, and orchestration.

This architecture allows a robotics system to divide responsibilities. ER 2 might interpret a request such as preparing a workstation, identify the sequence of subtasks, call a robot or tool for each step, inspect video to determine whether the step succeeded, and initiate recovery if necessary. The lower-level controller can then handle the detailed movement required to pick up, place, open, or manipulate an object.

What Gemini Robotics ER 2 can do

Spatial and physical reasoning

ER 2 is designed to reason about objects and their relationships in physical spaces. It can support object localization, spatially grounded planning, trajectory reasoning, and interpretation of changing surroundings. Example tasks include finding a particular item, identifying which object is closest to a target, reading an instrument, or determining how the arrangement of objects has changed.

These capabilities are useful when a robot must respond to an environment that is not completely known in advance. Instead of following only a fixed sequence, the system can use new visual or textual information to revise the next step. The model’s output still needs to be checked by the surrounding robotics system before an action is permitted.

Video understanding and progress monitoring

The model supports video moment finding and progress classification. It can help identify when an important event occurs in a continuous recording, determine whether a task has started or finished, and recognize when a step needs to be retried.

For a multi-step workflow, this means ER 2 can do more than produce an initial plan. It can also inspect evidence of what happened and help decide whether the system should continue, repeat a subtask, or recover from an unexpected result. This is particularly relevant to household, laboratory, warehouse, and workplace tasks where the physical state may change during execution.

Long-horizon task orchestration

ER 2 can combine reasoning with function calling, code execution, and custom robot APIs. These features allow it to coordinate several operations instead of treating every request as a single isolated response. It can delegate motor execution to a lower-level model while retaining responsibility for the larger task plan.

For example, a robot application could expose functions for checking a camera, moving an item, reading a sensor, or asking another robot to perform a task. ER 2 can select among those tools and use their results as part of the next planning step. The quality of the final behavior depends not only on the model but also on how carefully those tools, permissions, error states, and safety checks are designed.

Multi-robot collaboration

A notable focus of ER 2 is coordination among different robots. It can help systems communicate, account for different robot capabilities, divide work, and coordinate toward a shared objective. This could be useful when one robot can reach a location, another can manipulate an object, and a third system can inspect or transport the result.

The model does not remove the need for scheduling, communication protocols, collision avoidance, or hardware-level coordination. Instead, it provides a language- and vision-based reasoning layer that can help organize those activities.

Inputs, outputs, and model limits

The standard preview model accepts text, images, video, and audio. Its output is text, including reasoning-oriented responses, plans, tool calls, and instructions. It does not directly provide image, audio, video, speech, or motor-action output as a native model output.

SpecificationVerified detail
Model typeEmbodied reasoning vision-language model
Standard model identifiergemini-robotics-er-2-preview
Input modalitiesText, image, video, and audio
Output modalityText
Input token limit131,072 tokens
Maximum output65,536 tokens
Release statusPublic preview
Streaming variantgemini-robotics-er-2-streaming-preview through the Live API

The large input limit is relevant when a robotics application must combine instructions, visual context, previous observations, tool results, and task history. It does not guarantee that every long context will be interpreted perfectly. Robotics systems should still manage context carefully, retain the most relevant state, and validate important observations.

Tools and API support

The standard ER 2 preview endpoint supports function calling, code execution, computer use, file search, Google Search grounding, Google Maps grounding, URL context, and structured outputs. Batch API usage and context caching are also documented for the model.

Function calling is especially relevant to robotics because it lets an application expose controlled operations instead of asking the model to describe every action in ordinary prose. A tool might report sensor information, request a camera inspection, send a command to a lower-level controller, or query the status of another robot. Structured outputs can make responses easier for software to parse, but they do not by themselves make a proposed physical action safe.

Google also documents a streaming preview deployment for low-latency, bidirectional interaction through the Live API. This is a related deployment variant rather than a separate underlying model identity. Developers should distinguish the standard preview endpoint from the streaming endpoint when evaluating latency, integration requirements, and pricing.

Pricing and availability

Gemini Robotics ER 2 is available as a public preview through Google AI Studio and the Gemini API. Private or early preview access is also documented for Google Cloud’s Gemini Enterprise Agent Platform.

For the standard Gemini API preview endpoint, the supplied pricing is $1.00 per 1 million input tokens and $5.00 per 1 million output tokens, including reasoning tokens, through December 31, 2026. Google lists higher rates of $2.00 per 1 million input tokens and $10.00 per 1 million output tokens beginning January 1, 2027. Google AI Studio may provide usage without charge within its applicable limits, while paid API usage is billed according to the Gemini API pricing page.

These prices apply to token processing, not to the complete cost of operating a robot. Cameras, sensors, storage, network access, robot hardware, lower-level controllers, monitoring, and safety infrastructure can all add significant expense. The model is a preview release, so pricing, rate limits, regions, behavior, and availability may change.

Strengths and trade-offs

ER 2’s main strength is its focus on the parts of robotics that sit above direct motor control. It is designed for spatial reasoning, multi-step planning, video-based progress assessment, tool use, and coordination between robots or external systems. The combination of a large context limit and multimodal input can help an application maintain a broad view of a task and its changing physical environment.

Its trade-off is that this reasoning-oriented design can introduce more latency and system complexity than a narrowly scoped controller. A robot still needs additional components to transform text-based decisions into reliable physical behavior. The model’s output also needs independent checks because a plausible plan may be infeasible, incomplete, or unsafe in a particular environment.

In the supplied editorial assessment, ER 2 receives a high reasoning rating, a middle-range coding rating, slower speed rating, and moderate cost rating. These are editorial evaluations rather than Google-published benchmark results. They reflect the model’s intended role: careful orchestration and embodied reasoning are more important here than maximizing raw response speed or minimizing every API call.

When to choose Gemini Robotics ER 2

Choose ER 2 when a robotics application needs a high-level model to interpret multimodal observations and manage a task that unfolds over several steps. It is a strong candidate for:

  • Planning household, laboratory, warehouse, or workplace workflows
  • Reasoning about object locations, spatial relationships, and changing scenes
  • Detecting task progress or completion from video
  • Coordinating tools, custom robot APIs, and lower-level controllers
  • Dividing work between multiple robots with different capabilities
  • Providing a natural-language interface for human operators

Another option may be more appropriate when the requirement is direct low-level motor control, extremely predictable latency, or a compact controller for a narrowly defined movement. A dedicated motion planner, conventional control system, or vision-language-action model may be better suited to that part of the stack. ER 2 can still be used above those components when broader planning or recovery logic is needed.

Limitations and safety considerations

ER 2 should not be treated as a complete robotics system or as a safety authority. It does not replace perception validation, motion planning, collision avoidance, hardware interlocks, emergency stops, or human supervision. Its text output is not a guarantee that a proposed action is physically possible or safe.

Google advises discretion when using robotics models in production, commercial, public, or safety-critical environments. Deployments should use conservative tool permissions, independent safety systems, hardware-level safeguards, clear failure handling, and human oversight where appropriate. Testing should include unusual object positions, incomplete observations, failed actions, ambiguous instructions, communication loss, and disagreements between sensors and model assumptions.

As a preview model, ER 2 may also change over time. Teams should validate behavior against their own hardware and operating environment rather than assuming that a documented capability will perform consistently across all robots, scenes, languages, or task types.

Bottom line

Gemini Robotics ER 2 is best understood as a reasoning and orchestration model for embodied systems. It connects multimodal understanding with planning, progress monitoring, tool use, and multi-robot coordination, while leaving direct physical execution to other components. Its large context window and broad tool support make it suitable for complex, changing workflows, but its preview status, response latency, and need for independent safety controls make careful system integration essential.


Answers to Frequently Asked Questions

Is Gemini Robotics ER 2 safe to use as a standalone robot controller?
No. Gemini Robotics ER 2 should not be treated as a complete robotics system or safety authority. Deployments still require independent perception validation, motion planning, collision avoidance, hardware interlocks, emergency stops, failure handling, and appropriate human supervision.
What is the model identifier and context limit for Gemini Robotics ER 2?
The standard model identifier is `gemini-robotics-er-2-preview`. It supports up to 131,072 input tokens and a maximum of 65,536 output tokens. A streaming variant, `gemini-robotics-er-2-streaming-preview`, is available through the Live API.
What inputs and outputs does Gemini Robotics ER 2 support?
The standard preview model accepts text, images, video, and audio. Its native output is text, including reasoning-oriented responses, plans, tool calls, and instructions. It does not directly produce motor commands, images, audio, video, or speech as native model outputs.
What is the difference between Gemini Robotics ER 2 and a vision-language-action model?
Gemini Robotics ER 2 focuses on high-level planning, task decomposition, progress monitoring, tool use, and orchestration. A vision-language-action model generally maps visual and language inputs more directly to physical robot actions. ER 2 can delegate detailed movement to those lower-level systems.
What is Gemini Robotics ER 2 used for?
Gemini Robotics ER 2 is a high-level embodied reasoning model for robotics. It interprets text, images, video, and audio to create plans, monitor task progress, use tools, coordinate robots, and request recovery actions when steps fail. It is intended to work alongside lower-level controllers rather than directly control motors.


Sources 6
Provider

About Google DeepMind