What is Gemini Robotics ER 2?
Gemini Robotics ER 2 is Google DeepMind’s embodied reasoning vision-language model for robotics. It accepts text, images, video, and audio, and produces text-based reasoning, plans, and instructions that can be consumed by robot software, tools, or other models.
The model is intended to function as a robot’s high-level planning layer. In practical terms, it can help determine what is happening in a physical environment, decide which steps are needed to reach a goal, monitor whether those steps worked, and request another action when a step fails. A separate robot controller, vision-language-action model, or hardware system generally remains responsible for converting those instructions into safe motor movements.
This distinction is important. ER 2 is not presented as a general replacement for motion planning, hardware integration, perception pipelines, or safety systems. Its role is closer to an intelligent coordinator that connects perception, reasoning, tools, and robot actions over longer workflows.
Where ER 2 fits in Google DeepMind’s robotics lineup
Gemini Robotics ER 2 belongs to the Gemini Robotics 2 family and is specifically positioned as an embodied reasoning model. It can work alongside lower-level vision-language-action systems rather than replacing them. A vision-language-action model is designed to map visual and language information more directly to physical actions, while ER 2 concentrates on planning, task decomposition, progress monitoring, and orchestration.
This architecture allows a robotics system to divide responsibilities. ER 2 might interpret a request such as preparing a workstation, identify the sequence of subtasks, call a robot or tool for each step, inspect video to determine whether the step succeeded, and initiate recovery if necessary. The lower-level controller can then handle the detailed movement required to pick up, place, open, or manipulate an object.
What Gemini Robotics ER 2 can do
Spatial and physical reasoning
ER 2 is designed to reason about objects and their relationships in physical spaces. It can support object localization, spatially grounded planning, trajectory reasoning, and interpretation of changing surroundings. Example tasks include finding a particular item, identifying which object is closest to a target, reading an instrument, or determining how the arrangement of objects has changed.
These capabilities are useful when a robot must respond to an environment that is not completely known in advance. Instead of following only a fixed sequence, the system can use new visual or textual information to revise the next step. The model’s output still needs to be checked by the surrounding robotics system before an action is permitted.
Video understanding and progress monitoring
The model supports video moment finding and progress classification. It can help identify when an important event occurs in a continuous recording, determine whether a task has started or finished, and recognize when a step needs to be retried.
For a multi-step workflow, this means ER 2 can do more than produce an initial plan. It can also inspect evidence of what happened and help decide whether the system should continue, repeat a subtask, or recover from an unexpected result. This is particularly relevant to household, laboratory, warehouse, and workplace tasks where the physical state may change during execution.
Long-horizon task orchestration
ER 2 can combine reasoning with function calling, code execution, and custom robot APIs. These features allow it to coordinate several operations instead of treating every request as a single isolated response. It can delegate motor execution to a lower-level model while retaining responsibility for the larger task plan.
For example, a robot application could expose functions for checking a camera, moving an item, reading a sensor, or asking another robot to perform a task. ER 2 can select among those tools and use their results as part of the next planning step. The quality of the final behavior depends not only on the model but also on how carefully those tools, permissions, error states, and safety checks are designed.
Multi-robot collaboration
A notable focus of ER 2 is coordination among different robots. It can help systems communicate, account for different robot capabilities, divide work, and coordinate toward a shared objective. This could be useful when one robot can reach a location, another can manipulate an object, and a third system can inspect or transport the result.
The model does not remove the need for scheduling, communication protocols, collision avoidance, or hardware-level coordination. Instead, it provides a language- and vision-based reasoning layer that can help organize those activities.
Inputs, outputs, and model limits
The standard preview model accepts text, images, video, and audio. Its output is text, including reasoning-oriented responses, plans, tool calls, and instructions. It does not directly provide image, audio, video, speech, or motor-action output as a native model output.
| Specification | Verified detail |
|---|---|
| Model type | Embodied reasoning vision-language model |
| Standard model identifier | gemini-robotics-er-2-preview |
| Input modalities | Text, image, video, and audio |
| Output modality | Text |
| Input token limit | 131,072 tokens |
| Maximum output | 65,536 tokens |
| Release status | Public preview |
| Streaming variant | gemini-robotics-er-2-streaming-preview through the Live API |
The large input limit is relevant when a robotics application must combine instructions, visual context, previous observations, tool results, and task history. It does not guarantee that every long context will be interpreted perfectly. Robotics systems should still manage context carefully, retain the most relevant state, and validate important observations.
Tools and API support
The standard ER 2 preview endpoint supports function calling, code execution, computer use, file search, Google Search grounding, Google Maps grounding, URL context, and structured outputs. Batch API usage and context caching are also documented for the model.
Function calling is especially relevant to robotics because it lets an application expose controlled operations instead of asking the model to describe every action in ordinary prose. A tool might report sensor information, request a camera inspection, send a command to a lower-level controller, or query the status of another robot. Structured outputs can make responses easier for software to parse, but they do not by themselves make a proposed physical action safe.
Google also documents a streaming preview deployment for low-latency, bidirectional interaction through the Live API. This is a related deployment variant rather than a separate underlying model identity. Developers should distinguish the standard preview endpoint from the streaming endpoint when evaluating latency, integration requirements, and pricing.
Pricing and availability
Gemini Robotics ER 2 is available as a public preview through Google AI Studio and the Gemini API. Private or early preview access is also documented for Google Cloud’s Gemini Enterprise Agent Platform.
For the standard Gemini API preview endpoint, the supplied pricing is $1.00 per 1 million input tokens and $5.00 per 1 million output tokens, including reasoning tokens, through December 31, 2026. Google lists higher rates of $2.00 per 1 million input tokens and $10.00 per 1 million output tokens beginning January 1, 2027. Google AI Studio may provide usage without charge within its applicable limits, while paid API usage is billed according to the Gemini API pricing page.
These prices apply to token processing, not to the complete cost of operating a robot. Cameras, sensors, storage, network access, robot hardware, lower-level controllers, monitoring, and safety infrastructure can all add significant expense. The model is a preview release, so pricing, rate limits, regions, behavior, and availability may change.
Strengths and trade-offs
ER 2’s main strength is its focus on the parts of robotics that sit above direct motor control. It is designed for spatial reasoning, multi-step planning, video-based progress assessment, tool use, and coordination between robots or external systems. The combination of a large context limit and multimodal input can help an application maintain a broad view of a task and its changing physical environment.
Its trade-off is that this reasoning-oriented design can introduce more latency and system complexity than a narrowly scoped controller. A robot still needs additional components to transform text-based decisions into reliable physical behavior. The model’s output also needs independent checks because a plausible plan may be infeasible, incomplete, or unsafe in a particular environment.
In the supplied editorial assessment, ER 2 receives a high reasoning rating, a middle-range coding rating, slower speed rating, and moderate cost rating. These are editorial evaluations rather than Google-published benchmark results. They reflect the model’s intended role: careful orchestration and embodied reasoning are more important here than maximizing raw response speed or minimizing every API call.
When to choose Gemini Robotics ER 2
Choose ER 2 when a robotics application needs a high-level model to interpret multimodal observations and manage a task that unfolds over several steps. It is a strong candidate for:
- Planning household, laboratory, warehouse, or workplace workflows
- Reasoning about object locations, spatial relationships, and changing scenes
- Detecting task progress or completion from video
- Coordinating tools, custom robot APIs, and lower-level controllers
- Dividing work between multiple robots with different capabilities
- Providing a natural-language interface for human operators
Another option may be more appropriate when the requirement is direct low-level motor control, extremely predictable latency, or a compact controller for a narrowly defined movement. A dedicated motion planner, conventional control system, or vision-language-action model may be better suited to that part of the stack. ER 2 can still be used above those components when broader planning or recovery logic is needed.
Limitations and safety considerations
ER 2 should not be treated as a complete robotics system or as a safety authority. It does not replace perception validation, motion planning, collision avoidance, hardware interlocks, emergency stops, or human supervision. Its text output is not a guarantee that a proposed action is physically possible or safe.
Google advises discretion when using robotics models in production, commercial, public, or safety-critical environments. Deployments should use conservative tool permissions, independent safety systems, hardware-level safeguards, clear failure handling, and human oversight where appropriate. Testing should include unusual object positions, incomplete observations, failed actions, ambiguous instructions, communication loss, and disagreements between sensors and model assumptions.
As a preview model, ER 2 may also change over time. Teams should validate behavior against their own hardware and operating environment rather than assuming that a documented capability will perform consistently across all robots, scenes, languages, or task types.
Bottom line
Gemini Robotics ER 2 is best understood as a reasoning and orchestration model for embodied systems. It connects multimodal understanding with planning, progress monitoring, tool use, and multi-robot coordination, while leaving direct physical execution to other components. Its large context window and broad tool support make it suitable for complex, changing workflows, but its preview status, response latency, and need for independent safety controls make careful system integration essential.

