Gemini Robotics ER 2

Gemini Robotics ER 2 Streaming

by Google DeepMind · Public preview

A Google DeepMind preview endpoint for embodied reasoning in real-time robotic systems. It accepts text, images, video, and audio, streams text through the Gemini Live API, supports function calling, and is designed for continuous monitoring, task coordination, and multi-robot workflows rather than direct motor control.

Text Reasoning Coding
Gemini Robotics ER 2 Streaming is the real-time version of Google DeepMind's Gemini Robotics ER 2 embodied reasoning model. It is designed for robotic systems that continuously observe changing physical environments, reason about what is happening, and communicate decisions or invoke application-defined functions through the Gemini Live API.
Outputs

What Gemini Robotics ER 2 Streaming can produce

Text
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming
Model profile

Performance characteristics

8/10 Reasoning
4/10 Coding
9/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Gemini Robotics ER 2
Model type Reasoning
Context window 131K tokens
Maximum output 66K tokens
Release date 2026-07-30
Status Public preview
Knowledge cutoff notes

Google's public model documentation does not provide a verified knowledge-cutoff date for the exact Gemini Robotics ER 2 Streaming endpoint.

Model notes

The exact model endpoint is gemini-robotics-er-2-streaming-preview. It is a distinct real-time endpoint within the Gemini Robotics ER 2 family and uses the Gemini Live API for bidirectional streaming. The model accepts text, image, video, and audio input but produces text output. Function calling is supported, while the endpoint documentation lists structured outputs, context caching, batch API, code execution, computer use, file search, and URL context as unsupported. Google describes it as an embodied reasoning model rather than a lower-level vision-language-action motor controller. Pricing is subject to scheduled increases on January 1, 2027.

Cost

Model pricing

Input $1.00 per 1M tokens through December 31, 2026; $2.00 per 1M tokens from January 1, 2027
Output $5.00 per 1M tokens through December 31, 2026; $10.00 per 1M tokens from January 1, 2027
Model guide

Gemini Robotics ER 2 Streaming: Real-Time Reasoning for Robotic Agents

Gemini Robotics ER 2 Streaming is a Google DeepMind preview endpoint for robots that need to continuously interpret multimodal sensor input and respond through low-latency, bidirectional streaming. It accepts text, images, video, and audio, produces streamed text, and supports function calling for application-controlled robotic actions.

What is Gemini Robotics ER 2 Streaming?

Gemini Robotics ER 2 Streaming is a vision-language model endpoint from Google DeepMind for embodied reasoning. In practical terms, it is intended to help a robotic application interpret its surroundings and decide what information or action should come next. The model is not presented as a complete robot controller; instead, it provides higher-level reasoning that can be connected to a robot's control software.

The defining characteristic is streaming. Rather than waiting for one complete request and returning one complete response, the endpoint is designed for an ongoing exchange between a robot and the model. A system can send observations such as video, images, audio, or text, receive streamed textual guidance, and use function calls to connect the model with application-defined operations.

The exact preview model identifier is gemini-robotics-er-2-streaming-preview. It is available through Google AI Studio and the Gemini API as a public preview endpoint.

Where it fits in the Gemini Robotics family

Gemini Robotics ER 2 Streaming belongs to the Gemini Robotics ER 2 family, but it has a distinct role from the standard Gemini Robotics ER 2 endpoint. The streaming variant is specifically documented for Live API use and is optimized for continuous, low-latency interaction.

This positioning matters when selecting an endpoint. A conventional request-and-response model may be more suitable for occasional analysis, while the streaming variant is aimed at systems that need to keep a communication loop open as the robot's environment changes. The model's role is closer to an embodied reasoning service than to a traditional chatbot or a low-level vision-language-action controller.

How the streaming architecture works

Gemini Robotics ER 2 Streaming uses the Gemini Live API for bidirectional communication. “Bidirectional” means the application can continue sending input while the model is producing output, rather than treating every exchange as an isolated transaction.

For example, a warehouse robot could provide a continuing video stream, send an instruction in text, and receive streamed textual guidance as objects or task conditions change. When the model identifies a step that requires an external operation, function calling can connect that decision to an application-defined tool. The tool might represent a robot-control command, a monitoring action, a coordination service, or another operation implemented outside the model.

Function calling does not mean that the model directly provides deterministic motor commands. It provides a structured way for the surrounding application to decide how a suggested operation should be validated and executed. Safety checks, permissions, sensor validation, and low-level control remain responsibilities of the robotics system.

Inputs, outputs, and documented limits

SpecificationDocumented detail
Input typesText, images, video, and audio
Primary outputText streamed through the Gemini Live API
Input context limit131,072 tokens
Maximum output65,536 tokens
StreamingSupported
Function callingSupported
Model statusPublic preview

The model supports multimodal input but does not directly produce images, audio, video, or speech according to the supplied model profile. Its output is text, even when the input includes live or recorded audio and video. That text can then be interpreted by the application or used alongside function calling.

The 131,072-token context limit provides room for substantial task context, observations, and conversation history, while the maximum output limit is documented as 65,536 tokens. These are token limits rather than direct guarantees about the duration of a video stream or the number of camera frames a robotics application can process. Real-world throughput and latency will also depend on the implementation, media encoding, network conditions, and API behavior.

Robotics capabilities and reasoning

Google positions the endpoint for embodied reasoning: reasoning about situations in the physical world rather than only processing text. Its supported input types allow a robotics application to combine visual and auditory observations with instructions or task state.

  • Physical-world understanding: Interprets multimodal observations from a robotic environment.
  • Spatial reasoning: Helps reason about objects, locations, and relationships between elements in a scene.
  • Temporal reasoning: Supports interpretation of movement and changing task conditions over time.
  • Continuous monitoring: Allows an application to respond as observations arrive and events occur.
  • Task coordination: Can contribute to workflows in which multiple robots share state or delegate subtasks.
  • Function-based orchestration: Connects model decisions to tools and operations implemented by the application.

These capabilities make the endpoint relevant to tasks such as monitoring a warehouse scene, coordinating logistics steps, interpreting a changing work area, or helping multiple robots share task information. They should not be confused with guaranteed physical understanding or autonomous safety. The model can provide reasoning that is useful to a control system, but its responses still require validation before they affect physical equipment.

Speed, cost, and editorial assessment

The main performance trade-off is the choice of a specialized streaming endpoint over a general request-and-response workflow. Its Live API integration is intended to reduce interaction delays and support reactive agents. That makes it a better fit for continuing observation and coordination than a workflow that only needs an occasional written analysis.

The supplied editorial profile rates its speed at 9 out of 10, reasoning at 8 out of 10, coding at 4 out of 10, and cost at 7 out of 10. These are editorial assessments, not benchmark results or ratings published by Google. The relatively low coding assessment reflects the model's specialized robotics role rather than a claim that it cannot process code-related prompts.

For the paid Standard tier, the documented input price is $1.00 per 1 million tokens through December 31, 2026, increasing to $2.00 per 1 million tokens from January 1, 2027. Output costs $5.00 per 1 million tokens through December 31, 2026, and are scheduled to increase to $10.00 per 1 million tokens from January 1, 2027.

Streaming systems can generate ongoing input and output usage, so total cost depends on how much sensor data and textual context the application sends and how long sessions remain active. The per-token price alone does not determine the cost of operating a robot fleet.

Main strengths and limitations

Strengths

  • Designed for continuous interaction: Live API support and bidirectional streaming suit systems that need to react while observations are still arriving.
  • Broad multimodal input: Text, image, video, and audio can be used as inputs to an embodied reasoning workflow.
  • Useful application integration: Function calling provides a mechanism for connecting reasoning to robot software and coordination services.
  • Large documented limits: The endpoint supports a 131,072-token input context and up to 65,536 output tokens.
  • Robotics-specific positioning: Its purpose is clearer than that of a general-purpose language model when the problem involves physical environments and robot task state.

Limitations

  • Preview status: The endpoint is publicly available in preview, so developers should account for possible changes and should validate behavior before relying on it in production.
  • Text-only output: It does not directly generate images, audio, video, or speech. Applications needing those outputs require other components.
  • Not a low-level controller: It should not replace deterministic motor-control software, emergency stops, or validated safety layers.
  • Limited listed feature set: The supplied documentation profile does not list structured outputs, context caching, batch processing, code execution, computer use, file search, or URL context as supported for this endpoint.
  • Safety-critical limitations: Healthcare, transportation, industrial, and other high-risk deployments require independent safeguards and human oversight.

When to choose Gemini Robotics ER 2 Streaming

Choose Gemini Robotics ER 2 Streaming when the application needs an ongoing multimodal conversation with a robot and the value of fast reaction is more important than using a simple one-shot generation workflow. Good candidates include:

  • Robots that continuously inspect a work area through video or images.
  • Warehouse and logistics systems that need higher-level task interpretation.
  • Applications coordinating tasks across multiple robots.
  • Audio- and video-aware robotic assistants that need streamed responses.
  • Prototypes that connect embodied reasoning to application-defined functions.

Another option may be more appropriate when the task requires direct motor control, deterministic behavior, strict structured-output support, extensive code generation, batch processing, or a non-streaming analysis workflow. A conventional model endpoint may be simpler for occasional scene descriptions or planning requests, while a lower-level robotics controller is better suited to precise, safety-critical actuation. The standard Gemini Robotics ER 2 endpoint may also be preferable when the application does not need the streaming variant's Live API-specific behavior, although the supplied research does not provide a complete feature-by-feature comparison.

Practical safety guidance

Gemini Robotics ER 2 Streaming should be treated as one reasoning component in a larger robotics architecture. A responsible deployment should separate model-generated suggestions from the mechanisms that authorize and execute physical actions.

Robot-level emergency stops, deterministic control layers, sensor validation, permissions, simulation or staged testing, and operational safety procedures remain necessary. Developers should also define what happens when the model is uncertain, delayed, unavailable, or inconsistent with sensor data. These controls are especially important because a streamed response can arrive while the physical environment is changing.

Overall, the endpoint is most distinctive as a low-latency embodied reasoning interface. Its value lies in connecting continuous multimodal observations with streamed text and application functions, not in replacing the complete software and safety stack required to operate physical robots.


Answers to Frequently Asked Questions

Is Gemini Robotics ER 2 Streaming a robot controller?
No. It provides higher-level embodied reasoning and can suggest operations through function calling, but it does not replace deterministic motor-control software, emergency stops, sensor validation, permissions, or other robotics safety layers. All model-generated actions should be validated before execution.
What is Gemini Robotics ER 2 Streaming?
Gemini Robotics ER 2 Streaming is a Google DeepMind vision-language model endpoint for embodied reasoning in robotics. It helps applications interpret physical environments and determine what information or action should come next, while leaving low-level control to the surrounding robotics system.
What inputs and outputs does Gemini Robotics ER 2 Streaming support?
The endpoint accepts text, images, video, and audio as inputs. Its primary output is streamed text through the Gemini Live API; it does not directly generate images, audio, video, or speech. The documented input context limit is 131,072 tokens, and the maximum output is 65,536 tokens.
How does Gemini Robotics ER 2 Streaming work with the Gemini Live API?
It uses the Gemini Live API for bidirectional, continuous communication. A robot application can stream text, images, video, or audio to the model while receiving streamed textual guidance and connecting model decisions to application-defined tools through function calling.
When should developers choose Gemini Robotics ER 2 Streaming?
Developers should choose it when a robot needs an ongoing multimodal interaction with low-latency responses, such as continuous work-area inspection, warehouse coordination, multi-robot task management, or audio- and video-aware assistance. A conventional endpoint may be better for occasional analysis, while a dedicated low-level controller is more appropriate for precise or safety-critical actuation.


Sources 6
Provider

About Google DeepMind