Model catalog

Google DeepMind Models

Browse the AI models associated with Google DeepMind. Compare current and historical models by family, capabilities, context window, availability and intended use.

41 models tracked
41 Total models
19 Model families
7 Model types
32 Current / accessible
All models

Google DeepMind model catalog

Browser automation, visual UI interaction, repetitive web workflows, form filling, and user-interface testing

Type Multimodal
Context 128K
Reasoning 7/10
Speed 6/10
Multimodal Image input Tool use
Status

Legacy preview; currently documented and accessible, with newer Gemini 3.x models recommended for new computer-use applications

Input $1.25 per 1M tokens for prompts up to 200K tokens; $2.50 per 1M tokens for prompts over 200K tokens
Output $10.00 per 1M tokens for prompts up to 200K tokens; $15.00 per 1M tokens for prompts over 200K tokens
View model →

Large-scale multimodal processing, low-latency reasoning, coding, data extraction, tool-using agents, and applications requiring a very large context window.

Type General Purpose
Context 1.05M
Reasoning 8/10
Speed 9/10
Multimodal Image input Audio input
Status

Stable and currently served through the Gemini API with restricted access for users who have actively used Gemini 2.5 models; not deprecated; no shutdown date announced.

Input $0.30 per 1M text/image/video tokens; $1.00 per 1M audio tokens; batch input $0.15 per 1M text/image/video tokens.
Output $2.50 per 1M tokens, including thinking tokens.
View model →

High-volume classification, simple extraction, lightweight multimodal analysis, routing, tagging, summarization, and extremely latency-sensitive applications.

Type Lightweight
Context 1.05M
Reasoning 6/10
Speed 10/10
Multimodal Image input Audio input
Status

Stable; currently served through the Gemini API with access limited to users who have actively used Gemini 2.5 models. Google recommends newer models for new projects.

Input $0.10 per 1M text/image/video input tokens; $0.30 per 1M audio input tokens. Batch/Flex: $0.05 per 1M text/image/video tokens and $0.15 per 1M audio tokens.
Output $0.40 per 1M output tokens. Batch/Flex: $0.20 per 1M output tokens.
View model →

Real-time voice and video agents, speech-to-speech assistants, interactive customer support, tutoring, coaching, and multimodal Live API applications

Type Multimodal
Context 131K
Reasoning 7/10
Speed 9/10
Multimodal Image input Audio input
Status

Preview; currently listed by Google with limited access for users who have actively used Gemini 2.5 models; no shutdown date announced

Input $0.50 per 1M text tokens; $3.00 per 1M audio or video tokens
Output $2.00 per 1M text tokens; $12.00 per 1M audio tokens
View model →

Low-latency controllable text-to-speech, voice assistants, narration, read-aloud features, and multi-speaker audio generation

Type Other
Context 8K
Reasoning 2/10
Speed 8/10
Media output
Status

Preview; currently available with limited access conditions

Input $0.50 per 1M input text tokens; batch: $0.25 per 1M input text tokens
Output $10.00 per 1M output audio tokens; batch: $5.00 per 1M output audio tokens
View model →
Google DeepMind logo
Gemini 2.5

Gemini 2.5 Pro

Advanced coding, complex reasoning, mathematics, STEM analysis, long documents, large codebases, multimodal analysis, and tool-using agents.

Type Reasoning
Context 1.05M
Reasoning 9/10
Speed 6/10
Multimodal Image input Audio input
Status

Stable and generally available through the Gemini API; access is currently limited to users who have actively used Gemini 2.5 models. Google states that the model is not deprecated and will continue to be served until further notice.

Input $1.25 per 1M tokens for prompts up to 200K tokens; $2.50 per 1M tokens for prompts over 200K tokens. Batch/Flex rates are $0.625/$1.25 respectively.
Output $10.00 per 1M tokens for prompts up to 200K tokens; $15.00 per 1M tokens for prompts over 200K tokens. Batch/Flex rates are $5.00/$7.50 respectively. Prices include thinking tokens.
View model →

High-fidelity single-speaker and multi-speaker narration, audiobooks, podcasts, professional voiceovers, and scripted creative audio

Type Other
Context 8K
Reasoning 1/10
Speed 6/10
Media output
Status

Preview; currently listed and accessible through the Gemini API, with restricted access to Gemini 2.5 models

Input $1.00 per 1M text tokens standard; $0.50 per 1M text tokens batch
Output $20.00 per 1M audio tokens standard; $10.00 per 1M audio tokens batch
View model →
Google DeepMind logo
Gemini 2.5 Flash Image

Nano Banana

Fast image generation, conversational image editing, image transformation, and high-volume visual workflows

Type Multimodal
Context 66K
Reasoning 3/10
Speed 8/10
Multimodal Image input Media output
Status

Deprecated; scheduled to shut down on 2026-10-02

Input $0.30 per 1 million text or image input tokens; Batch/Flex $0.15; Priority $0.54
Output $0.039 per generated image; Batch/Flex $0.0195; Priority $0.0702
View model →

Agentic workflows, everyday coding, reasoning and planning, multimodal analysis, long-context document work, and cost-sensitive tool-using applications

Type Multimodal
Context 1.05M
Reasoning 9/10
Speed 9/10
Multimodal Image input Audio input
Status

Preview

Input $0.50 per 1 million input tokens
Output $3 per 1 million output tokens
View model →

Complex reasoning, advanced software engineering, long-context multimodal analysis, and agentic workflows requiring reliable tool use

Type Reasoning
Context 1.05M
Reasoning 9/10
Speed 7/10
Multimodal Image input Audio input
Status

Preview

Input $2.00 per 1M tokens for prompts up to 200K tokens; $4.00 per 1M tokens for prompts over 200K tokens
Output $12.00 per 1M tokens for prompts up to 200K tokens; $18.00 per 1M tokens for prompts over 200K tokens, including thinking tokens
View model →

Fast multimodal applications, coding assistants, long-context document and video analysis, tool-using agents, enterprise workflows, and rapid agentic execution

Type General Purpose
Context 1.05M
Reasoning 8/10
Speed 9/10
Multimodal Image input Audio input
Status

Stable; previous-generation Flash model; currently accessible; no shutdown date announced

Input $1.50 per 1 million input tokens
Output $7.50 per 1 million output tokens
View model →

Coding, software engineering, tool-using agents, long-context multimodal analysis, structured extraction and high-throughput enterprise workflows

Type General Purpose
Context 1.05M
Reasoning 8/10
Speed 9/10
Multimodal Image input Audio input
Status

Generally available; previous-generation Flash model, currently supported

Input $0.75 per 1M tokens through 2026-12-31; $1.50 per 1M tokens starting 2027-01-01
Output $3.75 per 1M tokens, including thinking tokens, through 2026-12-31; $7.50 per 1M tokens starting 2027-01-01
View model →

Long-horizon software engineering, autonomous agents, multimodal document workflows, enterprise knowledge work, and tool-using applications.

Type Multimodal
Context 1.05M
Reasoning 9/10
Speed 9/10
Multimodal Image input Audio input
Status

Generally available

Input $0.75 per 1 million tokens through December 31, 2026; $1.50 per 1 million tokens from January 1, 2027. Batch/Flex introductory input price: $0.375 per 1 million tokens.
Output $3.75 per 1 million tokens, including thinking tokens, through December 31, 2026; $7.50 per 1 million tokens from January 1, 2027. Batch/Flex introductory output price: $1.875 per 1 million tokens.
View model →

High-volume translation, classification, extraction, summarization, document processing, and lightweight tool-using agent workflows

Type Lightweight
Context 1.05M
Reasoning 7/10
Speed 10/10
Multimodal Image input Audio input
Status

Deprecated; generally available and accessible until scheduled shutdown on May 7, 2027

Input $0.25 per 1M text/image/video tokens; $0.50 per 1M audio tokens; batch/flex: $0.125 per 1M text/image/video tokens and $0.25 per 1M audio tokens
Output $1.50 per 1M tokens standard; $0.75 per 1M tokens for batch/flex inference
View model →

Low-latency voice agents, real-time dialogue, multimodal live sessions, and interactive audio applications

Type Multimodal
Context 131K
Reasoning 7/10
Speed 9/10
Multimodal Image input Audio input
Status

Legacy preview; currently accessible; Google recommends Gemini 3.8 Live for most new low-latency voice-agent deployments

Input $0.75 per 1M text tokens; $3.00 per 1M audio tokens or $0.005 per audio minute; $1.00 per 1M image/video tokens or $0.002 per image/video minute
Output $4.50 per 1M text tokens; $12.00 per 1M audio tokens or $0.018 per audio minute
View model →
Google DeepMind logo
Gemini 3.1 Flash Audio

Gemini 3.1 Flash TTS

Controllable expressive speech, narration, accessibility, scripted audio, and multi-speaker TTS prototypes

Type Other
Context 8K
Reasoning 2/10
Speed 8/10
Media output Streaming
Status

Legacy preview; currently accessible; no shutdown date announced

Input $1.00 per 1 million text tokens standard; $0.50 per 1 million text tokens batch
Output $20.00 per 1 million audio tokens standard; $10.00 per 1 million audio tokens batch
View model →
Google DeepMind logo
Gemini 3.1 Flash Lite Image

Nano Banana 2 Lite

Fast, low-cost 1K image generation and editing, rapid visual prototyping, interactive applications, high-volume image variations, storyboarding, and lightweight creative workflows

Type Multimodal
Context 66K
Reasoning 3/10
Speed 10/10
Multimodal Image input Video input
Status

Generally available; scheduled retirement June 28, 2027 or later

Input $0.25 per 1 million input tokens for text, image, and video input; batch pricing $0.125 per 1 million input tokens
Output $1.50 per 1 million text/thought output tokens and $30 per 1 million image output tokens; approximately $0.0336 per 1K image; batch image output pricing $15 per 1 million image tokens
View model →

Agentic workflows, coding agents, long-context multimodal analysis, tool use, and scaled production applications

Type General Purpose
Context 1.05M
Reasoning 9/10
Speed 8/10
Multimodal Image input Audio input
Status

Generally available; stable

Input $1.50 per 1M tokens standard; $0.75 per 1M tokens Batch and Flex; $2.70 per 1M tokens Priority
Output $9.00 per 1M tokens standard; $4.50 per 1M tokens Batch and Flex; $16.20 per 1M tokens Priority
View model →

High-volume, latency-sensitive agentic workflows, document parsing, translation, classification, data extraction, search-backed applications, and multimodal sub-agents.

Type Lightweight
Context 1.05M
Reasoning 8/10
Speed 10/10
Multimodal Image input Audio input
Status

General availability

Input $0.30 per 1 million tokens for standard text, image, video, and audio input; $0.15 per 1 million input tokens for Batch and Flex; $0.54 per 1 million input tokens for Priority.
Output $2.50 per 1 million output tokens for standard usage; $1.25 per 1 million output tokens for Batch and Flex; $4.50 per 1 million output tokens for Priority.
View model →

Low-latency, real-time speech-to-speech translation for calls, meetings, travel, customer support, and multilingual voice applications

Type Multimodal
Context 131K
Reasoning 2/10
Speed 10/10
Audio input Media output Streaming
Status

Preview

Input $3.50 per 1 million input audio tokens, approximately $0.0053 per minute
Output $21.00 per 1 million output audio tokens, approximately $0.0315 per minute
View model →
Google DeepMind logo
Gemini 3.5 Audio

Gemini 3.5 Transcribe

Pre-recorded audio transcription, multilingual speech recognition, speaker-labeled transcripts, timestamped transcripts, smart dictation, and domain-specific vocabulary

Type Other
Context 96K
Reasoning 2/10
Speed 9/10
Multimodal Audio input
Status

Generally available (GA)

Input $2.00 per 1M audio input tokens, approximately $0.003 per minute; free tier available
Output $12.00 per 1M text output tokens, approximately $0.002 per minute; free tier available
View model →

High-volume text-to-speech production, low-latency voice-agent cascades, read-aloud applications, voice replication, and everyday single-speaker speech

Type Other
Context 8K
Reasoning 1/10
Speed 9/10
Media output Streaming
Status

Generally available

Input $0.50 per 1 million text tokens through December 31, 2026; $1.00 per 1 million text tokens starting January 1, 2027. Batch and Flex input: $0.25 through December 31, 2026; $0.50 starting January 1, 2027.
Output $6.00 per 1 million audio tokens through December 31, 2026; $12.00 per 1 million audio tokens starting January 1, 2027. Standard equivalent: approximately $0.0015 per 10 seconds of audio through December 31, 2026.
View model →

Studio-quality narration, audiobooks, expressive voice acting, complex multi-speaker dialogue, regional accents, difficult pronunciations, long-form narration, voice design, and voice replication.

Type Other
Context 8K
Reasoning 1/10
Speed 8/10
Media output Streaming
Status

Generally available (GA); no shutdown date announced

Input $0.50 per 1 million text input tokens through December 31, 2026; $1.00 per 1 million text input tokens from January 1, 2027
Output $9.00 per 1 million audio output tokens through December 31, 2026; $18.00 per 1 million audio output tokens from January 1, 2027. Equivalent to $0.00225 per 10 seconds of audio through December 31, 2026 and $0.0045 per 10 seconds thereafter.
View model →

Low-latency voice agents, real-time audio-to-audio dialogue, multimodal assistants, and interactive tool-using applications

Type Multimodal
Context 131K
Reasoning 7/10
Speed 10/10
Multimodal Image input Audio input
Status

Stable; generally available

Input $0.75 per 1M text tokens; $3.00 per 1M audio tokens or approximately $0.005 per minute; $1.00 per 1M image/video tokens or approximately $0.002 per minute
Output $4.50 per 1M text tokens; $12.00 per 1M audio tokens or approximately $0.018 per minute
View model →

Complex real-time voice agents, multi-step problem solving, asynchronous tool workflows, technical support, travel coordination, and spoken STEM or coding tutoring

Type Reasoning
Context 131K
Reasoning 9/10
Speed 7/10
Multimodal Image input Audio input
Status

Stable; generally available

Input $0.75 per 1M text tokens; $3.00 per 1M audio tokens or $0.005 per audio minute; $1.00 per 1M image/video tokens or $0.002 per minute
Output $4.50 per 1M text tokens; $12.00 per 1M audio tokens or $0.018 per audio minute
View model →
Google DeepMind logo
Gemini Deep Research

Gemini Deep Research

Autonomous market research, due diligence, literature reviews, competitive analysis, source-heavy investigations, and cited research reports

Type Other
Context 1.05M
Reasoning 8/10
Speed 8/10
Multimodal Image input Audio input
Status

Preview

Input Variable pay-as-you-go pricing based on underlying model usage and tools; no fixed per-input-token price is specified for the agent.
Output Variable pay-as-you-go pricing based on research depth and generated output; typical total cost is estimated at approximately $1–$3 per standard task.
View model →
Google DeepMind logo
Gemini Deep Research

Gemini Deep Research Max

Comprehensive market research, competitive analysis, due diligence, literature reviews, and source-rich investigative reports

Type Reasoning
Context 1.05M
Reasoning 10/10
Speed 3/10
Multimodal Image input Audio input
Status

Preview; currently available through the Interactions API in the Gemini API and Google AI Studio

Input Pay-as-you-go based on underlying Gemini model inference and tool usage; no fixed model-specific input rate published
Output Pay-as-you-go based on underlying Gemini model inference and tool usage; no fixed model-specific output rate published
View model →
Google DeepMind logo
Gemini Embedding

Gemini Embedding

Text semantic search, RAG retrieval, document matching, classification, clustering, and recommendation systems

Type Other
Context 2K
Reasoning 1/10
Speed 8/10
Status

Current, scheduled for shutdown on 2028-05-14

Input $0.00015 per 1,000 input tokens for online Vertex AI requests; $0.00012 per 1,000 input tokens for batch requests
Output No charge for embedding output
View model →
Google DeepMind logo
Gemini Embedding

Gemini Embedding 2

Cross-modal semantic search, multimodal RAG, vector retrieval, recommendations, classification, clustering, and indexing mixed text and media collections.

Type Embedding
Context 8K
Reasoning 2/10
Speed 7/10
Multimodal Image input Audio input
Status

Generally available

Input Standard paid tier: text $0.20 per 1M tokens; images $0.45 per 1M tokens or $0.00012 per image; audio $6.50 per 1M tokens or $0.00016 per second; video $12.00 per 1M tokens or $0.00079 per frame. Batch pricing is 50% lower.
Output No separate output-token price; the model returns embeddings. Output is included in the input-modality pricing structure.
View model →
Google DeepMind logo
Gemini Image

Nano Banana 2

Fast, high-volume image generation and editing, visual iteration, marketing assets, diagrams, infographics, localization, and applications requiring image-search grounding

Type Multimodal
Context 131K
Reasoning 7/10
Speed 9/10
Multimodal Image input Video input
Status

Generally available; the stable Gemini API model is gemini-3.1-flash-image

Input $0.50 per 1 million tokens for text/image input; batch input $0.25 per 1 million tokens
Output $3 per 1 million text/thinking tokens and $60 per 1 million image-output tokens; approximately $0.067 per 1K image, $0.101 per 2K image, and $0.151 per 4K image; batch image output $30 per 1 million image-output tokens
View model →
Google DeepMind logo
Gemini Image

Nano Banana Pro

Professional image generation and editing, complex compositions, product mockups, infographics, branded creative, multilingual localization, and high-fidelity visual prototyping

Type Multimodal
Context 66K
Reasoning 8/10
Speed 6/10
Multimodal Image input Media output
Status

Current stable model

Input $2.00 per 1M text/image input tokens; approximately $0.0011 per input image
Output $12.00 per 1M text and thinking tokens; $120.00 per 1M image tokens, equivalent to approximately $0.134 per 1K/2K image and $0.24 per 4K image
View model →

Fast text-to-video, image-to-video, conversational video editing, video extension, interpolation, marketing content, and short-form cinematic production

Type Multimodal
Context 1.05M
Reasoning 5/10
Speed 9/10
Multimodal Image input Video input
Status

Current; stable model available as gemini-omni-1.1-flash, with gemini-omni-flash-preview also documented

Input $1.50 per 1M input tokens for text, image, video, and audio inputs under standard pricing
Output $17.50 per 1M video output tokens; $9.00 per 1M text output tokens where applicable
View model →
Google DeepMind logo
Gemini Robotics ER

Gemini Robotics ER 2

High-level robot planning, spatial reasoning, video progress tracking, tool orchestration, and multi-robot collaboration

Type Reasoning
Context 131K
Reasoning 8/10
Speed 4/10
Multimodal Image input Audio input
Status

Public preview

Input $1.00 per 1 million tokens for the standard preview endpoint through December 31, 2026; $2.00 per 1 million tokens starting January 1, 2027
Output $5.00 per 1 million tokens, including reasoning tokens, through December 31, 2026; $10.00 per 1 million tokens starting January 1, 2027
View model →

Low-latency robotic agents, continuous audio/video monitoring, function-based robot orchestration, warehouse workflows, and multi-robot coordination

Type Reasoning
Context 131K
Reasoning 8/10
Speed 9/10
Multimodal Image input Audio input
Status

Public preview

Input $1.00 per 1M tokens through December 31, 2026; $2.00 per 1M tokens from January 1, 2027
Output $5.00 per 1M tokens through December 31, 2026; $10.00 per 1M tokens from January 1, 2027
View model →

Full-length AI-generated songs, vocal music, instrumental arrangements, songwriting experiments, soundtracks, and image-inspired music creation

Type Other
Context 131K
Reasoning 1/10
Speed 7/10
Multimodal Image input Media output
Status

Generally available

Input $0.08 per full song; no free tier
Output $0.08 per full song
View model →

Interactive instrumental music generation, live musical improvisation, prompt-driven DJ tools, MIDI-controlled experiences, and real-time creative audio applications

Type Other
Reasoning 1/10
Speed 9/10
Media output Streaming
Status

Experimental and currently documented through the Gemini API as lyria-realtime-exp

Input Not publicly listed for this model
Output Not publicly listed for this model
View model →

Full-length AI music, soundtrack creation, songwriting, structured compositions, advertising audio, games, and creative production workflows

Type Other
Reasoning 1/10
Speed 7/10
Multimodal Image input Media output
Status

Public Preview

Input $0.08 per full-song generation
View model →
Google DeepMind logo
Veo 3.1

Veo 3.1

Cinematic text-to-video and image-to-video generation, short-form storytelling, storyboarding, advertising concepts, visual effects exploration, and creative previsualization with synchronized audio

Type Multimodal
Context 1K
Reasoning 1/10
Speed 5/10
Multimodal Image input Video input
Status

Preview; currently accessible through Google APIs and Google products

Input $0.40 per second for standard Veo 3.1 video with audio at 720p or 1080p; $0.60 per second at 4K
Output Generated video with native audio; pricing is charged per generated video second
View model →

Fast, high-volume video generation, creative iteration, social content, advertising concepts, and automated production workflows

Type Video Generation
Context 1K
Speed 9/10
Multimodal Image input Video input
Status

Generally available on Vertex AI; preview model on the Gemini API

Input $0.10 per video second at 720p; $0.12 per video second at 1080p; $0.30 per video second at 4K where supported
Output Video with natively generated audio, priced per generated video second
View model →

High-volume text-to-video and image-to-video generation, rapid creative iteration, social content, advertising variations, and cost-sensitive production workflows

Type Other
Reasoning 1/10
Speed 8/10
Multimodal Image input Media output
Status

Preview; currently available through the Gemini API and Google Cloud Vertex AI

Input $0.05 per second for 720p video with audio; $0.08 per second for 1080p video with audio
Output Video with native audio, priced per generated-video second
View model →