Browser automation, visual UI interaction, repetitive web workflows, form filling, and user-interface testing
Type
Multimodal
Context
128K
Reasoning
7/10
Speed
6/10
Multimodal
Image input
Tool use
Status
Legacy preview; currently documented and accessible, with newer Gemini 3.x models recommended for new computer-use applications
Input
$1.25 per 1M tokens for prompts up to 200K tokens; $2.50 per 1M tokens for prompts over 200K tokens
Output
$10.00 per 1M tokens for prompts up to 200K tokens; $15.00 per 1M tokens for prompts over 200K tokens
View model
→
Large-scale multimodal processing, low-latency reasoning, coding, data extraction, tool-using agents, and applications requiring a very large context window.
Type
General Purpose
Context
1.05M
Reasoning
8/10
Speed
9/10
Multimodal
Image input
Audio input
Status
Stable and currently served through the Gemini API with restricted access for users who have actively used Gemini 2.5 models; not deprecated; no shutdown date announced.
Input
$0.30 per 1M text/image/video tokens; $1.00 per 1M audio tokens; batch input $0.15 per 1M text/image/video tokens.
Output
$2.50 per 1M tokens, including thinking tokens.
View model
→
High-volume classification, simple extraction, lightweight multimodal analysis, routing, tagging, summarization, and extremely latency-sensitive applications.
Type
Lightweight
Context
1.05M
Reasoning
6/10
Speed
10/10
Multimodal
Image input
Audio input
Status
Stable; currently served through the Gemini API with access limited to users who have actively used Gemini 2.5 models. Google recommends newer models for new projects.
Input
$0.10 per 1M text/image/video input tokens; $0.30 per 1M audio input tokens. Batch/Flex: $0.05 per 1M text/image/video tokens and $0.15 per 1M audio tokens.
Output
$0.40 per 1M output tokens. Batch/Flex: $0.20 per 1M output tokens.
View model
→
Real-time voice and video agents, speech-to-speech assistants, interactive customer support, tutoring, coaching, and multimodal Live API applications
Type
Multimodal
Context
131K
Reasoning
7/10
Speed
9/10
Multimodal
Image input
Audio input
Status
Preview; currently listed by Google with limited access for users who have actively used Gemini 2.5 models; no shutdown date announced
Input
$0.50 per 1M text tokens; $3.00 per 1M audio or video tokens
Output
$2.00 per 1M text tokens; $12.00 per 1M audio tokens
View model
→
Low-latency controllable text-to-speech, voice assistants, narration, read-aloud features, and multi-speaker audio generation
Type
Other
Context
8K
Reasoning
2/10
Speed
8/10
Media output
Status
Preview; currently available with limited access conditions
Input
$0.50 per 1M input text tokens; batch: $0.25 per 1M input text tokens
Output
$10.00 per 1M output audio tokens; batch: $5.00 per 1M output audio tokens
View model
→
Advanced coding, complex reasoning, mathematics, STEM analysis, long documents, large codebases, multimodal analysis, and tool-using agents.
Type
Reasoning
Context
1.05M
Reasoning
9/10
Speed
6/10
Multimodal
Image input
Audio input
Status
Stable and generally available through the Gemini API; access is currently limited to users who have actively used Gemini 2.5 models. Google states that the model is not deprecated and will continue to be served until further notice.
Input
$1.25 per 1M tokens for prompts up to 200K tokens; $2.50 per 1M tokens for prompts over 200K tokens. Batch/Flex rates are $0.625/$1.25 respectively.
Output
$10.00 per 1M tokens for prompts up to 200K tokens; $15.00 per 1M tokens for prompts over 200K tokens. Batch/Flex rates are $5.00/$7.50 respectively. Prices include thinking tokens.
View model
→
High-fidelity single-speaker and multi-speaker narration, audiobooks, podcasts, professional voiceovers, and scripted creative audio
Type
Other
Context
8K
Reasoning
1/10
Speed
6/10
Media output
Status
Preview; currently listed and accessible through the Gemini API, with restricted access to Gemini 2.5 models
Input
$1.00 per 1M text tokens standard; $0.50 per 1M text tokens batch
Output
$20.00 per 1M audio tokens standard; $10.00 per 1M audio tokens batch
View model
→
Fast image generation, conversational image editing, image transformation, and high-volume visual workflows
Type
Multimodal
Context
66K
Reasoning
3/10
Speed
8/10
Multimodal
Image input
Media output
Status
Deprecated; scheduled to shut down on 2026-10-02
Input
$0.30 per 1 million text or image input tokens; Batch/Flex $0.15; Priority $0.54
Output
$0.039 per generated image; Batch/Flex $0.0195; Priority $0.0702
View model
→
Agentic workflows, everyday coding, reasoning and planning, multimodal analysis, long-context document work, and cost-sensitive tool-using applications
Type
Multimodal
Context
1.05M
Reasoning
9/10
Speed
9/10
Multimodal
Image input
Audio input
Input
$0.50 per 1 million input tokens
Output
$3 per 1 million output tokens
View model
→
Complex reasoning, advanced software engineering, long-context multimodal analysis, and agentic workflows requiring reliable tool use
Type
Reasoning
Context
1.05M
Reasoning
9/10
Speed
7/10
Multimodal
Image input
Audio input
Input
$2.00 per 1M tokens for prompts up to 200K tokens; $4.00 per 1M tokens for prompts over 200K tokens
Output
$12.00 per 1M tokens for prompts up to 200K tokens; $18.00 per 1M tokens for prompts over 200K tokens, including thinking tokens
View model
→
Fast multimodal applications, coding assistants, long-context document and video analysis, tool-using agents, enterprise workflows, and rapid agentic execution
Type
General Purpose
Context
1.05M
Reasoning
8/10
Speed
9/10
Multimodal
Image input
Audio input
Status
Stable; previous-generation Flash model; currently accessible; no shutdown date announced
Input
$1.50 per 1 million input tokens
Output
$7.50 per 1 million output tokens
View model
→
Coding, software engineering, tool-using agents, long-context multimodal analysis, structured extraction and high-throughput enterprise workflows
Type
General Purpose
Context
1.05M
Reasoning
8/10
Speed
9/10
Multimodal
Image input
Audio input
Status
Generally available; previous-generation Flash model, currently supported
Input
$0.75 per 1M tokens through 2026-12-31; $1.50 per 1M tokens starting 2027-01-01
Output
$3.75 per 1M tokens, including thinking tokens, through 2026-12-31; $7.50 per 1M tokens starting 2027-01-01
View model
→
Long-horizon software engineering, autonomous agents, multimodal document workflows, enterprise knowledge work, and tool-using applications.
Type
Multimodal
Context
1.05M
Reasoning
9/10
Speed
9/10
Multimodal
Image input
Audio input
Status
Generally available
Input
$0.75 per 1 million tokens through December 31, 2026; $1.50 per 1 million tokens from January 1, 2027. Batch/Flex introductory input price: $0.375 per 1 million tokens.
Output
$3.75 per 1 million tokens, including thinking tokens, through December 31, 2026; $7.50 per 1 million tokens from January 1, 2027. Batch/Flex introductory output price: $1.875 per 1 million tokens.
View model
→
High-volume translation, classification, extraction, summarization, document processing, and lightweight tool-using agent workflows
Type
Lightweight
Context
1.05M
Reasoning
7/10
Speed
10/10
Multimodal
Image input
Audio input
Status
Deprecated; generally available and accessible until scheduled shutdown on May 7, 2027
Input
$0.25 per 1M text/image/video tokens; $0.50 per 1M audio tokens; batch/flex: $0.125 per 1M text/image/video tokens and $0.25 per 1M audio tokens
Output
$1.50 per 1M tokens standard; $0.75 per 1M tokens for batch/flex inference
View model
→
Low-latency voice agents, real-time dialogue, multimodal live sessions, and interactive audio applications
Type
Multimodal
Context
131K
Reasoning
7/10
Speed
9/10
Multimodal
Image input
Audio input
Status
Legacy preview; currently accessible; Google recommends Gemini 3.8 Live for most new low-latency voice-agent deployments
Input
$0.75 per 1M text tokens; $3.00 per 1M audio tokens or $0.005 per audio minute; $1.00 per 1M image/video tokens or $0.002 per image/video minute
Output
$4.50 per 1M text tokens; $12.00 per 1M audio tokens or $0.018 per audio minute
View model
→
Controllable expressive speech, narration, accessibility, scripted audio, and multi-speaker TTS prototypes
Type
Other
Context
8K
Reasoning
2/10
Speed
8/10
Media output
Streaming
Status
Legacy preview; currently accessible; no shutdown date announced
Input
$1.00 per 1 million text tokens standard; $0.50 per 1 million text tokens batch
Output
$20.00 per 1 million audio tokens standard; $10.00 per 1 million audio tokens batch
View model
→
Gemini 3.1 Flash Lite Image
Fast, low-cost 1K image generation and editing, rapid visual prototyping, interactive applications, high-volume image variations, storyboarding, and lightweight creative workflows
Type
Multimodal
Context
66K
Reasoning
3/10
Speed
10/10
Multimodal
Image input
Video input
Status
Generally available; scheduled retirement June 28, 2027 or later
Input
$0.25 per 1 million input tokens for text, image, and video input; batch pricing $0.125 per 1 million input tokens
Output
$1.50 per 1 million text/thought output tokens and $30 per 1 million image output tokens; approximately $0.0336 per 1K image; batch image output pricing $15 per 1 million image tokens
View model
→
Agentic workflows, coding agents, long-context multimodal analysis, tool use, and scaled production applications
Type
General Purpose
Context
1.05M
Reasoning
9/10
Speed
8/10
Multimodal
Image input
Audio input
Status
Generally available; stable
Input
$1.50 per 1M tokens standard; $0.75 per 1M tokens Batch and Flex; $2.70 per 1M tokens Priority
Output
$9.00 per 1M tokens standard; $4.50 per 1M tokens Batch and Flex; $16.20 per 1M tokens Priority
View model
→
High-volume, latency-sensitive agentic workflows, document parsing, translation, classification, data extraction, search-backed applications, and multimodal sub-agents.
Type
Lightweight
Context
1.05M
Reasoning
8/10
Speed
10/10
Multimodal
Image input
Audio input
Status
General availability
Input
$0.30 per 1 million tokens for standard text, image, video, and audio input; $0.15 per 1 million input tokens for Batch and Flex; $0.54 per 1 million input tokens for Priority.
Output
$2.50 per 1 million output tokens for standard usage; $1.25 per 1 million output tokens for Batch and Flex; $4.50 per 1 million output tokens for Priority.
View model
→
Low-latency, real-time speech-to-speech translation for calls, meetings, travel, customer support, and multilingual voice applications
Type
Multimodal
Context
131K
Reasoning
2/10
Speed
10/10
Audio input
Media output
Streaming
Input
$3.50 per 1 million input audio tokens, approximately $0.0053 per minute
Output
$21.00 per 1 million output audio tokens, approximately $0.0315 per minute
View model
→
Pre-recorded audio transcription, multilingual speech recognition, speaker-labeled transcripts, timestamped transcripts, smart dictation, and domain-specific vocabulary
Type
Other
Context
96K
Reasoning
2/10
Speed
9/10
Multimodal
Audio input
Status
Generally available (GA)
Input
$2.00 per 1M audio input tokens, approximately $0.003 per minute; free tier available
Output
$12.00 per 1M text output tokens, approximately $0.002 per minute; free tier available
View model
→
High-volume text-to-speech production, low-latency voice-agent cascades, read-aloud applications, voice replication, and everyday single-speaker speech
Type
Other
Context
8K
Reasoning
1/10
Speed
9/10
Media output
Streaming
Status
Generally available
Input
$0.50 per 1 million text tokens through December 31, 2026; $1.00 per 1 million text tokens starting January 1, 2027. Batch and Flex input: $0.25 through December 31, 2026; $0.50 starting January 1, 2027.
Output
$6.00 per 1 million audio tokens through December 31, 2026; $12.00 per 1 million audio tokens starting January 1, 2027. Standard equivalent: approximately $0.0015 per 10 seconds of audio through December 31, 2026.
View model
→
Studio-quality narration, audiobooks, expressive voice acting, complex multi-speaker dialogue, regional accents, difficult pronunciations, long-form narration, voice design, and voice replication.
Type
Other
Context
8K
Reasoning
1/10
Speed
8/10
Media output
Streaming
Status
Generally available (GA); no shutdown date announced
Input
$0.50 per 1 million text input tokens through December 31, 2026; $1.00 per 1 million text input tokens from January 1, 2027
Output
$9.00 per 1 million audio output tokens through December 31, 2026; $18.00 per 1 million audio output tokens from January 1, 2027. Equivalent to $0.00225 per 10 seconds of audio through December 31, 2026 and $0.0045 per 10 seconds thereafter.
View model
→
Low-latency voice agents, real-time audio-to-audio dialogue, multimodal assistants, and interactive tool-using applications
Type
Multimodal
Context
131K
Reasoning
7/10
Speed
10/10
Multimodal
Image input
Audio input
Status
Stable; generally available
Input
$0.75 per 1M text tokens; $3.00 per 1M audio tokens or approximately $0.005 per minute; $1.00 per 1M image/video tokens or approximately $0.002 per minute
Output
$4.50 per 1M text tokens; $12.00 per 1M audio tokens or approximately $0.018 per minute
View model
→
Complex real-time voice agents, multi-step problem solving, asynchronous tool workflows, technical support, travel coordination, and spoken STEM or coding tutoring
Type
Reasoning
Context
131K
Reasoning
9/10
Speed
7/10
Multimodal
Image input
Audio input
Status
Stable; generally available
Input
$0.75 per 1M text tokens; $3.00 per 1M audio tokens or $0.005 per audio minute; $1.00 per 1M image/video tokens or $0.002 per minute
Output
$4.50 per 1M text tokens; $12.00 per 1M audio tokens or $0.018 per audio minute
View model
→
Autonomous market research, due diligence, literature reviews, competitive analysis, source-heavy investigations, and cited research reports
Type
Other
Context
1.05M
Reasoning
8/10
Speed
8/10
Multimodal
Image input
Audio input
Input
Variable pay-as-you-go pricing based on underlying model usage and tools; no fixed per-input-token price is specified for the agent.
Output
Variable pay-as-you-go pricing based on research depth and generated output; typical total cost is estimated at approximately $1–$3 per standard task.
View model
→
Comprehensive market research, competitive analysis, due diligence, literature reviews, and source-rich investigative reports
Type
Reasoning
Context
1.05M
Reasoning
10/10
Speed
3/10
Multimodal
Image input
Audio input
Status
Preview; currently available through the Interactions API in the Gemini API and Google AI Studio
Input
Pay-as-you-go based on underlying Gemini model inference and tool usage; no fixed model-specific input rate published
Output
Pay-as-you-go based on underlying Gemini model inference and tool usage; no fixed model-specific output rate published
View model
→
Text semantic search, RAG retrieval, document matching, classification, clustering, and recommendation systems
Type
Other
Context
2K
Reasoning
1/10
Speed
8/10
Status
Current, scheduled for shutdown on 2028-05-14
Input
$0.00015 per 1,000 input tokens for online Vertex AI requests; $0.00012 per 1,000 input tokens for batch requests
Output
No charge for embedding output
View model
→
Cross-modal semantic search, multimodal RAG, vector retrieval, recommendations, classification, clustering, and indexing mixed text and media collections.
Type
Embedding
Context
8K
Reasoning
2/10
Speed
7/10
Multimodal
Image input
Audio input
Status
Generally available
Input
Standard paid tier: text $0.20 per 1M tokens; images $0.45 per 1M tokens or $0.00012 per image; audio $6.50 per 1M tokens or $0.00016 per second; video $12.00 per 1M tokens or $0.00079 per frame. Batch pricing is 50% lower.
Output
No separate output-token price; the model returns embeddings. Output is included in the input-modality pricing structure.
View model
→
Fast, high-volume image generation and editing, visual iteration, marketing assets, diagrams, infographics, localization, and applications requiring image-search grounding
Type
Multimodal
Context
131K
Reasoning
7/10
Speed
9/10
Multimodal
Image input
Video input
Status
Generally available; the stable Gemini API model is gemini-3.1-flash-image
Input
$0.50 per 1 million tokens for text/image input; batch input $0.25 per 1 million tokens
Output
$3 per 1 million text/thinking tokens and $60 per 1 million image-output tokens; approximately $0.067 per 1K image, $0.101 per 2K image, and $0.151 per 4K image; batch image output $30 per 1 million image-output tokens
View model
→
Professional image generation and editing, complex compositions, product mockups, infographics, branded creative, multilingual localization, and high-fidelity visual prototyping
Type
Multimodal
Context
66K
Reasoning
8/10
Speed
6/10
Multimodal
Image input
Media output
Status
Current stable model
Input
$2.00 per 1M text/image input tokens; approximately $0.0011 per input image
Output
$12.00 per 1M text and thinking tokens; $120.00 per 1M image tokens, equivalent to approximately $0.134 per 1K/2K image and $0.24 per 4K image
View model
→
Fast text-to-video, image-to-video, conversational video editing, video extension, interpolation, marketing content, and short-form cinematic production
Type
Multimodal
Context
1.05M
Reasoning
5/10
Speed
9/10
Multimodal
Image input
Video input
Status
Current; stable model available as gemini-omni-1.1-flash, with gemini-omni-flash-preview also documented
Input
$1.50 per 1M input tokens for text, image, video, and audio inputs under standard pricing
Output
$17.50 per 1M video output tokens; $9.00 per 1M text output tokens where applicable
View model
→
High-level robot planning, spatial reasoning, video progress tracking, tool orchestration, and multi-robot collaboration
Type
Reasoning
Context
131K
Reasoning
8/10
Speed
4/10
Multimodal
Image input
Audio input
Input
$1.00 per 1 million tokens for the standard preview endpoint through December 31, 2026; $2.00 per 1 million tokens starting January 1, 2027
Output
$5.00 per 1 million tokens, including reasoning tokens, through December 31, 2026; $10.00 per 1 million tokens starting January 1, 2027
View model
→
Low-latency robotic agents, continuous audio/video monitoring, function-based robot orchestration, warehouse workflows, and multi-robot coordination
Type
Reasoning
Context
131K
Reasoning
8/10
Speed
9/10
Multimodal
Image input
Audio input
Input
$1.00 per 1M tokens through December 31, 2026; $2.00 per 1M tokens from January 1, 2027
Output
$5.00 per 1M tokens through December 31, 2026; $10.00 per 1M tokens from January 1, 2027
View model
→
Full-length AI-generated songs, vocal music, instrumental arrangements, songwriting experiments, soundtracks, and image-inspired music creation
Type
Other
Context
131K
Reasoning
1/10
Speed
7/10
Multimodal
Image input
Media output
Status
Generally available
Input
$0.08 per full song; no free tier
Output
$0.08 per full song
View model
→
Interactive instrumental music generation, live musical improvisation, prompt-driven DJ tools, MIDI-controlled experiences, and real-time creative audio applications
Type
Other
Reasoning
1/10
Speed
9/10
Media output
Streaming
Status
Experimental and currently documented through the Gemini API as lyria-realtime-exp
Input
Not publicly listed for this model
Output
Not publicly listed for this model
View model
→
Generating short music or audio clips
Media output
View model
→
Full-length AI music, soundtrack creation, songwriting, structured compositions, advertising audio, games, and creative production workflows
Type
Other
Reasoning
1/10
Speed
7/10
Multimodal
Image input
Media output
Input
$0.08 per full-song generation
View model
→
Cinematic text-to-video and image-to-video generation, short-form storytelling, storyboarding, advertising concepts, visual effects exploration, and creative previsualization with synchronized audio
Type
Multimodal
Context
1K
Reasoning
1/10
Speed
5/10
Multimodal
Image input
Video input
Status
Preview; currently accessible through Google APIs and Google products
Input
$0.40 per second for standard Veo 3.1 video with audio at 720p or 1080p; $0.60 per second at 4K
Output
Generated video with native audio; pricing is charged per generated video second
View model
→
Fast, high-volume video generation, creative iteration, social content, advertising concepts, and automated production workflows
Type
Video Generation
Context
1K
Speed
9/10
Multimodal
Image input
Video input
Status
Generally available on Vertex AI; preview model on the Gemini API
Input
$0.10 per video second at 720p; $0.12 per video second at 1080p; $0.30 per video second at 4K where supported
Output
Video with natively generated audio, priced per generated video second
View model
→
High-volume text-to-video and image-to-video generation, rapid creative iteration, social content, advertising variations, and cost-sensitive production workflows
Type
Other
Reasoning
1/10
Speed
8/10
Multimodal
Image input
Media output
Status
Preview; currently available through the Gemini API and Google Cloud Vertex AI
Input
$0.05 per second for 720p video with audio; $0.08 per second for 1080p video with audio
Output
Video with native audio, priced per generated-video second
View model
→