Long-context analysis, enterprise agents, research, coding assistance, structured extraction and tool-enabled workflows
Type
Multimodal
Context
1M
Reasoning
9/10
Speed
9/10
Multimodal
Image input
Tool use
Input
$1.25 per 1M tokens below 200K prompt tokens; $2.50 per 1M tokens at or above 200K prompt tokens
Output
$2.50 per 1M tokens below 200K prompt tokens; $5.00 per 1M tokens at or above 200K prompt tokens
View model
→
Software engineering, codebase analysis, technical reasoning, long-context work, tool-using agents, document analysis, and structured workflow automation
Type
Coding
Context
500K
Reasoning
9/10
Speed
8/10
Multimodal
Image input
Tool use
Input
$2.00 per 1M tokens; $0.30 per 1M cached tokens; long context of 200K tokens or more: $4.00 input and $0.60 cached input per 1M tokens
Output
$6.00 per 1M tokens; long context of 200K tokens or more: $12.00 per 1M tokens
View model
→
Advanced software engineering, long-context reasoning, agentic tool use, research with web or X search, and professional knowledge work.
Type
Reasoning
Context
500K
Reasoning
9/10
Speed
8/10
Multimodal
Image input
Tool use
Status
Current; available through the xAI API, Grok Build, Cursor, and selected model gateways.
Input
$2.00 per 1 million input tokens below 200,000 prompt tokens; $0.50 per 1 million cached input tokens; $4.00 per 1 million input tokens above 200,000 prompt tokens; $1.00 per 1 million cached input tokens above 200,000 prompt tokens.
Output
$6.00 per 1 million output tokens below 200,000 prompt tokens; $12.00 per 1 million output tokens above 200,000 prompt tokens.
View model
→
Agentic coding, long-context software engineering, research, knowledge work, visual analysis, structured extraction, and tool-using workflows.
Type
Reasoning
Context
500K
Reasoning
9/10
Speed
8/10
Multimodal
Image input
Tool use
Status
Current and available; superseded as xAI's flagship by Grok 4.7 but still supported on the xAI API.
Input
$2.00 per 1M tokens below 200k prompt tokens; $4.00 per 1M tokens at 200k tokens or more. Cached input is $0.50 per 1M tokens below 200k and $1.00 per 1M tokens at 200k or more.
Output
$6.00 per 1M tokens below 200k prompt tokens; $12.00 per 1M tokens at 200k tokens or more.
View model
→
Low-latency coding, interactive development, and agentic workflows in Cursor or Grok Build
Type
General Purpose
Context
500K
Reasoning
9/10
Speed
10/10
Multimodal
Image input
Tool use
Status
Available as a faster-serving variant in Cursor and Grok Build; not available as a separate public xAI API model
Input
2x standard Grok 4.7 input-token rate where Fast pricing applies; standard Grok 4.7 API rate is $2 per 1M input tokens
Output
2x standard Grok 4.7 output-token rate where Fast pricing applies; standard Grok 4.7 API rate is $6 per 1M output tokens
View model
→
Deep research, parallel investigation, long-context analysis, and tool-assisted synthesis
Type
Reasoning
Context
1M
Reasoning
9/10
Speed
5/10
Multimodal
Image input
Tool use
Status
Beta; currently available through the xAI API
Input
$1.25 per 1M tokens; $2.50 per 1M tokens for requests at or above 200K input tokens
Output
$2.50 per 1M tokens; $5.00 per 1M tokens for requests at or above 200K input tokens
View model
→
Fast general-purpose text generation, image-aware analysis, coding assistance, structured extraction, tool-calling agents and large-context workflows
Type
General Purpose
Context
1M
Reasoning
6/10
Speed
9/10
Multimodal
Image input
Tool use
Input
$1.25 per 1M tokens; $0.20 per 1M cached tokens. Long-context requests at or above 200K prompt tokens: $2.50 per 1M input tokens; $0.40 per 1M cached tokens.
Output
$2.50 per 1M tokens; $5.00 per 1M tokens for long-context requests at or above 200K prompt tokens.
View model
→
Complex reasoning, coding, technical research, long-context document analysis, image understanding, structured responses, and tool-enabled agentic workflows
Type
Reasoning
Context
1M
Reasoning
8/10
Speed
8/10
Multimodal
Image input
Tool use
Input
$1.25 per 1M tokens for prompts under 200K tokens; $2.50 per 1M tokens for prompts at or above 200K tokens; cached input is $0.20 or $0.40 per 1M tokens respectively
Output
$2.50 per 1M tokens for prompts under 200K tokens; $5.00 per 1M tokens for prompts at or above 200K tokens
View model
→
Agentic coding, web development, debugging, software engineering workflows, MCP integrations, and fast tool-calling applications.
Type
Coding
Context
256K
Reasoning
8/10
Speed
9/10
Multimodal
Image input
Tool use
Status
Current; public beta
Input
$1.00 per 1M input tokens; $0.20 per 1M cached input tokens. Long-context pricing for prompts exceeding 200K tokens is $2.00 per 1M input tokens and $0.40 per 1M cached input tokens.
Output
$2.00 per 1M output tokens; $4.00 per 1M output tokens when the prompt exceeds 200K tokens.
View model
→
API-based text-to-image generation, image editing, visual prototyping, creative applications, and per-image billing
Type
Other
Context
1K
Reasoning
1/10
Speed
7/10
Multimodal
Image input
Media output
Status
Current and available through the xAI API as of September 24, 2026
Input
$0.002 per input image; text prompt pricing is not separately stated
Output
$0.02 per generated image for 1K and 2K resolution
View model
→
High-quality text-to-image generation, image editing, multi-reference compositing, marketing graphics, product imagery, typography, and design assets
Multimodal
Image input
Media output
Status
Generally available
Input
$0.01 per input image
Output
$0.04-$0.08 per generated image, depending on resolution and quality
View model
→
High-quality image generation and editing, realistic product and marketing imagery, detailed scenes, creative assets, and images requiring stronger text rendering or prompt adherence.
Type
Multimodal
Reasoning
1/10
Speed
7/10
Multimodal
Image input
Media output
Status
Currently available; scheduled for API retirement on November 2, 2026. After retirement, requests using the slug are served by Grok Imagine Image 2.0 with quality set to low.
Input
$0.01 per input image
Output
$0.05 per 1K image; $0.07 per 2K image
View model
→
Short-form text-to-video, image-to-video, reference-guided video, cinematic prototyping, marketing clips, and audiovisual creative workflows
Multimodal
Image input
Audio input
Status
Generally available
Input
$0.01 per image; preset audio input is free
Output
$0.08/sec at 480p; $0.14/sec at 720p; $0.25/sec at 1080p
View model
→
Low-cost text-to-video and image-to-video drafts, social clips, rapid creative iteration, and high-volume generation
Type
Lightweight
Reasoning
0/10
Speed
8/10
Multimodal
Image input
Media output
Status
Current and available through the xAI Imagine API
Input
$0.01 per image input; video input is charged per second where applicable
Output
$0.02 per second at 480p; $0.03 per second at 720p; $0.14 per second at 1080p
View model
→
Expressive speech synthesis, voice agents, narration, podcasts, audiobooks, accessibility, and interactive audio applications
Type
Other
Reasoning
1/10
Speed
8/10
Media output
Streaming
Input
$15.00 per 1 million input characters
Output
Included in character-based pricing; no separate output charge documented
Model page unavailable
Realtime voice agents, customer support, telephony, sales, multilingual conversations, and tool-enabled spoken workflows
Type
Multimodal
Reasoning
8/10
Speed
9/10
Multimodal
Audio input
Media output
Status
Current and available through the xAI Speech to Speech API
Input
$0.08 per minute of audio; $0.004 per text input
Output
$0.08 per minute of audio
View model
→
Batch and real-time multilingual speech transcription, dictation, voice assistants, accessibility, meetings, and customer-support audio
Type
Speech Recognition
Speed
8/10
Audio input
Streaming
Status
Currently accessible; original speech-to-text model; scheduled for deprecation in the coming weeks
Input
$0.10 per hour of audio for REST batch transcription; $0.20 per hour of audio for streaming
Output
Included in the transcription service price; output is text transcript data
View model
→
Batch and real-time transcription of multilingual, noisy, conversational, telephony, and voice-agent audio
Type
Other
Reasoning
0/10
Speed
9/10
Audio input
Streaming
Input
$0.10 per hour of audio for batch REST transcription; $0.20 per hour of audio for streaming transcription
View model
→