Grok 4.20

Grok 4.20-0309-reasoning

by xAI · current

xAI’s Grok 4.20-0309-reasoning is a reasoning-focused model for complex analysis, coding, research and agentic workflows. It accepts text and images, produces text, supports structured outputs and function calling, and provides a 1-million-token context window. Standard pricing is $1.25 per million input tokens and $2.50 per million output tokens, while prompts of at least 200,000 tokens use higher long-context rates.

Text Reasoning Coding
Grok 4.20-0309-reasoning is the reasoning variant in xAI’s Grok 4.20 API family. Its dated model identifier is grok-4.20-0309-reasoning. The model accepts text and image inputs, returns text, supports tool calls and structured responses, and can process up to 1,000,000 tokens of context. It is designed for tasks where careful multi-step analysis matters more than the lowest possible latency, including technical research, coding, large-document analysis and agentic workflows.
Outputs

What Grok 4.20-0309-reasoning can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching Batch API
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Grok 4.20
Model type Reasoning
Context window 1M tokens
Release date 2026-04-07
Status current
Knowledge cutoff notes

xAI’s available official model documentation and Grok 4.20 model card do not provide a specific knowledge-cutoff date for the exact grok-4.20-0309-reasoning model. Server-side web or X search can provide newer external information during a request but does not change the model’s underlying knowledge cutoff.

Model notes

The canonical API identifier is grok-4.20-0309-reasoning. xAI lists aliases including grok-4.20, grok-4.20-reasoning, and grok-4.20-reasoning-latest, plus beta and experimental aliases. Aliases are not separate model entities. xAI documents text and image input with text output, function calling, structured outputs, a 1-million-token context window, caching, streaming workflows, and Batch API support. The model has a 20% batch discount. Long-context pricing applies once a prompt reaches 200K tokens. The release date is recorded as April 7, 2026 based on the dated official Grok 4.20 model card; the exact first API availability date was not separately verified. Reasoning, coding, speed, and cost scores are editorial comparative estimates, not provider-published ratings.

Cost

Model pricing

Input $1.25 per 1M tokens for prompts under 200K tokens; $2.50 per 1M tokens for prompts at or above 200K tokens; cached input is $0.20 or $0.40 per 1M tokens respectively
Output $2.50 per 1M tokens for prompts under 200K tokens; $5.00 per 1M tokens for prompts at or above 200K tokens
Model guide

Grok 4.20-0309-reasoning: xAI’s 1M-Token Model for Complex Analysis

Grok 4.20-0309-reasoning is xAI’s reasoning-focused API model for complex analysis, coding, research, long-context document work, image understanding and tool-enabled workflows. It accepts text and images, produces text, supports function calling and structured outputs, and provides a 1-million-token context window. Standard pricing is $1.25 per million input tokens and $2.50 per million output tokens for prompts below 200,000 tokens, with higher long-context rates and discounted cached-input and batch-processing options.

What is Grok 4.20-0309-reasoning?

Grok 4.20-0309-reasoning is a reasoning-focused language model provided by xAI. The model’s canonical API identifier is grok-4.20-0309-reasoning, and it is listed in xAI’s Grok 4.20 model catalog alongside separately listed non-reasoning and multi-agent variants.

In practical terms, it is intended for problems that benefit from several stages of analysis before an answer is produced. Examples include comparing technical documents, tracing a difficult software bug, extracting conclusions from a large collection of files, planning a tool-based workflow, or interpreting an image together with detailed written instructions.

xAI describes Grok 4.20 as a high-performance model focused on speed, prompt adherence, low hallucination rates, reasoning and agentic tool calling. Those are provider claims rather than independent benchmark results. The supplied documentation does not provide a benchmark table for this exact dated model, so its comparative scores should be treated cautiously.

Where it fits in xAI’s lineup

This model is the reasoning member of the Grok 4.20 API family. The dated identifier is useful when reproducibility matters because xAI also lists rolling aliases such as grok-4.20, grok-4.20-reasoning and grok-4.20-reasoning-latest. Aliases can point to changing versions, while the dated identifier is intended to identify a specific release.

Grok 4.20-0309-reasoning should not be confused with the Grok assistant available through consumer applications. Consumer Grok may expose web search, X search, voice, image generation, video generation or other application-level features, but those features are not native output modalities of this API model. The model itself produces text and accepts text or images according to the supplied model documentation.

Inputs, outputs and modalities

The model supports text and image input. Text input can include ordinary prompts, technical material, instructions and content supplied as part of a larger workflow. Image input allows the model to analyze visual information alongside a written request.

Its native output is text. It does not natively generate images, audio or video, so it is not the right choice when the central requirement is media creation. A surrounding application can use its text output to control another service or tool, but that does not make image, audio or video generation an inherent capability of Grok 4.20-0309-reasoning.

  • Text input: Supported.
  • Image input: Supported.
  • Audio input: Not listed as supported.
  • Video input: Not listed as supported.
  • Text output: Supported.
  • Image, audio and video output: Not supported as native model outputs.

This distinction matters when evaluating multimodal systems. Grok 4.20-0309-reasoning is multimodal because it can understand images as well as text, not because it directly creates non-text media.

Context window and output limit

The model has a context window of 1,000,000 tokens. A context window is the amount of input and generated conversation material the model can consider within a request. A million-token limit can accommodate very large documents, extensive codebases, long research collections or multi-step agent workflows, subject to the application’s own request construction and file-handling implementation.

The large context window does not guarantee that every very long prompt will receive equally detailed attention. For practical use, it remains sensible to organize documents, identify the relevant task clearly and ask for focused outputs rather than sending large amounts of unrelated material.

xAI’s available documentation does not publish a fixed model-specific maximum output-token value for this model. That value should therefore be treated as unknown rather than assumed to be equal to the context window. A request or API response may report no explicit maximum when one has not been supplied.

Reasoning, coding and tool support

The model’s main distinction is its reasoning orientation. It is designed for tasks requiring decomposition, comparison, inference and multi-step problem solving. Reasoning can be useful when the correct answer depends on connecting evidence across a long prompt or when an application needs the model to decide which action to take next.

It is also suitable for coding assistance. Supported uses include explaining unfamiliar code, proposing implementations, reviewing changes, diagnosing errors and helping coordinate code-related tools. The available research assigns editorial scores of 8 out of 10 for reasoning and coding, but these are comparative estimates created for the catalog and are not ratings published by xAI.

Function calling lets the model request an operation from an external application using a defined schema. For example, an application could expose a database lookup, document retrieval function or business workflow action. The application remains responsible for executing the function, validating arguments and applying permissions. Structured outputs can help the model return data in a predictable machine-readable format instead of unrestricted prose.

The model supports streaming, caching and batch processing. Streaming can let an application display generated text as it arrives. Caching can reduce the cost of repeated input material where the provider’s caching rules apply. Batch processing is intended for workloads that do not need immediate responses and receives a documented discount for the Grok 4.20 reasoning, non-reasoning and multi-agent models.

Pricing for Grok 4.20-0309-reasoning

xAI lists separate input, cached-input and output rates. For prompts below 200,000 tokens, the standard rates are:

Usage typePrice per 1 million tokens
Input$1.25
Cached input$0.20
Output$2.50

When a prompt reaches at least 200,000 tokens, the long-context rates apply:

Usage typePrice per 1 million tokens
Input$2.50
Cached input$0.40
Output$5.00

The threshold applies to the prompt length, not simply to the model’s maximum capacity. A short request and response therefore use the standard rates even though the model can accept much larger contexts. xAI also documents a 20% batch discount for the relevant Grok 4.20 models. Requests sent through xAI’s US regional endpoint carry a stated 10% regional pricing premium.

These prices are usage rates rather than a recurring consumer subscription price. Actual costs depend on input volume, output volume, cached material, long-context requests, batch use and regional routing.

Main strengths and limitations

The strongest reason to choose this model is the combination of reasoning capability, long context, image understanding and tool support. It can bring together written instructions, visual evidence, large files and external functions in one workflow. That makes it a plausible fit for research assistants, coding systems, document-analysis products and agentic applications.

  • Long-context analysis: The 1-million-token window is useful for large document sets, lengthy code and extended workflows.
  • Multimodal understanding: Text and image input support allows visual evidence to be considered with written context.
  • Structured application integration: Function calling and structured outputs are useful when model responses must connect reliably to software.
  • Operational flexibility: Streaming, caching and batch processing support different latency and cost requirements.
  • Reasoning-oriented behavior: The model is positioned for complex analysis rather than only short conversational responses.

There are also important limitations. Reasoning-oriented models may be less appropriate than a faster non-reasoning option when an application needs very low latency for simple requests. The higher long-context rates can make large prompts expensive, and generated output is billed separately. The model cannot directly create images, audio or video, and the maximum output-token limit is not specified in the available documentation.

The model’s underlying knowledge-cutoff date is not disclosed in the supplied official documentation. Web or X search tools, where available through an application, can provide newer external information during a request, but they do not change the model’s underlying training cutoff. As with other generative models, important outputs should be checked rather than treated as authoritative without verification.

When to choose Grok 4.20-0309-reasoning

Choose this model when the task benefits from careful multi-step reasoning and the input may be unusually large. Good examples include:

  • Analyzing a long technical specification or collection of reports.
  • Reviewing code across multiple files and explaining interactions between components.
  • Extracting structured records from documents or images.
  • Building an agent that chooses among external tools or business actions.
  • Combining image interpretation with detailed written analysis.
  • Running offline or non-urgent large-scale processing through the Batch API.

A different option may be better for straightforward classification, short answers or latency-sensitive interactions where extended reasoning is unnecessary. A non-reasoning sibling may offer a better speed or cost profile for those requests, although the supplied research does not provide a direct price or performance comparison. A dedicated image, audio or video model is more appropriate when generating those media types is the primary objective. Similarly, a specialized embedding or speech model should be preferred for workloads designed specifically around semantic indexing or voice processing.

Availability and identifiers

The canonical identifier is grok-4.20-0309-reasoning. xAI lists the model as available in the US East and US West regions. The dated model was recorded with a release date of April 7, 2026 based on xAI’s dated Grok 4.20 model card. The exact first API availability date was not separately verified.

No provider-published deprecation or shutdown date is verified in the supplied research. For applications that need stable behavior, the dated identifier is preferable to a rolling alias. For applications that prioritize automatically receiving the newest compatible version, an alias may be more convenient but can reduce reproducibility.

Bottom line

Grok 4.20-0309-reasoning is best understood as a long-context, reasoning-oriented text model with image understanding and software-integration features. Its 1-million-token context, structured outputs and function calling make it suited to complex research, coding and agentic workflows. Its main trade-offs are higher costs for very long prompts, an unspecified maximum output limit, the lack of native media generation and the additional latency or expense that may accompany reasoning-heavy work.


Answers to Frequently Asked Questions

When should I choose Grok 4.20-0309-reasoning?
Choose it for complex reasoning, long-context document analysis, multi-file code review, image-and-text interpretation, structured extraction and agentic workflows using external tools. A faster non-reasoning model may be more suitable for simple or latency-sensitive requests, while dedicated media models are better for generating images, audio or video.
How much does Grok 4.20-0309-reasoning cost?
For prompts under 200,000 tokens, pricing is $1.25 per million input tokens, $0.20 per million cached-input tokens and $2.50 per million output tokens. For prompts of at least 200,000 tokens, the rates increase to $2.50 for input, $0.40 for cached input and $5.00 for output per million tokens. Batch requests receive a documented 20% discount.
What input and output modalities does Grok 4.20-0309-reasoning support?
Grok 4.20-0309-reasoning accepts text and images and generates text. It does not natively generate or accept audio and video, nor does it natively generate images. External applications can connect its text output to other tools or media services.
What is Grok 4.20-0309-reasoning?
Grok 4.20-0309-reasoning is xAI’s reasoning-focused language model for complex, multi-step analysis. It supports text and image input, produces text, and is designed for tasks such as technical document comparison, code debugging, large-scale research and tool-based workflows.
How large is Grok 4.20-0309-reasoning’s context window?
The model has a context window of 1,000,000 tokens, which can accommodate very large document collections, extensive codebases and long agent workflows. However, xAI does not specify a fixed maximum output-token limit for this model.


Sources 6
Provider

About xAI