Grok 4.20

Grok 4.20-0309-non-reasoning

by xAI · current

A practical guide to xAI’s Grok 4.20-0309-non-reasoning API model, covering its text and image inputs, text-only output, 1-million-token context, pricing, long-context rates, tool support, Batch API, speed-oriented design and limitations compared with reasoning-focused options.

Text Reasoning Coding
Grok 4.20-0309-non-reasoning is an xAI API model built for fast general-purpose generation rather than extended deliberation. It accepts text and images, produces text, supports function calling and structured outputs, and can process up to 1,000,000 tokens of context. Standard pricing is $1.25 per million input tokens and $2.50 per million output tokens, with separate cached-input and long-context rates.
Outputs

What Grok 4.20-0309-non-reasoning can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching Batch API
Model profile

Performance characteristics

6/10 Reasoning
8/10 Coding
9/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Grok 4.20
Model type General Purpose
Context window 1M tokens
Release date 2026-04-07
Status current
Knowledge cutoff notes

No direct authoritative knowledge-cutoff date was published for the exact Grok 4.20-0309-non-reasoning model in the reviewed xAI documentation. Server-side web and X search tools can provide external current information but do not change the underlying model cutoff.

Model notes

The exact canonical model ID is grok-4.20-0309-non-reasoning. Official aliases include grok-4.20-non-reasoning, grok-4.20-non-reasoning-latest, grok-4.20-beta-non-reasoning, grok-4.20-beta-latest-non-reasoning, grok-4.20-beta-0309-non-reasoning and grok-4.20-non-reasoning-gv2. The model accepts text and image inputs and returns text. xAI documents function calling and structured outputs. Batch API is supported and the current pricing documentation lists a 20% batch discount. Requests sent to the US regional endpoint receive a 10% token-pricing premium. The April 7, 2026 date is the publication date of the Grok 4.20 system card and is used as the model family's documented release reference; a separate exact API launch date for this non-reasoning snapshot was not found. Editorial scores are comparative estimates, not provider-published benchmarks. The knowledge cutoff and maximum output token limit were not directly verified for this exact model.

Cost

Model pricing

Input $1.25 per 1M tokens; $0.20 per 1M cached tokens. Long-context requests at or above 200K prompt tokens: $2.50 per 1M input tokens; $0.40 per 1M cached tokens.
Output $2.50 per 1M tokens; $5.00 per 1M tokens for long-context requests at or above 200K prompt tokens.
Model guide

Grok 4.20-0309-non-reasoning: Fast Long-Context API Inference

Grok 4.20-0309-non-reasoning is xAI’s fast, text-generating API model for general language, coding, image-aware analysis, tool calling and structured workflows. It accepts text and images, returns text, supports a 1-million-token context window and Batch API processing, and omits the dedicated test-time reasoning mode used by reasoning-focused models.

What is Grok 4.20-0309-non-reasoning?

Grok 4.20-0309-non-reasoning is xAI’s non-reasoning variant in the Grok 4.20 model family. Its design priority is fast response generation for applications that need capable language processing without the extra test-time deliberation associated with a dedicated reasoning model.

The canonical API model identifier is grok-4.20-0309-non-reasoning. xAI also documents aliases including grok-4.20-non-reasoning and grok-4.20-non-reasoning-latest. The supplied model documentation identifies the model as a current API offering, while the separate Grok consumer product and other Grok models serve different access and workload requirements.

In practical terms, this model is intended for applications such as chat, text transformation, coding assistance, visual question answering, document analysis, structured extraction and tool-connected agents. It is not an image, audio or video generation model.

Inputs, outputs and modalities

The model accepts both text and images. Image input allows an application to provide a photograph, diagram, screenshot or other supported visual material alongside instructions. This makes the model useful for image-aware question answering, visual inspection and document workflows where the source includes pages or figures.

Its native output is text only. Although xAI offers separate media capabilities elsewhere in the broader Grok ecosystem, Grok 4.20-0309-non-reasoning does not directly generate images, audio or video. Applications that require those output types should use a dedicated media model or another appropriate service.

CapabilitySupport
Text inputSupported
Image inputSupported
Audio or video inputNot documented for this model
Text outputSupported
Image, audio or video outputNot supported
Function callingSupported
Structured outputsSupported

Structured outputs are useful when an application needs responses that follow a specified machine-readable structure. Function calling serves a different purpose: it lets the model request actions or information from external tools, such as business systems, search services or application functions. The supplied research verifies structured outputs and function calling, but it does not establish that every structured-output configuration is a separate provider-defined JSON mode.

Context window and output limits

Grok 4.20-0309-non-reasoning has a documented context window of 1,000,000 tokens. A context window is the amount of input and conversation material the model can consider in one request, subject to the API’s request rules. This unusually large limit is useful for long documents, large codebases, extended conversation histories and knowledge workflows that would otherwise require aggressive splitting or summarization.

xAI’s pricing documentation treats requests at or above 200,000 prompt tokens as long-context requests. That threshold is a pricing boundary rather than the model’s maximum context size: the documented maximum remains 1,000,000 tokens.

A maximum output-token limit for this exact model was not directly verified in the supplied documentation. Applications should therefore obtain the current limit from the model’s API documentation or enforce their own output budget rather than assuming that the full context window is available for generated output.

Pricing and cost trade-offs

Standard pricing is charged per million tokens:

Usage typeStandard rate
Input tokens$1.25 per million
Cached input tokens$0.20 per million
Output tokens$2.50 per million
Long-context input at or above 200,000 prompt tokens$2.50 per million
Long-context cached input at or above 200,000 prompt tokens$0.40 per million
Long-context output at or above 200,000 prompt tokens$5.00 per million

These prices make ordinary requests substantially cheaper than long-context requests, so sending a very large prompt has a direct cost consequence even when it remains below the 1-million-token capacity. Cached-input pricing can reduce the cost of repeated prompt material when the API recognizes it as cacheable. Requests sent through the US regional endpoint incur a 10% token-pricing premium.

xAI also documents Batch API support and a 20% batch discount for this model. Batch processing is suited to asynchronous workloads such as bulk classification, document extraction or offline content transformation where an immediate response is not required. It is less suitable for interactive chat or user-facing requests that need a prompt response.

Speed, reasoning and coding

The non-reasoning designation is the model’s most important positioning detail. It is intended to answer directly without a dedicated reasoning mode that spends additional inference effort on difficult multi-step problems. That generally makes it a better fit when latency, throughput and predictable per-request cost matter more than maximum deliberation.

This does not mean the model cannot perform multi-step tasks or write code. It can generate explanations, transform code, help debug, produce scripts and participate in tool-connected workflows. However, the supplied research does not provide a benchmark proving a particular coding rank or reasoning advantage. The editorial assessment rates its coding suitability as strong and its speed as very strong; those are comparative evaluations, not xAI-published benchmark results.

For difficult mathematical proofs, deeply nested planning, or problems where careful extended deliberation is more valuable than response speed, a dedicated reasoning model may be more appropriate. The non-reasoning model is better understood as a fast general-purpose worker than as a specialist for maximum-depth inference.

API features and availability

The model is documented for the xAI API in the us-east-1 and us-west-2 regions. The supplied rate-limit research lists a baseline limit of 37 requests per second and 10 million tokens per minute, with higher limits available at higher account tiers.

Streaming is supported through xAI’s text-generation API surfaces, allowing an application to display generated text progressively instead of waiting for the complete response. Prompt caching is reflected in the separate cached-input price. Batch processing is available for asynchronous jobs, and function calling allows the model to participate in application workflows beyond simple text generation.

Exact request syntax can vary depending on the xAI API surface and the combination of streaming, tools, structured outputs and long-context input. Implementers should use the current xAI documentation for the endpoint and request format rather than copying assumptions from an unrelated SDK generation.

Best use cases

  • Fast general-purpose generation: Use it for chat responses, rewriting, summarization, classification and other text tasks where low latency is important.
  • Image-aware analysis: Provide screenshots, diagrams or document pages with text instructions for visual question answering and inspection.
  • Large-document workflows: Use the 1-million-token context window for long reports, extensive source material or large knowledge inputs, while accounting for the higher long-context rates.
  • Coding assistance: Use it for code generation, explanation, transformation and debugging when rapid iteration is more important than maximum reasoning depth.
  • Tool-connected agents: Combine function calling with external services, application functions or data systems.
  • Structured extraction: Request structured outputs for downstream processing of documents, records and other semi-structured content.
  • High-volume offline processing: Use the Batch API when jobs can run asynchronously and the 20% batch discount is valuable.

Limitations and when to choose another option

The model’s text-only output rules out native image, audio and video generation. A media-generation model is a better choice when the application must create those formats rather than analyze them. Audio and video input are also not documented for this exact model.

Its non-reasoning configuration is another important limitation. For complex proofs, demanding multi-stage planning or tasks that benefit from extensive internal deliberation, choose a reasoning-focused option instead. Conversely, selecting a reasoning model for every request may increase latency and cost without improving routine text work.

Long context is useful but not automatically economical. Once a prompt reaches the 200,000-token threshold, both input and output rates increase. Large context also does not guarantee that every detail will receive equal attention, so applications should still organize source material clearly and test retrieval or extraction quality on their own data.

The model’s exact knowledge cutoff was not published in the authoritative documentation reviewed for this record. xAI search tools can provide current external information when explicitly enabled, but tool access should not be treated as proof that the underlying model has a permanently current training dataset.

Bottom line

Grok 4.20-0309-non-reasoning is a strong fit for developers who need fast text generation with image understanding, tool calling, structured responses and unusually long context. Its principal trade-off is deliberate: it prioritizes speed and general-purpose throughput over the deeper deliberation of reasoning-oriented models. It is most compelling for interactive applications, coding support, document processing and agent workflows that can use text output and can manage the higher prices associated with very large prompts.


Answers to Frequently Asked Questions

When should I choose a reasoning model instead of Grok 4.20-0309-non-reasoning?
Choose a reasoning-focused model for complex mathematical proofs, demanding multi-stage planning or tasks that benefit from extended internal deliberation. Grok 4.20-0309-non-reasoning is better suited to fast responses, high throughput and routine text or coding tasks.
How much does Grok 4.20-0309-non-reasoning cost?
Standard pricing is $1.25 per million input tokens, $0.20 per million cached input tokens and $2.50 per million output tokens. Long-context requests of at least 200,000 prompt tokens cost $2.50 per million input tokens, $0.40 per million cached input tokens and $5.00 per million output tokens.
What is the context window of Grok 4.20-0309-non-reasoning?
The model has a documented context window of 1,000,000 tokens. Requests with 200,000 or more prompt tokens are classified as long-context requests for pricing purposes and incur higher rates.
What is Grok 4.20-0309-non-reasoning best used for?
It is designed for fast general-purpose text generation, chat, coding assistance, document analysis, structured extraction, image-aware question answering and tool-connected agent workflows.
Does Grok 4.20-0309-non-reasoning support images and multimodal input?
Yes. The model accepts text and image inputs, including photographs, diagrams, screenshots and document pages. Its output is text only, and it does not generate images, audio or video.


Sources 5
Provider

About xAI