What is Grok 4.20-0309-non-reasoning?
Grok 4.20-0309-non-reasoning is xAI’s non-reasoning variant in the Grok 4.20 model family. Its design priority is fast response generation for applications that need capable language processing without the extra test-time deliberation associated with a dedicated reasoning model.
The canonical API model identifier is grok-4.20-0309-non-reasoning. xAI also documents aliases including grok-4.20-non-reasoning and grok-4.20-non-reasoning-latest. The supplied model documentation identifies the model as a current API offering, while the separate Grok consumer product and other Grok models serve different access and workload requirements.
In practical terms, this model is intended for applications such as chat, text transformation, coding assistance, visual question answering, document analysis, structured extraction and tool-connected agents. It is not an image, audio or video generation model.
Inputs, outputs and modalities
The model accepts both text and images. Image input allows an application to provide a photograph, diagram, screenshot or other supported visual material alongside instructions. This makes the model useful for image-aware question answering, visual inspection and document workflows where the source includes pages or figures.
Its native output is text only. Although xAI offers separate media capabilities elsewhere in the broader Grok ecosystem, Grok 4.20-0309-non-reasoning does not directly generate images, audio or video. Applications that require those output types should use a dedicated media model or another appropriate service.
| Capability | Support |
|---|---|
| Text input | Supported |
| Image input | Supported |
| Audio or video input | Not documented for this model |
| Text output | Supported |
| Image, audio or video output | Not supported |
| Function calling | Supported |
| Structured outputs | Supported |
Structured outputs are useful when an application needs responses that follow a specified machine-readable structure. Function calling serves a different purpose: it lets the model request actions or information from external tools, such as business systems, search services or application functions. The supplied research verifies structured outputs and function calling, but it does not establish that every structured-output configuration is a separate provider-defined JSON mode.
Context window and output limits
Grok 4.20-0309-non-reasoning has a documented context window of 1,000,000 tokens. A context window is the amount of input and conversation material the model can consider in one request, subject to the API’s request rules. This unusually large limit is useful for long documents, large codebases, extended conversation histories and knowledge workflows that would otherwise require aggressive splitting or summarization.
xAI’s pricing documentation treats requests at or above 200,000 prompt tokens as long-context requests. That threshold is a pricing boundary rather than the model’s maximum context size: the documented maximum remains 1,000,000 tokens.
A maximum output-token limit for this exact model was not directly verified in the supplied documentation. Applications should therefore obtain the current limit from the model’s API documentation or enforce their own output budget rather than assuming that the full context window is available for generated output.
Pricing and cost trade-offs
Standard pricing is charged per million tokens:
| Usage type | Standard rate |
|---|---|
| Input tokens | $1.25 per million |
| Cached input tokens | $0.20 per million |
| Output tokens | $2.50 per million |
| Long-context input at or above 200,000 prompt tokens | $2.50 per million |
| Long-context cached input at or above 200,000 prompt tokens | $0.40 per million |
| Long-context output at or above 200,000 prompt tokens | $5.00 per million |
These prices make ordinary requests substantially cheaper than long-context requests, so sending a very large prompt has a direct cost consequence even when it remains below the 1-million-token capacity. Cached-input pricing can reduce the cost of repeated prompt material when the API recognizes it as cacheable. Requests sent through the US regional endpoint incur a 10% token-pricing premium.
xAI also documents Batch API support and a 20% batch discount for this model. Batch processing is suited to asynchronous workloads such as bulk classification, document extraction or offline content transformation where an immediate response is not required. It is less suitable for interactive chat or user-facing requests that need a prompt response.
Speed, reasoning and coding
The non-reasoning designation is the model’s most important positioning detail. It is intended to answer directly without a dedicated reasoning mode that spends additional inference effort on difficult multi-step problems. That generally makes it a better fit when latency, throughput and predictable per-request cost matter more than maximum deliberation.
This does not mean the model cannot perform multi-step tasks or write code. It can generate explanations, transform code, help debug, produce scripts and participate in tool-connected workflows. However, the supplied research does not provide a benchmark proving a particular coding rank or reasoning advantage. The editorial assessment rates its coding suitability as strong and its speed as very strong; those are comparative evaluations, not xAI-published benchmark results.
For difficult mathematical proofs, deeply nested planning, or problems where careful extended deliberation is more valuable than response speed, a dedicated reasoning model may be more appropriate. The non-reasoning model is better understood as a fast general-purpose worker than as a specialist for maximum-depth inference.
API features and availability
The model is documented for the xAI API in the us-east-1 and us-west-2 regions. The supplied rate-limit research lists a baseline limit of 37 requests per second and 10 million tokens per minute, with higher limits available at higher account tiers.
Streaming is supported through xAI’s text-generation API surfaces, allowing an application to display generated text progressively instead of waiting for the complete response. Prompt caching is reflected in the separate cached-input price. Batch processing is available for asynchronous jobs, and function calling allows the model to participate in application workflows beyond simple text generation.
Exact request syntax can vary depending on the xAI API surface and the combination of streaming, tools, structured outputs and long-context input. Implementers should use the current xAI documentation for the endpoint and request format rather than copying assumptions from an unrelated SDK generation.
Best use cases
- Fast general-purpose generation: Use it for chat responses, rewriting, summarization, classification and other text tasks where low latency is important.
- Image-aware analysis: Provide screenshots, diagrams or document pages with text instructions for visual question answering and inspection.
- Large-document workflows: Use the 1-million-token context window for long reports, extensive source material or large knowledge inputs, while accounting for the higher long-context rates.
- Coding assistance: Use it for code generation, explanation, transformation and debugging when rapid iteration is more important than maximum reasoning depth.
- Tool-connected agents: Combine function calling with external services, application functions or data systems.
- Structured extraction: Request structured outputs for downstream processing of documents, records and other semi-structured content.
- High-volume offline processing: Use the Batch API when jobs can run asynchronously and the 20% batch discount is valuable.
Limitations and when to choose another option
The model’s text-only output rules out native image, audio and video generation. A media-generation model is a better choice when the application must create those formats rather than analyze them. Audio and video input are also not documented for this exact model.
Its non-reasoning configuration is another important limitation. For complex proofs, demanding multi-stage planning or tasks that benefit from extensive internal deliberation, choose a reasoning-focused option instead. Conversely, selecting a reasoning model for every request may increase latency and cost without improving routine text work.
Long context is useful but not automatically economical. Once a prompt reaches the 200,000-token threshold, both input and output rates increase. Large context also does not guarantee that every detail will receive equal attention, so applications should still organize source material clearly and test retrieval or extraction quality on their own data.
The model’s exact knowledge cutoff was not published in the authoritative documentation reviewed for this record. xAI search tools can provide current external information when explicitly enabled, but tool access should not be treated as proof that the underlying model has a permanently current training dataset.
Bottom line
Grok 4.20-0309-non-reasoning is a strong fit for developers who need fast text generation with image understanding, tool calling, structured responses and unusually long context. Its principal trade-off is deliberate: it prioritizes speed and general-purpose throughput over the deeper deliberation of reasoning-oriented models. It is most compelling for interactive applications, coding support, document processing and agent workflows that can use text output and can manage the higher prices associated with very large prompts.

