Grok 4.6

Grok 4.6

by xAI · Current and available; superseded as xAI's flagship by Grok 4.7 but still supported on the xAI API.

Grok 4.6 is xAI’s reasoning model for complex coding, research, knowledge work, and agentic applications. It accepts text and images, supports configurable reasoning effort, function calling, structured outputs, streaming, and web or X search tools. Its 500,000-token context window suits large codebases and long documents, while pricing rises for prompts of 200,000 tokens or more. The model returns text only and does not support xAI’s Batch API.

Text Reasoning Coding
Grok 4.6 is an xAI model built for complex coding, research, knowledge work, and tool-driven applications. Its defining practical advantages are a 500,000-token context window, always-on but configurable reasoning, image understanding, and support for functions and structured responses. It produces text rather than images, audio, or video, so it is best suited to software and workflow tasks where analysis and tool use matter more than native media generation.
Outputs

What Grok 4.6 can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching
Model profile

Performance characteristics

9/10 Reasoning
9/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Grok 4.6
Model type Reasoning
Context window 500K tokens
Release date 2026-08-12
Status Current and available; superseded as xAI's flagship by Grok 4.7 but still supported on the xAI API.
Knowledge cutoff notes

The reviewed first-party Grok 4.6 documentation does not state an exact knowledge cutoff. Current information requires enabling xAI server-side web or X search tools or supplying external context.

Model notes

The canonical API identifier is grok-4.6. The model accepts text and image inputs and returns text only. Reasoning is always enabled and supports low, medium, high, and xhigh effort, with high as the default. Function calling and structured outputs are supported. Server-side web and X search can provide current information, but these tools do not change the model's underlying knowledge cutoff. Batch API is explicitly not supported. Pricing is higher for prompts at or above 200,000 tokens. Grok 4.6 remains available even though Grok 4.7 is now xAI's latest flagship model.

Cost

Model pricing

Input $2.00 per 1M tokens below 200k prompt tokens; $4.00 per 1M tokens at 200k tokens or more. Cached input is $0.50 per 1M tokens below 200k and $1.00 per 1M tokens at 200k or more.
Output $6.00 per 1M tokens below 200k prompt tokens; $12.00 per 1M tokens at 200k tokens or more.
Model guide

Grok 4.6: xAI’s Long-Context Reasoning Model for Coding and Agents

Grok 4.6 is xAI’s frontier reasoning model for coding, research, agentic workflows, and long-running interactive tasks. It accepts text and images, supports configurable reasoning effort, function calling, structured outputs, web and X search tools, and a 500,000-token context window. API pricing starts at $2 per million input tokens and $6 per million output tokens, although prompts of 200,000 tokens or more are charged at higher rates.

What is Grok 4.6?

Grok 4.6 is a reasoning-focused model from xAI, released on August 12, 2026, with the API identifier grok-4.6. It is designed for demanding multi-step work rather than only short conversational answers. Typical tasks include software engineering, codebase analysis, research, structured extraction, application development, and assistants that call external tools.

xAI positions Grok 4.6 as a frontier model for coding, agentic workflows, knowledge work, and long-running interactive tasks. In practical terms, that means it can spend more computation working through a problem before returning its answer, while still allowing developers to select how much reasoning effort to use. The model is currently supported on the xAI API and selected xAI and partner products. The supplied documentation identifies Grok 4.7 as xAI’s newer flagship, but Grok 4.6 remains available through the API.

Core capabilities and supported modalities

Grok 4.6 accepts text and images and returns text. Image input allows it to inspect screenshots, diagrams, documents, interfaces, and other visual material alongside written instructions. This makes it suitable for tasks such as explaining a user-interface screenshot, reviewing a visual document, or connecting an image to a coding or research question.

Its output is text-only. Although the wider Grok ecosystem includes image, video, and other media capabilities, those capabilities should not be attributed to Grok 4.6 itself. Applications that need native image, video, audio, speech, or music generation require a different model or API.

  • Text input and text generation
  • Image input for visual analysis and image-grounded reasoning
  • Configurable reasoning effort
  • Function calling and external tool integration
  • Structured machine-readable outputs
  • Streaming responses
  • Server-side web and X search tools when enabled

500,000-token context and configurable reasoning

Grok 4.6 has a documented 500,000-token context window. A context window is the amount of material the model can consider in one request, including instructions, conversation history, uploaded content, tool results, and the requested answer. This large limit is particularly relevant to codebase analysis, long documents, research collections, and multi-step agent sessions.

The supplied research does not specify a maximum output-token limit, so there is no verified maximum output value to report. Developers should distinguish the 500,000-token context capacity from the length of the generated response: the context figure does not automatically mean that every request can produce a 500,000-token answer.

Reasoning is always enabled in Grok 4.6 and cannot be turned off. Developers can select low, medium, high, or xhigh reasoning effort. High is the default. Lower settings are intended for simpler tasks and latency-sensitive tool calls, while high and xhigh provide more internal computation for difficult coding, analysis, and planning problems.

This creates a direct quality, speed, and cost trade-off. A low-effort request may return faster and consume fewer reasoning tokens, while an xhigh request may be more appropriate for a difficult software-engineering problem where additional latency is acceptable. The model’s reasoning setting should therefore be treated as an application-level control rather than a fixed rating of its intelligence.

Coding, tools, and agentic workflows

Coding is one of Grok 4.6’s clearest intended uses. Its combination of long context, reasoning, image input, function calling, and structured outputs can support tasks such as reviewing a large repository, tracing a bug across multiple files, proposing an implementation plan, generating code changes, or returning machine-readable results to an orchestration system.

Function calling allows the model to request actions from application-defined tools. The application, rather than the model, executes those functions and supplies the results back to the conversation. This pattern can connect Grok 4.6 to databases, internal services, code utilities, search systems, or business workflows. Structured outputs are useful when the response must follow a defined schema instead of being interpreted as free-form prose.

xAI also documents server-side web and X search tools for supported requests. These tools can provide more current information than the model’s underlying training knowledge alone. They do not eliminate the need to check sources, and the reviewed Grok 4.6 documentation does not state an exact knowledge cutoff.

Streaming is supported, which allows an application to display a response as it is produced rather than waiting for the entire completion. That can improve perceived responsiveness, although it does not remove the latency associated with higher reasoning effort.

Pricing and API availability

Standard global API pricing for Grok 4.6 is based on token usage. For prompts below 200,000 tokens, the documented rates are:

Usage typePrice
Input tokens$2 per 1 million tokens
Cached input tokens$0.50 per 1 million tokens
Output tokens$6 per 1 million tokens

For prompts of 200,000 tokens or more, the rates increase to $4 per million input tokens, $1 per million cached input tokens, and $12 per million output tokens. The threshold makes long-context planning and codebase analysis more expensive than ordinary requests, even though the model’s 500,000-token context is useful for those workloads.

The referenced API tier has a documented limit of 150 requests per second and 50 million tokens per minute. Grok 4.6 does not support xAI’s Batch API. A US regional endpoint adds a 10 percent token-pricing premium. These limits and prices should be checked against current xAI documentation before deployment because provider pricing, quotas, and model availability can change.

Main strengths and limitations

The strongest verified characteristics of Grok 4.6 are its long context, reasoning controls, coding focus, visual input, and tool integration. A 500,000-token context can reduce the need to divide a large repository or document set into many smaller requests. Function calling and structured outputs make it easier to place the model inside a repeatable software workflow rather than using it only as a chat interface. Web and X search can also help applications retrieve current information when those tools are enabled.

Its main limitations are equally important. Grok 4.6 only produces text, so it is not a direct replacement for xAI models that generate images, video, or audio. Reasoning is always enabled, which can make simple requests slower or more expensive than using a non-reasoning model. Higher effort levels increase the practical cost and latency of difficult tasks. Long prompts also cross into a higher pricing tier at 200,000 tokens.

The lack of Batch API support is a constraint for large asynchronous workloads. Teams processing many independent requests may need to build their own queueing and retry system around standard API calls. The supplied research also does not provide a verified maximum output-token limit or an exact knowledge-cutoff date.

Best use cases for Grok 4.6

  • Agentic coding: Use reasoning, function calls, and structured responses to plan and execute multi-step software tasks.
  • Large codebase analysis: Use the 500,000-token context to provide broad repository or documentation context where supported by the request design.
  • Research assistants: Combine reasoning with web or X search when current information is important.
  • Visual software work: Supply screenshots, diagrams, or interface references for analysis alongside code and instructions.
  • Structured extraction: Return predictable fields from documents or tool results for downstream systems.
  • Interactive application development: Use streaming, tool calls, and long-running context for assistants and prototypes.

For simple classification, short transformations, or extremely latency-sensitive operations, a smaller or non-reasoning model may be more economical. For native media generation, a specialized image, video, or audio model is more appropriate. For bulk asynchronous processing, an option with a supported batch interface may reduce implementation overhead.

When to choose Grok 4.6

Choose Grok 4.6 when the task benefits from deep reasoning, substantial context, coding ability, visual understanding, or interaction with external tools. It is a particularly suitable choice when a single workflow needs to combine repository-scale context, planning, function calls, and structured results.

Choose another option when the priority is the lowest possible latency or cost for straightforward prompts, when reasoning must be disabled, when the application needs native non-text output, or when a batch-processing API is essential. Grok 4.7 may be the more appropriate xAI choice when access to the provider’s latest flagship model is more important than retaining Grok 4.6 specifically. The relevant comparison is not simply which model is newest: it is whether the current task justifies Grok 4.6’s long context and reasoning overhead.

Bottom line

Grok 4.6 is a text-output reasoning model aimed at complex coding, research, and agentic work. Its 500,000-token context, configurable reasoning effort, image input, function calling, structured outputs, and optional web and X search make it more suitable for substantial workflows than for basic chat alone. Its always-on reasoning, long-context price increase, lack of Batch API support, and inability to generate media are the main trade-offs to evaluate before adoption.


Answers to Frequently Asked Questions

What is Grok 4.6 designed for?
Grok 4.6 is a reasoning-focused xAI model designed for complex coding, codebase analysis, research, structured extraction, application development, and agentic workflows that use external tools.
Can Grok 4.6 analyze images and generate media?
Grok 4.6 accepts image inputs and can analyze screenshots, diagrams, documents, and interfaces, but it only produces text. Native image, video, audio, speech, or music generation requires a different model or API.
How much does Grok 4.6 API access cost?
For prompts below 200,000 tokens, Grok 4.6 costs $2 per million input tokens, $0.50 per million cached input tokens, and $6 per million output tokens. For prompts of 200,000 tokens or more, the rates increase to $4 per million input tokens, $1 per million cached input tokens, and $12 per million output tokens. Pricing and quotas should be checked against current xAI documentation before deployment.
What reasoning settings are available in Grok 4.6?
Reasoning is always enabled in Grok 4.6. Developers can choose low, medium, high, or xhigh reasoning effort, with high as the default. Lower settings generally reduce latency and cost, while high and xhigh provide more computation for difficult coding, analysis, and planning tasks.
How large is Grok 4.6’s context window?
Grok 4.6 has a documented 500,000-token context window. This allows it to consider large codebases, documents, conversation histories, tool results, and research collections in a single request, although it does not mean the model can generate a 500,000-token response.


Sources 6
Provider

About xAI