Grok 4.20

Grok 4.20 Multi-Agent Beta

by xAI · Beta; currently available through the xAI API

Grok 4.20 Multi-Agent Beta is xAI’s specialized model for deep research through coordinated parallel agents. It accepts text and image inputs, returns text, supports a 1-million-token context window, built-in search and code tools, structured outputs, streaming, caching, and batch processing. Its main trade-offs are higher latency and token usage, long-context pricing, tighter rate limits, and no current support for custom tools or max_tokens.

Text Reasoning Coding
Grok 4.20 Multi-Agent Beta is xAI’s specialized model for research tasks that benefit from parallel investigation rather than a single linear response. Using the pinned identifier grok-4.20-multi-agent-0309, it can coordinate multiple agents, use supported server-side tools, and combine their findings into a final answer. It accepts text and image inputs and offers a 1-million-token context window, with standard API pricing of $1.25 per million input tokens and $2.50 per million output tokens below the long-context threshold.
Outputs

What Grok 4.20 Multi-Agent Beta can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching Batch API
Model profile

Performance characteristics

9/10 Reasoning
8/10 Coding
5/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family Grok 4.20
Model type Reasoning
Context window 1M tokens
Status Beta; currently available through the xAI API
Knowledge cutoff notes

xAI's reviewed model documentation does not publish a knowledge cutoff for this exact model. Server-side search tools can provide current information during use without changing the underlying cutoff.

Model notes

The exact pinned model ID is grok-4.20-multi-agent-0309. Official aliases include grok-4.20-multi-agent, grok-4.20-multi-agent-latest, grok-4.20-multi-agent-beta-latest, and grok-4.20-multi-agent-beta-0309. The model coordinates multiple agents and returns the leader agent's result. xAI documents built-in server-side tools such as web search, X search, code execution, and collections search, but client-side and custom tools are not currently supported by the multi-agent variant. The Chat Completions API and max_tokens parameter are unsupported. Long-context pricing applies when a request reaches 200K input tokens. Batch processing is supported with a 20% discount. Release date and knowledge cutoff were not directly published in the reviewed first-party sources.

Cost

Model pricing

Input $1.25 per 1M tokens; $2.50 per 1M tokens for requests at or above 200K input tokens
Output $2.50 per 1M tokens; $5.00 per 1M tokens for requests at or above 200K input tokens
Model guide

Grok 4.20 Multi-Agent Beta: Parallel Deep Research with a 1M-Token Context

Grok 4.20 Multi-Agent Beta is an xAI reasoning model built for deep research workflows. It coordinates multiple agents that investigate different parts of a request and then produces a synthesized answer. The model accepts text and images, returns text, supports built-in research tools and structured outputs, and provides a 1-million-token context window. Its main trade-offs are higher latency, potentially higher token usage, tighter rate limits, and the lack of support for custom client-side tools and max_tokens.

What is Grok 4.20 Multi-Agent Beta?

Grok 4.20 Multi-Agent Beta is an xAI model designed for deep research and complex information-gathering tasks. Rather than treating every request as a single-agent conversation, it can coordinate several agents that explore different aspects of the problem. A leader agent then synthesizes those investigations into the response returned to the caller.

The model’s pinned identifier is grok-4.20-multi-agent-0309. xAI also documents aliases including grok-4.20-multi-agent, grok-4.20-multi-agent-latest, grok-4.20-multi-agent-beta-latest, and grok-4.20-multi-agent-beta-0309. The model is currently described as beta and is available through the xAI API.

In practical terms, this is not primarily a fast chat endpoint. It is intended for requests where the extra work of exploring multiple lines of inquiry can improve the completeness or organization of the result.

How the multi-agent workflow works

For a complicated research question, different agents can investigate separate subproblems in parallel. For example, a request to compare several technologies might require separate research into specifications, pricing, limitations, and market context. The leader agent can then bring those findings together into one response.

The model’s reasoning configuration affects the number of collaborating agents, rather than simply exposing a longer single-agent chain of thought. The internal state of sub-agents is not ordinarily returned to the caller. Instead, the caller receives the leader agent’s tool calls and final response.

This architecture can be useful when a request naturally divides into independent research paths. It can also increase latency and total token consumption, so parallelism should be treated as a capability with a cost rather than as an automatic improvement for every prompt.

Inputs, outputs, and context window

Grok 4.20 Multi-Agent Beta supports text and image inputs and produces text output. It does not provide direct image, audio, or video output. The model has a 1,000,000-token context window, which is substantially more useful for large research prompts, extensive reference material, and long document-based investigations than a short-context chat model.

A context window is the amount of input and generated material the model can process as part of a request. The documented one-million-token capacity does not mean that every request will be inexpensive or fast. Long prompts consume input tokens, and requests that reach the long-context pricing threshold are charged at higher rates.

xAI does not publish a maximum output-token limit for this exact model in the supplied documentation. The max_tokens parameter is documented as unsupported, so applications should not assume they can control output length using that parameter.

Tools and API support

The model supports reasoning, function calling, structured outputs, streaming, cached input, and batch processing. Structured outputs are useful when an application needs a predictable machine-readable response, such as a research report with fixed fields, a comparison matrix, or a list of findings and sources.

xAI documents built-in server-side tools for multi-agent research workflows, including web search, X search, code execution, and collections search. These tools can help the model investigate current information, public posts, calculations, or content stored in supported collections. Search access can improve the freshness of a response during use, but it does not change the model’s underlying knowledge cutoff.

The multi-agent variant has important tool restrictions. Client-side tools and custom tools are not currently supported according to the supplied xAI documentation. The model is accessed through the xAI Responses API or xAI SDK, and the Chat Completions API is not supported for this variant. Applications should therefore verify that their integration uses the supported API generation before implementation.

Pricing and rate limits

For requests below 200,000 input tokens, the documented standard prices are:

Usage typePrice
Input tokens$1.25 per million tokens
Cached input tokens$0.20 per million tokens
Output tokens$2.50 per million tokens

When a request reaches the 200,000-input-token long-context threshold, the rates increase to $2.50 per million input tokens, $0.40 per million cached input tokens, and $5.00 per million output tokens. Batch processing is supported with a documented 20% discount.

The base-tier rate limit is 9 requests per second and 2.5 million tokens per minute. Higher account tiers can provide higher limits. These limits matter for applications that process many research jobs concurrently, especially because multi-agent requests may use more tokens and take longer than simpler single-agent requests.

Strengths and practical capabilities

  • Parallel research: The model is designed to divide complex investigations among collaborating agents and synthesize the results.
  • Long-context analysis: A 1-million-token context window supports large collections of reference material and extended research prompts.
  • Current-information workflows: Built-in web and X search can help retrieve information beyond the model’s static knowledge.
  • Structured research output: Structured outputs and function calling can make results easier to integrate into software.
  • Multimodal research input: Image inputs allow visual material to be included alongside text in supported research tasks.
  • Operational flexibility: Streaming, caching, and batch processing support different application designs and cost controls.

Its coding capability is best understood as part of a research workflow rather than as the model’s sole purpose. Code execution is among the documented built-in tools, so the model can use computation or code-assisted investigation where the supported workflow calls for it. The available research does not provide a separate coding benchmark or a claim that it is the best choice for software development.

Limitations and trade-offs

The main limitation is that multi-agent research adds overhead. Coordinating parallel agents can increase response time and token usage, making the model less suitable for rapid, high-volume conversational applications. Its base rate limits are also tighter than those documented for standard Grok 4.20 variants.

The model is therefore a poor fit when an application needs consistently low latency, a predictable single-agent execution path, or precise control over the number of generated tokens. The unsupported max_tokens parameter may be particularly inconvenient for systems that require strict output budgets.

Custom client-side and other custom tools are not currently supported by the multi-agent variant. Teams that need to connect their own tool implementations should verify whether a different xAI endpoint or model is more appropriate. Similarly, the model’s long-context pricing can make very large requests substantially more expensive than ordinary short-context calls.

The model’s knowledge cutoff is not published in the reviewed first-party documentation. Web and X search can provide newer information during a request, but search results still need evaluation, and they do not guarantee that every answer is complete or correct.

When to choose Grok 4.20 Multi-Agent Beta

Choose Grok 4.20 Multi-Agent Beta when the task benefits from several research paths being explored and reconciled. Suitable examples include preparing a detailed market comparison, examining a large set of documents, investigating a question with multiple factual subtopics, producing a structured research brief, or combining web and X searches with analysis.

The model is especially defensible when completeness and breadth matter more than immediate response speed. Its large context window can also reduce the need to divide a substantial document set into many smaller requests, although the long-context price tier should be included in the cost calculation.

A conventional single-agent model may be more appropriate for short questions, routine chat, simple extraction, or latency-sensitive workloads. A standard Grok endpoint may also be preferable when an application needs higher throughput or a simpler request path. A different option should be considered when custom client-side tools, strict output-token limits, or highly predictable execution are central requirements.

Bottom line

Grok 4.20 Multi-Agent Beta is a specialized research model rather than a general-purpose low-latency chat model. Its defining feature is coordinated investigation: multiple agents can explore different parts of a difficult request before a leader produces the final text response. A 1-million-token context window, built-in search and code tools, structured outputs, streaming, caching, and batch support make it suitable for substantial research workflows.

The trade-off is operational complexity. Users must account for higher possible latency, increased token consumption, long-context pricing, tighter base rate limits, and restrictions on custom tools and max_tokens. For deep, tool-assisted research these compromises may be worthwhile; for quick answers or tightly controlled production interactions, a simpler model can be a better fit.


Answers to Frequently Asked Questions

What is the context window of Grok 4.20 Multi-Agent Beta?
Grok 4.20 Multi-Agent Beta has a 1,000,000-token context window, making it suitable for large research prompts, extensive reference material, and long document-based investigations. Requests using 200,000 or more input tokens are charged at higher long-context rates.
What is Grok 4.20 Multi-Agent Beta designed for?
Grok 4.20 Multi-Agent Beta is designed for complex research and information-gathering tasks. It coordinates multiple agents that investigate different parts of a question, then uses a leader agent to synthesize the findings into a final response.
Which tools and APIs does Grok 4.20 Multi-Agent Beta support?
The model supports reasoning, function calling, structured outputs, streaming, cached input, batch processing, and built-in web search, X search, code execution, and collections search. It is accessed through the xAI Responses API or SDK; Chat Completions, client-side tools, and custom tools are not currently supported.
When should I use Grok 4.20 Multi-Agent Beta instead of a standard chat model?
Use it when a task benefits from parallel research, large document analysis, current-information searches, or structured research reports. A conventional single-agent model is generally better for short questions, routine chat, low-latency workloads, strict output-token limits, or applications requiring custom tools.
How much does Grok 4.20 Multi-Agent Beta cost?
For requests below 200,000 input tokens, pricing is $1.25 per million input tokens, $0.20 per million cached input tokens, and $2.50 per million output tokens. At or above the long-context threshold, the rates increase to $2.50, $0.40, and $5.00 per million tokens respectively. Batch processing has a documented 20% discount.


Sources 5
Provider

About xAI