Grok 4.7

Grok 4.7 Fast

by xAI · Available as a faster-serving variant in Cursor and Grok Build; not available as a separate public xAI API model

Grok 4.7 Fast is a faster-serving variant of xAI's Grok 4.7 for Cursor and Grok Build. It targets interactive coding and agentic workflows, supports text and image input with text output, uses a 500,000-token context window, and costs twice standard rates where Fast pricing applies. It is not a separate public xAI API model.

Text Reasoning Coding
Grok 4.7 Fast is xAI's low-latency serving option for Grok 4.7. It is intended for interactive coding, rapid development iterations, and agentic workflows in products such as Cursor and Grok Build. The important distinction is that Fast describes how the model is served, not a separately documented model checkpoint or standalone public xAI API model. It retains Grok 4.7's underlying capabilities, including text and image input, text generation, reasoning, coding, tool use, and a 500,000-token context window, while trading higher usage cost for faster responses where Fast pricing applies.
Outputs

What Grok 4.7 Fast can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Structured output Prompt caching
Model profile

Performance characteristics

9/10 Reasoning
9/10 Coding
10/10 Speed
5/10 Cost efficiency
Specifications

Technical details

Model family Grok 4.7
Model type General Purpose
Context window 500K tokens
Knowledge cutoff May 2026
Release date 2026-09-21
Status Available as a faster-serving variant in Cursor and Grok Build; not available as a separate public xAI API model
Knowledge cutoff notes

The knowledge cutoff belongs to the underlying Grok 4.7 model. Grok 4.7 Fast is a faster-serving variant rather than a separately trained model with its own documented cutoff.

Model notes

Grok 4.7 Fast is the same underlying model as Grok 4.7 served on faster infrastructure. It is not a distinct public xAI API model ID. The provider documents availability in Cursor and Grok Build, with Fast usage billed at twice standard token rates where applicable. Grok Build's free tier does not include the Fast variant. Existing Grok 4.7 records represent the canonical underlying model and should not be treated as a separate model family. The model supports text and image input and text output through the underlying Grok 4.7 capability set; non-text model output is not documented.

Cost

Model pricing

Input 2x standard Grok 4.7 input-token rate where Fast pricing applies; standard Grok 4.7 API rate is $2 per 1M input tokens
Output 2x standard Grok 4.7 output-token rate where Fast pricing applies; standard Grok 4.7 API rate is $6 per 1M output tokens
Model guide

Grok 4.7 Fast: Low-Latency Coding and Agentic Workflows

Grok 4.7 Fast is a faster-serving variant of xAI's Grok 4.7 for supported products including Cursor and Grok Build. It uses the same underlying model rather than a separate public API model, prioritizing lower latency for interactive coding and agentic tasks. Fast usage costs twice the standard Grok 4.7 token rates where that pricing applies.

What is Grok 4.7 Fast?

Grok 4.7 Fast is a low-latency serving variant of Grok 4.7 from xAI. According to the supplied provider documentation, it uses the same underlying model as standard Grok 4.7 but is served on faster infrastructure. This makes the distinction operational rather than architectural: Fast is not presented as a separately trained model family or an independently documented model checkpoint.

That distinction matters when evaluating the product. Users should expect the capabilities associated with Grok 4.7, but they should not assume that “Grok 4.7 Fast” is a model identifier that can be selected through the public xAI API. The public API exposes the canonical grok-4.7 model, while the Fast variant is documented for supported host products.

The primary reason to choose Fast is shorter waiting time during interactive work. A developer repeatedly asking for code changes, reviewing generated output, or supervising an agent can benefit more from faster turnaround than a batch job that processes many requests without human interaction.

Where Grok 4.7 Fast is available

The documented availability of Grok 4.7 Fast is in Cursor and Grok Build. It is not available as a separate Fast model through the public xAI API. This means access depends on the host product's model routing, subscription or plan rules, and billing behavior rather than on a normal API model selection alone.

xAI's documentation also states that the Fast variant is not included in Grok Build's free tier. Users should therefore check the current rules of the product they are using before assuming that access to standard Grok 4.7 also includes faster serving.

Capabilities and supported modalities

Because Grok 4.7 Fast uses the same underlying model as Grok 4.7, the supplied research associates it with a broad set of model capabilities:

  • Input: text and images.
  • Output: text. No direct image, audio, or video output is documented.
  • Reasoning: configurable reasoning effort is supported through the underlying Grok 4.7 capability set.
  • Coding: software engineering and coding assistance are central intended uses.
  • Tools: function calling and tool use are supported through relevant interfaces.
  • Structured responses: structured outputs are supported by the underlying model.
  • Context: a 500,000-token context window is listed for Grok 4.7.

These are capabilities of the underlying Grok 4.7 model as documented by the provider. They should not be read as evidence that Grok 4.7 Fast has separate API parameters, limits, or a distinct model card. In particular, the research does not provide a separate maximum output-token limit for the Fast serving variant.

The model's image-input support can be useful when a coding or agentic task depends on screenshots, diagrams, or other visual references. Its output remains text, so it is better understood as a text-generating model with multimodal input rather than a system for creating images, audio, or video.

Coding and agentic work

Grok 4.7 Fast is most clearly differentiated by the workflow it targets. In Cursor, faster serving can reduce the time between a developer's instruction and a proposed code change, explanation, or debugging step. This is especially relevant when the user is actively reviewing each response and immediately submitting a follow-up request.

In Grok Build, the intended use includes agentic tasks. An agentic workflow is one in which the model performs multiple steps, uses tools, or continues working toward a result rather than answering with a single short response. Faster responses can make these workflows feel more responsive, although the supplied research does not provide independent benchmark results showing how much faster Fast is in a particular workload.

Relevant examples include:

  • Generating and revising code during an interactive development session.
  • Investigating errors while a developer iterates on a solution.
  • Running tool-using development workflows where each step depends on the previous response.
  • Building or refining projects in Grok Build.
  • Handling long-running tasks where repeated response delays materially affect productivity.

Function calling and structured outputs can help a supported host application interpret the model's responses and connect them to tools or application logic. However, the exact tool availability and behavior depend on the interface in which Grok 4.7 Fast is used.

Speed versus cost

Fast serving is a performance-versus-cost choice. Where separate Fast pricing applies, usage is billed at twice the standard Grok 4.7 token rates. Standard Grok 4.7 API pricing is listed as $2 per million input tokens and $6 per million output tokens. Applying the documented two-times relationship gives an implied equivalent of $4 per million input tokens and $12 per million output tokens for Fast-priced usage, but these figures should not be treated as a separate public API price for Grok 4.7 Fast.

Serving optionInput pricing referenceOutput pricing referenceBest fit
Standard Grok 4.7 API$2 per 1 million tokens$6 per 1 million tokensDirect API access and cost-sensitive workloads
Grok 4.7 FastTwice the standard rate where Fast pricing appliesTwice the standard rate where Fast pricing appliesInteractive coding and agentic workflows where latency matters

The practical value of Fast depends on whether quicker responses save enough developer or operator time to justify the added expense. For a person waiting on each response, lower latency may improve productivity. For a batch process that can run unattended, paying a premium for faster serving may provide little benefit.

Strengths and limitations

Strengths

  • Lower-latency positioning: Fast serving is designed for workflows where response time affects the user's ability to continue.
  • Underlying Grok 4.7 capability set: The variant retains the documented model family capabilities rather than being described as a reduced feature version.
  • Strong fit for coding: Software engineering, interactive development, and tool-using workflows are its clearest use cases.
  • Large context window: The underlying model is documented with a 500,000-token context window, which can support large working contexts when the host product makes that capacity available.
  • Image input: Visual references can be supplied alongside text in supported interfaces.

Limitations

  • No separate public API model: Developers looking for a standalone grok-4.7-fast-style API identifier should not assume one exists. The public API model identity is grok-4.7.
  • Product-dependent access: Availability is documented in Cursor and Grok Build, and access may depend on each product's plan and routing rules.
  • Higher cost: Fast usage costs twice standard rates where the separate Fast pricing rule applies.
  • Free-tier exclusion: Grok Build's free tier does not include the Fast variant according to the supplied research.
  • No documented non-text output: The model is documented for text output, not image, audio, or video generation.
  • Unspecified separate limits: The research does not state a Fast-specific maximum output size, streaming behavior, fine-tuning availability, or independent API quota.

When to choose Grok 4.7 Fast

Choose Grok 4.7 Fast when the work is interactive and the cost of waiting is meaningful. It is a reasonable fit for a developer working in Cursor, a user building in Grok Build, or an operator supervising a multi-step agent that benefits from quick turnarounds. The same underlying model capabilities make it suitable for coding assistance, difficult knowledge-work tasks, document creation, presentations, and tool-supported development when those features are exposed by the host interface.

Choose standard Grok 4.7 instead when direct public API access is required, when cost is more important than latency, or when a workload can run in batches without a person waiting for each response. Standard serving is also the clearer option for applications that need the canonical public model identifier rather than product-specific Fast access.

Fast should not be selected solely because it sounds like a more capable model. The supplied research characterizes it as the same underlying Grok 4.7 model on faster infrastructure. Its principal benefit is responsiveness, not a separately documented improvement in reasoning quality, coding accuracy, context capacity, or output modality.

Bottom line

Grok 4.7 Fast is best understood as a premium, lower-latency access mode for Grok 4.7. It is aimed at Cursor and Grok Build users who value rapid iterations in coding and agentic workflows. Its 500,000-token context, text-and-image input, text output, reasoning, coding, function-calling, and structured-output capabilities come from the underlying Grok 4.7 model. The trade-off is limited availability and twice the standard token cost where Fast pricing applies.


Answers to Frequently Asked Questions

When should you choose Grok 4.7 Fast instead of standard Grok 4.7?
Choose Grok 4.7 Fast for interactive coding, Cursor development, Grok Build projects, and multi-step agentic workflows where shorter response times improve productivity. Choose standard Grok 4.7 when direct public API access, lower cost, or unattended batch processing is more important than latency.
How much does Grok 4.7 Fast cost compared with standard Grok 4.7?
Where separate Fast pricing applies, usage costs twice the standard Grok 4.7 token rates. Based on standard rates of $2 per million input tokens and $6 per million output tokens, the implied Fast-equivalent rates are $4 per million input tokens and $12 per million output tokens. These are not separate public API prices for Grok 4.7 Fast.
What can Grok 4.7 Fast do?
Grok 4.7 Fast supports the documented capabilities of the underlying Grok 4.7 model, including text and image input, text output, configurable reasoning, coding assistance, function calling, tool use, structured outputs, and a listed 500,000-token context window.
What is Grok 4.7 Fast?
Grok 4.7 Fast is a lower-latency serving variant of Grok 4.7 from xAI. It uses the same underlying model on faster infrastructure, so its main distinction is responsiveness rather than a separate model architecture or independently trained checkpoint.
Where is Grok 4.7 Fast available?
Grok 4.7 Fast is documented for use in Cursor and Grok Build. It is not available as a separate Fast model identifier through the public xAI API, which exposes the canonical grok-4.7 model.


Sources 4
Provider

About xAI