Llama 4

Llama 4 Maverick

by Meta AI · Available; open-weight static checkpoint

Meta’s Llama 4 Maverick is an open-weight multimodal mixture-of-experts model designed for image understanding, visual reasoning, coding, multilingual applications, tool calling, and long-context text processing. It accepts text and images and produces text and code, but does not natively generate images, audio, or video. The instruction-tuned version has a documented 1-million-token context window, while pricing depends on self-hosting or third-party infrastructure rather than a single official Meta token rate.

Text Reasoning Coding
Llama 4 Maverick is Meta’s April 2025 open-weight multimodal model with 17 billion active parameters, 400 billion total parameters, 128 routed experts, and a reported 1 million-token context window for the instruction-tuned version. It accepts text and images and produces text and code, but it does not natively generate images, audio, or video.
Outputs

What Llama 4 Maverick can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Streaming Fine-tuning
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Llama 4
Model type Multimodal
Context window 1M tokens
Maximum output tokens
Knowledge cutoff August 2024
Release date 2025-04-05
Status Available; open-weight static checkpoint
Knowledge cutoff notes

Meta's Llama 4 model card identifies August 2024 as the cutoff for the pretraining data. This is the underlying model knowledge cutoff and is not extended by external retrieval, tools or user-provided context.

Model notes

Llama 4 Maverick is the 17B-active-parameter, 400B-total-parameter member of Meta's Llama 4 family and uses 128 routed experts plus a shared expert. The 1M-token context length applies to the instruction-tuned model; the base model is documented with a shorter 256K context length. Meta reports text and image input with text and code output. The model is distributed under the Llama 4 Community License rather than a standard permissive open-source license. Official documentation describes zero-shot function and tool-calling formats, but does not establish a separate JSON-schema constrained-output or legacy JSON-mode capability. No official Meta per-token hosted price was identified because the primary distribution model is downloadable weights; third-party hosting prices vary. The model's pretraining data cutoff is August 2024.

Model guide

Llama 4 Maverick: Meta’s Open-Weight Multimodal Model for Long Context

Llama 4 Maverick is Meta’s open-weight, natively multimodal mixture-of-experts model for text and image understanding, general assistant tasks, visual reasoning, coding, multilingual applications, creative writing, and long-context processing.

What is Llama 4 Maverick?

Llama 4 Maverick is an open-weight artificial intelligence model from Meta. It belongs to the Llama 4 family and is designed for applications that need both language understanding and image understanding. In practical terms, it can respond to text prompts, inspect images, answer questions about visual content, write and explain code, summarize information, and handle other general assistant tasks.

Unlike a conventional dense model, Maverick uses a mixture-of-experts architecture. It has 400 billion total parameters, but 17 billion active parameters are used for an individual input according to Meta’s published specifications. The model contains 128 routed experts and a shared expert. This design allows the model to have a large overall capacity without activating every parameter for every token.

Meta released Llama 4 Maverick as an open-weight model on April 5, 2025. “Open-weight” means that the model weights are available for download and deployment under Meta’s Llama 4 Community License. This is different from a conventional hosted-only service: users or organizations can arrange their own infrastructure or use a third-party provider, subject to the license and the technical requirements of running the model.

Where Maverick fits in Meta’s model lineup

Maverick is the 17-billion-active-parameter, 400-billion-total-parameter member of Meta’s Llama 4 family. Its positioning is between a general-purpose language model and a vision-language model: it is intended to work with text and images rather than text alone.

The model is available as both a base model and an instruction-tuned model. The instruction-tuned version is the practical choice for assistant-style interactions because it is prepared to follow user instructions and use the documented Llama 4 prompt and tool-calling formats. The reported 1-million-token context length applies to the instruction-tuned model; the base model is documented with a shorter 256,000-token context length.

The model should not be confused with Meta AI, the consumer assistant available through Meta’s applications and services. Llama 4 Maverick is a downloadable model checkpoint for developers and organizations, while Meta AI is a consumer-facing product that may use different systems and features.

Inputs, outputs, and core capabilities

CapabilityVerified detail
Text inputSupported
Image inputSupported
Audio inputNot documented for this model
Video inputNot documented for this model
Text and code outputSupported
Image, audio, or video outputNot natively supported
Context lengthUp to 1 million tokens for the instruction-tuned model; 256,000 tokens for the base model
Knowledge cutoffAugust 2024

Maverick’s multimodality is primarily about understanding images rather than generating media. A user can provide a photograph, diagram, screenshot, chart, or other supported image and ask the model to describe it, extract information, answer questions about it, or connect its visual content with written instructions. The model returns text or code, not a newly generated image, video, or audio file.

The model’s general language capabilities cover assistant conversations, creative writing, summarization, multilingual applications, visual question answering, and coding. These are broad intended uses rather than a guarantee that every task will be equally reliable. Results depend on the quality of the prompt, the deployment environment, the input image, and the complexity of the task.

Long context and knowledge limitations

The instruction-tuned model’s reported 1-million-token context window is one of Maverick’s most significant practical distinctions. A context window is the amount of input and conversation material the model can consider at once. A large window can be useful for long documents, collections of files, extensive codebases, transcripts, or image-and-text workflows that would otherwise need to be divided into many smaller requests.

The context limit is not the same as a knowledge cutoff or a guarantee of perfect recall. Meta’s model documentation identifies August 2024 as the pretraining-data cutoff. Information created after that date is not part of the model’s underlying learned knowledge unless it is supplied by the user or connected through external retrieval and tools.

No maximum output-token limit is identified in the supplied official materials. The 1-million-token figure should therefore be treated as the documented context capacity, not as a promise that a single response can contain one million generated tokens. Actual input and output limits may also depend on the runtime, hardware, serving provider, quantization, and deployment configuration.

Reasoning, coding, and tool support

Maverick is intended for visual reasoning and general reasoning tasks such as interpreting an image, comparing evidence in a document, following multi-step instructions, or explaining a technical problem. The supplied evaluation record rates its reasoning capability as 8 out of 10, but that is an editorial assessment rather than a score published by Meta. It should be used as a relative guide, not as a standardized benchmark result.

The model is also suited to coding assistance. It can generate code, explain existing code, suggest changes, and help analyze technical material. The supplied assessment rates coding at 8 out of 10, again as an editorial evaluation rather than a provider-published fact. It should not be interpreted as proof that Maverick will outperform every other coding model or that generated code is safe to run without review.

Official Llama 4 documentation describes zero-shot function and tool-calling formats. This means an application can prompt the model to produce a structured request for an external function, such as retrieving information, querying a database, or performing an application action. The model itself does not automatically provide web search, a database, or an execution environment. Developers must supply the tools, validate the model’s arguments, and decide whether an action is safe to execute.

The supplied research confirms tool use and streaming support, but it does not establish a separate JSON-schema-constrained output mode or a distinct legacy JSON mode. Tool-call formatting should therefore not be presented as equivalent to guaranteed schema-constrained generation.

Architecture, license, and deployment

Because Maverick is distributed as open weights, it can be considered by teams that need more control over deployment than a hosted-only model provides. Potential advantages include choosing infrastructure, integrating the model into an existing application stack, and keeping model-serving decisions closer to the organization. Those advantages come with operational responsibilities: the deployer must provide suitable hardware or select a hosting partner, manage latency and capacity, monitor outputs, and comply with Meta’s Llama 4 Community License.

The license is not described in the supplied sources as a standard permissive open-source license. Organizations should review the official Llama 4 Community License before commercial deployment, redistribution, or use in regulated environments. Open weights also do not remove the need for privacy, security, copyright, and data-governance reviews.

No official per-token hosted price from Meta was identified. Since the primary distribution model is downloadable weights, there is no single Meta-published input or output price to quote. Third-party hosting services may charge for compute, requests, or tokens, and their prices can vary with hardware, quantization, region, concurrency, and service level.

Main strengths and trade-offs

  • Native image understanding: Maverick can combine written instructions with image inputs, which is useful for screenshots, diagrams, documents, and visual question answering.
  • Large context capacity: The instruction-tuned model is documented with a 1-million-token context window, enabling long-document and large-codebase workflows when the deployment supports it.
  • Open-weight availability: Teams can download the model and choose how to deploy it rather than relying exclusively on a provider-hosted endpoint.
  • Broad general-purpose coverage: The model targets assistant work, creative writing, multilingual applications, coding, and visual reasoning rather than a single narrow task.
  • Tool-calling support: The documented prompt format can support function-oriented application workflows when developers provide and supervise external tools.

The main trade-off is that a large open-weight model can require substantial infrastructure and engineering effort. A hosted model may be simpler to start with, while a smaller model may offer lower latency and lower operating cost for straightforward requests. The supplied editorial assessment gives Maverick a speed score of 7 and a cost score of 8 out of 10, but these are not fixed provider specifications: actual speed and cost depend heavily on deployment choices.

Maverick also has capability boundaries. It does not natively generate images, audio, or video; its documented knowledge cutoff is August 2024; and no official maximum output-token value is identified in the supplied material. It may also produce inaccurate answers, so visual interpretations, code, tool arguments, and factual claims require appropriate validation.

Best use cases for Llama 4 Maverick

Maverick is a strong candidate when an application needs open-weight multimodal reasoning and the team can support model deployment. Suitable examples include:

  • Answering questions about images, screenshots, diagrams, charts, or scanned material.
  • Building internal assistants that combine long documents with text instructions.
  • Analyzing large technical repositories or extensive documentation within a single context, subject to runtime limits.
  • Generating, explaining, and reviewing code.
  • Creating multilingual assistants or creative-writing applications.
  • Connecting model responses to controlled business functions through tool calling.
  • Deploying a model in an environment where infrastructure and data-handling choices need to remain under the organization’s control.

When to choose this model

Choose Llama 4 Maverick when image understanding, long context, and open-weight deployment matter more than having the simplest possible hosted API. It is particularly relevant for teams that want to experiment with a large multimodal model, customize the serving stack, or connect model reasoning to their own tools and data.

A smaller or more specialized model may be more appropriate when requests are simple, response speed is the primary concern, or infrastructure costs must be minimized. A hosted service may be preferable when the team does not want to manage hardware, model serving, scaling, or updates. A dedicated media-generation model is the better option when the required output is an image, audio clip, or video rather than text or code.

For current facts beyond August 2024, Maverick needs user-provided information or an external retrieval system. For high-stakes decisions, production code, or actions that change external systems, it should be used with verification and explicit safeguards rather than treated as an autonomous authority.

Bottom line

Llama 4 Maverick combines open-weight distribution, image understanding, a mixture-of-experts architecture, coding support, tool-calling formats, and a very large documented context window. Its strongest practical fit is a custom or self-managed assistant that must work across text and images and process substantial amounts of context. Its limitations are equally important: there is no official Meta token price, no documented maximum output limit in the supplied research, no native image or audio generation, and no built-in guarantee of current information or factual accuracy.


Answers to Frequently Asked Questions

What is Llama 4 Maverick?
Llama 4 Maverick is Meta’s open-weight multimodal AI model for understanding text and images. It can answer questions about visual content, summarize information, write and explain code, and perform general assistant tasks.
How many parameters and how much context does Llama 4 Maverick support?
Maverick has 400 billion total parameters, with 17 billion active parameters used for an individual input through its mixture-of-experts architecture. The instruction-tuned model supports a documented context window of up to 1 million tokens, while the base model supports 256,000 tokens.
What is Llama 4 Maverick’s knowledge cutoff?
The model’s documented pretraining-data cutoff is August 2024. For newer information, it must receive user-provided content or connect to an external retrieval system or tool.
Is Llama 4 Maverick free to use and available for self-hosting?
Meta released Llama 4 Maverick as an open-weight model that can be downloaded and deployed under the Llama 4 Community License. It does not have a single official Meta per-token hosting price, but self-hosting requires suitable infrastructure and compliance with the license. Third-party providers may charge for compute, requests, or tokens.
Can Llama 4 Maverick process and generate images, audio, or video?
Llama 4 Maverick can process image inputs such as photos, screenshots, diagrams, and charts, but it primarily produces text or code. Audio and video input are not documented for this model, and it does not natively generate images, audio, or video.


Sources 5
Provider

About Meta AI