What is Llama 4 Maverick?
Llama 4 Maverick is an open-weight artificial intelligence model from Meta. It belongs to the Llama 4 family and is designed for applications that need both language understanding and image understanding. In practical terms, it can respond to text prompts, inspect images, answer questions about visual content, write and explain code, summarize information, and handle other general assistant tasks.
Unlike a conventional dense model, Maverick uses a mixture-of-experts architecture. It has 400 billion total parameters, but 17 billion active parameters are used for an individual input according to Meta’s published specifications. The model contains 128 routed experts and a shared expert. This design allows the model to have a large overall capacity without activating every parameter for every token.
Meta released Llama 4 Maverick as an open-weight model on April 5, 2025. “Open-weight” means that the model weights are available for download and deployment under Meta’s Llama 4 Community License. This is different from a conventional hosted-only service: users or organizations can arrange their own infrastructure or use a third-party provider, subject to the license and the technical requirements of running the model.
Where Maverick fits in Meta’s model lineup
Maverick is the 17-billion-active-parameter, 400-billion-total-parameter member of Meta’s Llama 4 family. Its positioning is between a general-purpose language model and a vision-language model: it is intended to work with text and images rather than text alone.
The model is available as both a base model and an instruction-tuned model. The instruction-tuned version is the practical choice for assistant-style interactions because it is prepared to follow user instructions and use the documented Llama 4 prompt and tool-calling formats. The reported 1-million-token context length applies to the instruction-tuned model; the base model is documented with a shorter 256,000-token context length.
The model should not be confused with Meta AI, the consumer assistant available through Meta’s applications and services. Llama 4 Maverick is a downloadable model checkpoint for developers and organizations, while Meta AI is a consumer-facing product that may use different systems and features.
Inputs, outputs, and core capabilities
| Capability | Verified detail |
|---|---|
| Text input | Supported |
| Image input | Supported |
| Audio input | Not documented for this model |
| Video input | Not documented for this model |
| Text and code output | Supported |
| Image, audio, or video output | Not natively supported |
| Context length | Up to 1 million tokens for the instruction-tuned model; 256,000 tokens for the base model |
| Knowledge cutoff | August 2024 |
Maverick’s multimodality is primarily about understanding images rather than generating media. A user can provide a photograph, diagram, screenshot, chart, or other supported image and ask the model to describe it, extract information, answer questions about it, or connect its visual content with written instructions. The model returns text or code, not a newly generated image, video, or audio file.
The model’s general language capabilities cover assistant conversations, creative writing, summarization, multilingual applications, visual question answering, and coding. These are broad intended uses rather than a guarantee that every task will be equally reliable. Results depend on the quality of the prompt, the deployment environment, the input image, and the complexity of the task.
Long context and knowledge limitations
The instruction-tuned model’s reported 1-million-token context window is one of Maverick’s most significant practical distinctions. A context window is the amount of input and conversation material the model can consider at once. A large window can be useful for long documents, collections of files, extensive codebases, transcripts, or image-and-text workflows that would otherwise need to be divided into many smaller requests.
The context limit is not the same as a knowledge cutoff or a guarantee of perfect recall. Meta’s model documentation identifies August 2024 as the pretraining-data cutoff. Information created after that date is not part of the model’s underlying learned knowledge unless it is supplied by the user or connected through external retrieval and tools.
No maximum output-token limit is identified in the supplied official materials. The 1-million-token figure should therefore be treated as the documented context capacity, not as a promise that a single response can contain one million generated tokens. Actual input and output limits may also depend on the runtime, hardware, serving provider, quantization, and deployment configuration.
Reasoning, coding, and tool support
Maverick is intended for visual reasoning and general reasoning tasks such as interpreting an image, comparing evidence in a document, following multi-step instructions, or explaining a technical problem. The supplied evaluation record rates its reasoning capability as 8 out of 10, but that is an editorial assessment rather than a score published by Meta. It should be used as a relative guide, not as a standardized benchmark result.
The model is also suited to coding assistance. It can generate code, explain existing code, suggest changes, and help analyze technical material. The supplied assessment rates coding at 8 out of 10, again as an editorial evaluation rather than a provider-published fact. It should not be interpreted as proof that Maverick will outperform every other coding model or that generated code is safe to run without review.
Official Llama 4 documentation describes zero-shot function and tool-calling formats. This means an application can prompt the model to produce a structured request for an external function, such as retrieving information, querying a database, or performing an application action. The model itself does not automatically provide web search, a database, or an execution environment. Developers must supply the tools, validate the model’s arguments, and decide whether an action is safe to execute.
The supplied research confirms tool use and streaming support, but it does not establish a separate JSON-schema-constrained output mode or a distinct legacy JSON mode. Tool-call formatting should therefore not be presented as equivalent to guaranteed schema-constrained generation.
Architecture, license, and deployment
Because Maverick is distributed as open weights, it can be considered by teams that need more control over deployment than a hosted-only model provides. Potential advantages include choosing infrastructure, integrating the model into an existing application stack, and keeping model-serving decisions closer to the organization. Those advantages come with operational responsibilities: the deployer must provide suitable hardware or select a hosting partner, manage latency and capacity, monitor outputs, and comply with Meta’s Llama 4 Community License.
The license is not described in the supplied sources as a standard permissive open-source license. Organizations should review the official Llama 4 Community License before commercial deployment, redistribution, or use in regulated environments. Open weights also do not remove the need for privacy, security, copyright, and data-governance reviews.
No official per-token hosted price from Meta was identified. Since the primary distribution model is downloadable weights, there is no single Meta-published input or output price to quote. Third-party hosting services may charge for compute, requests, or tokens, and their prices can vary with hardware, quantization, region, concurrency, and service level.
Main strengths and trade-offs
- Native image understanding: Maverick can combine written instructions with image inputs, which is useful for screenshots, diagrams, documents, and visual question answering.
- Large context capacity: The instruction-tuned model is documented with a 1-million-token context window, enabling long-document and large-codebase workflows when the deployment supports it.
- Open-weight availability: Teams can download the model and choose how to deploy it rather than relying exclusively on a provider-hosted endpoint.
- Broad general-purpose coverage: The model targets assistant work, creative writing, multilingual applications, coding, and visual reasoning rather than a single narrow task.
- Tool-calling support: The documented prompt format can support function-oriented application workflows when developers provide and supervise external tools.
The main trade-off is that a large open-weight model can require substantial infrastructure and engineering effort. A hosted model may be simpler to start with, while a smaller model may offer lower latency and lower operating cost for straightforward requests. The supplied editorial assessment gives Maverick a speed score of 7 and a cost score of 8 out of 10, but these are not fixed provider specifications: actual speed and cost depend heavily on deployment choices.
Maverick also has capability boundaries. It does not natively generate images, audio, or video; its documented knowledge cutoff is August 2024; and no official maximum output-token value is identified in the supplied material. It may also produce inaccurate answers, so visual interpretations, code, tool arguments, and factual claims require appropriate validation.
Best use cases for Llama 4 Maverick
Maverick is a strong candidate when an application needs open-weight multimodal reasoning and the team can support model deployment. Suitable examples include:
- Answering questions about images, screenshots, diagrams, charts, or scanned material.
- Building internal assistants that combine long documents with text instructions.
- Analyzing large technical repositories or extensive documentation within a single context, subject to runtime limits.
- Generating, explaining, and reviewing code.
- Creating multilingual assistants or creative-writing applications.
- Connecting model responses to controlled business functions through tool calling.
- Deploying a model in an environment where infrastructure and data-handling choices need to remain under the organization’s control.
When to choose this model
Choose Llama 4 Maverick when image understanding, long context, and open-weight deployment matter more than having the simplest possible hosted API. It is particularly relevant for teams that want to experiment with a large multimodal model, customize the serving stack, or connect model reasoning to their own tools and data.
A smaller or more specialized model may be more appropriate when requests are simple, response speed is the primary concern, or infrastructure costs must be minimized. A hosted service may be preferable when the team does not want to manage hardware, model serving, scaling, or updates. A dedicated media-generation model is the better option when the required output is an image, audio clip, or video rather than text or code.
For current facts beyond August 2024, Maverick needs user-provided information or an external retrieval system. For high-stakes decisions, production code, or actions that change external systems, it should be used with verification and explicit safeguards rather than treated as an autonomous authority.
Bottom line
Llama 4 Maverick combines open-weight distribution, image understanding, a mixture-of-experts architecture, coding support, tool-calling formats, and a very large documented context window. Its strongest practical fit is a custom or self-managed assistant that must work across text and images and process substantial amounts of context. Its limitations are equally important: there is no official Meta token price, no documented maximum output limit in the supplied research, no native image or audio generation, and no built-in guarantee of current information or factual accuracy.

