Llama 4

Llama 4 Scout

by Meta AI · Available; open-weight model

Meta’s Llama 4 Scout is a sparse mixture-of-experts model with 17 billion active and 109 billion total parameters. It accepts text and images, produces text and code, supports tool-call formatting, and offers a 10-million-token context window. The downloadable model is suited to long-document analysis, codebase exploration, visual reasoning, multilingual assistants, and customized self-hosted deployments, but it does not independently execute tools or generate media such as images, audio, or video.

Text Reasoning Coding
Llama 4 Scout is the smaller member of Meta’s Llama 4 model release, but its defining advantage is not simply size: it combines image understanding with an exceptionally large 10-million-token context window. Available as downloadable base and instruction-tuned weights, Scout is intended for developers and organizations that want to analyze very large collections of text, code, documents, or images within their own infrastructure.
Outputs

What Llama 4 Scout can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Streaming Fine-tuning Structured output
Model profile

Performance characteristics

7/10 Reasoning
7/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Llama 4
Model type Multimodal
Context window 10M tokens
Maximum output tokens
Knowledge cutoff August 2024
Release date 2025-04-05
Status Available; open-weight model
Knowledge cutoff notes

Meta's model card identifies August 2024 as the cutoff for the pretraining data. External retrieval or web tools may provide newer information at runtime but do not change the model's underlying cutoff.

Model notes

Llama 4 Scout is a sparse mixture-of-experts model with 17 billion active parameters, 16 experts, and 109 billion total parameters. Meta documents multilingual text and image input and multilingual text and code output. The model has a 10-million-token context length and an August 2024 pretraining-data cutoff. Meta reports that Scout can fit on a single 80 GB H100 with on-the-fly int4 quantization; BF16 deployment requires more memory. The model is available as base and instruction-tuned checkpoints, including the official identifier meta-llama/Llama-4-Scout-17B-16E-Instruct. It is distributed under the Llama 4 Community License Agreement rather than a conventional open-source license. Pricing is not populated because the canonical Meta distribution is downloadable weights rather than a single official per-token hosted price. Editorial scores are comparative estimates, not vendor specifications.

Model guide

Llama 4 Scout: Meta’s Open-Weight Model for Long-Context Vision and Code

Llama 4 Scout is Meta’s open-weight, natively multimodal mixture-of-experts model. It has 17 billion active parameters, 109 billion total parameters, native text-and-image input, text-and-code output, tool-calling support, and a 10-million-token context window aimed at long-document, codebase, and visual-analysis workloads.

What is Llama 4 Scout?

Llama 4 Scout is an open-weight multimodal language model released by Meta on April 5, 2025. It is designed to accept text and images and produce text and code. The instruction-tuned version is intended for assistant-style use, including visual question answering, document analysis, coding assistance, multilingual interaction, and tool-oriented workflows.

Scout uses a mixture-of-experts architecture. Instead of activating every parameter for every token, the model routes each token through a subset of specialized neural networks called experts. Scout has 109 billion total parameters, but approximately 17 billion are active for a given token, with 16 experts in the model. This sparse design is intended to provide more capability than a dense model with a similar active size while reducing the computation required for each token.

The model is part of Meta’s Llama 4 family and is distributed as base and instruction-tuned checkpoints. The documented instruction-tuned identifier is meta-llama/Llama-4-Scout-17B-16E-Instruct. Meta distributes the weights under the Llama 4 Community License Agreement, which is a custom license rather than a conventional open-source license. Organizations should review the license and acceptable-use requirements before deploying Scout commercially or at large scale.

The main differentiator: a 10-million-token context window

Scout’s most distinctive specification is its 10-million-token context window. A context window is the amount of text, image information, conversation history, or other prompt content that a model can consider in one interaction. For comparison, the research supplied for this model identifies 128,000 tokens as the context length of earlier Llama 3 models.

In practical terms, Scout can be used for tasks such as examining very large document collections, navigating substantial codebases, comparing multiple long reports, or retaining an unusually large amount of conversation and activity history. It can also support workflows in which image content is analyzed alongside lengthy textual instructions or reference material.

The maximum context length is not the same as a guarantee that every deployment will process 10 million tokens quickly or cheaply. Memory consumption, attention implementation, image-token overhead, batching, quantization, hardware, and serving software all affect real-world performance. A deployment may support the nominal limit while still becoming slow or expensive near that limit, so organizations should benchmark representative workloads rather than treating the headline number as a practical default.

Supported modalities and capabilities

Scout is natively multimodal on the input side. It supports multilingual text and images within a unified model architecture, rather than relying only on a separate image-captioning system. Its native output is text and code; it does not directly generate images, audio, video, or speech.

  • Text input: Supported.
  • Image input: Supported, including visual questions, captions, recognition, documents, charts, and image grounding.
  • Text output: Supported.
  • Code output: Supported.
  • Image, audio, video, music, speech, and embedding output: Not identified as native Scout outputs in the supplied specifications.

Meta reports that Scout was pretrained with prompts containing up to 48 images and was tested successfully during post-training with up to eight images. These figures describe reported training and evaluation conditions, not necessarily a universal practical limit for every serving implementation.

Image understanding is useful for more than describing photographs. The documented use cases include visual question answering, chart interpretation, document understanding, image recognition, captioning, and grounding. For example, an application could provide a scanned report, ask Scout to locate a particular figure, and then request a textual explanation of the chart. The quality of such workflows will depend on image resolution, prompt design, preprocessing, and the serving environment.

Reasoning, coding, and tool use

Scout’s instruction-tuned version is intended for reasoning over text and images, including multi-step analysis of documents, charts, and visual questions. Meta’s model card reports results on reasoning, coding, multilingual, image-understanding, and long-context evaluations. Reported instruction-model results include 74.3 on MMLU-Pro, 57.2 on GPQA Diamond, 32.8 on LiveCodeBench, and 94.4 on DocVQA. These are provider-reported benchmark results under the associated evaluation setup; they should not be treated as a guarantee for a particular application.

The model also supports coding tasks such as code generation, explanation, transformation, and codebase exploration. Its unusually large context is especially relevant when a coding assistant needs to inspect many files or preserve more project context than a conventional context window allows. Scout does not automatically make a complete software engineering system, however. Repository indexing, execution, testing, permissions, and result verification still belong to the surrounding application.

Scout’s instruction format documents zero-shot function and tool calling. It can produce a tool-call expression when functions are supplied in the prompt, but the model does not execute external tools by itself. An application must interpret the expression, call the relevant service, and return the result to the model. Scout therefore supports tool-oriented orchestration rather than independent web browsing or autonomous tool execution.

Deployment, memory, and cost trade-offs

Meta’s reference materials state that the BF16 model can be quantized on the fly to int4 and fit on a single 80 GB H100 GPU. BF16 deployment requires substantially more memory, while FP8 and int4 quantization can reduce memory requirements with a reported small quality trade-off. The exact speed, throughput, and quality impact will depend on the inference engine, batch size, prompt length, hardware, and quantization method.

This deployment profile gives Scout a different cost structure from a hosted per-token model. The canonical Meta distribution is downloadable weights, and no single official Meta per-token input or output price is provided in the supplied research. The operator instead pays for infrastructure, storage, engineering, power, and serving operations. A self-hosted deployment can be attractive when data control, customization, or high-volume usage matters, but it also transfers operational responsibility to the organization.

The 17-billion active-parameter design may reduce per-token computation relative to a dense 109-billion-parameter model, but the total model still has substantial memory and infrastructure requirements. Quantization can make deployment more accessible, although it may introduce quality or compatibility trade-offs. Scout should therefore be evaluated on total workload cost rather than parameter count alone.

Main strengths and limitations

Strengths

  • Very long context: The 10-million-token context window is well suited to large documents, codebases, and multi-document workflows.
  • Native image understanding: Text and images can be processed together for visual question answering, chart analysis, document understanding, and image grounding.
  • Open-weight deployment: Developers can download and operate the model through supported distribution channels instead of relying on one mandatory hosted endpoint.
  • Efficient sparse architecture: Only about 17 billion of 109 billion total parameters are active for each token.
  • Tool-oriented prompting: The instruction format supports function and tool-call expressions for applications that provide their own orchestration layer.
  • Customization potential: The model is available for fine-tuning workflows according to the supplied model data, subject to the license and technical requirements.

Limitations

  • Text-only output: Scout does not natively generate images, audio, video, or speech.
  • No independent browsing or execution: Web search, external APIs, code execution, and other actions require an application layer.
  • Operational complexity: Self-hosting requires suitable hardware, quantization or serving expertise, monitoring, and performance testing.
  • Context is not free: Very long prompts can increase latency and memory use, even when they fit within the stated limit.
  • Knowledge cutoff: The pretraining-data cutoff is August 2024. Retrieval or external tools can provide newer information at runtime, but do not update the underlying model.
  • License review is necessary: The Llama 4 Community License Agreement and acceptable-use requirements should be checked for the intended deployment.

Best use cases for Llama 4 Scout

Scout is a strong candidate when an application needs several of its distinguishing properties at the same time: long context, image understanding, downloadable weights, and text or code generation. Suitable examples include:

  • Summarizing and querying very large document collections.
  • Analyzing reports, scanned documents, diagrams, and charts.
  • Exploring large software repositories or reviewing related files together.
  • Building multilingual assistants that combine text and images.
  • Creating visual customer-support workflows with a separate retrieval or business-system layer.
  • Generating synthetic data or transforming documents inside a controlled environment.
  • Deploying customized inference where an organization wants greater control over model hosting and data flow.

When to choose this model

Choose Llama 4 Scout when the 10-million-token context window is materially useful, when image input is important, or when downloadable weights and self-managed deployment are preferable to a fixed hosted API. It is particularly compelling for long-document and codebase analysis where repeatedly selecting small excerpts would lose useful context.

A smaller or more conventional model may be more appropriate when requests are short, latency and infrastructure simplicity matter more than context size, or the application does not need image understanding. A hosted model may also be preferable when the team wants predictable API operations and does not want to manage GPUs, quantization, scaling, and monitoring. Conversely, a dedicated image, speech, video, or embedding model is a better fit when those outputs are the primary requirement.

Scout should not be selected solely because its context limit is large. Test it with the documents, images, languages, code, and latency targets that matter to the application. The most useful comparison is not just benchmark performance, but the complete trade-off among answer quality, prompt length, inference cost, deployment control, and engineering effort.

Pricing and availability

Llama 4 Scout is available as open weights through Meta and distribution partners including Hugging Face. The supplied research does not identify a canonical official hosted input or output price from Meta, so there is no verified per-token price to report. Costs depend on how the weights are obtained and served, including hardware, hosting, storage, quantization, and operational requirements.

In summary, Scout is best understood as a long-context, vision-capable open-weight model rather than a turnkey multimodal media generator or autonomous agent. Its combination of a sparse architecture, image understanding, tool-call formatting, and a 10-million-token context window makes it especially relevant to organizations building specialized assistants and analysis systems around their own infrastructure.


Answers to Frequently Asked Questions

How large is Llama 4 Scout’s context window?
Llama 4 Scout has a reported context window of up to 10 million tokens, enabling analysis of very large document collections, codebases, reports, and conversation histories. Real-world speed, memory use, and cost may be significantly higher near the maximum.
What is Llama 4 Scout?
Llama 4 Scout is an open-weight multimodal language model from Meta that accepts text and images and generates text and code. It is designed for tasks such as visual question answering, document analysis, coding assistance, multilingual interaction, and tool-oriented workflows.
What modalities does Llama 4 Scout support?
Scout supports text and image inputs and produces text and code outputs. It can analyze documents, charts, diagrams, and photographs, but it does not natively generate images, audio, video, speech, or embeddings.
How can Llama 4 Scout be deployed and what does it cost?
Scout is distributed as downloadable weights through Meta and partners such as Hugging Face. Meta’s reference materials state that the BF16 model can be quantized to int4 and fit on a single 80 GB H100 GPU. There is no canonical official per-token price in the supplied information; total costs depend on hardware, hosting, storage, quantization, and operations.
Can Llama 4 Scout use tools or browse the web autonomously?
Scout supports zero-shot function and tool-call expressions, but it does not execute tools or browse the web independently. An external application must interpret the model’s tool call, invoke the relevant service, and return the result.


Sources 7
Provider

About Meta AI