What is Llama 4 Scout?
Llama 4 Scout is an open-weight multimodal language model released by Meta on April 5, 2025. It is designed to accept text and images and produce text and code. The instruction-tuned version is intended for assistant-style use, including visual question answering, document analysis, coding assistance, multilingual interaction, and tool-oriented workflows.
Scout uses a mixture-of-experts architecture. Instead of activating every parameter for every token, the model routes each token through a subset of specialized neural networks called experts. Scout has 109 billion total parameters, but approximately 17 billion are active for a given token, with 16 experts in the model. This sparse design is intended to provide more capability than a dense model with a similar active size while reducing the computation required for each token.
The model is part of Meta’s Llama 4 family and is distributed as base and instruction-tuned checkpoints. The documented instruction-tuned identifier is meta-llama/Llama-4-Scout-17B-16E-Instruct. Meta distributes the weights under the Llama 4 Community License Agreement, which is a custom license rather than a conventional open-source license. Organizations should review the license and acceptable-use requirements before deploying Scout commercially or at large scale.
The main differentiator: a 10-million-token context window
Scout’s most distinctive specification is its 10-million-token context window. A context window is the amount of text, image information, conversation history, or other prompt content that a model can consider in one interaction. For comparison, the research supplied for this model identifies 128,000 tokens as the context length of earlier Llama 3 models.
In practical terms, Scout can be used for tasks such as examining very large document collections, navigating substantial codebases, comparing multiple long reports, or retaining an unusually large amount of conversation and activity history. It can also support workflows in which image content is analyzed alongside lengthy textual instructions or reference material.
The maximum context length is not the same as a guarantee that every deployment will process 10 million tokens quickly or cheaply. Memory consumption, attention implementation, image-token overhead, batching, quantization, hardware, and serving software all affect real-world performance. A deployment may support the nominal limit while still becoming slow or expensive near that limit, so organizations should benchmark representative workloads rather than treating the headline number as a practical default.
Supported modalities and capabilities
Scout is natively multimodal on the input side. It supports multilingual text and images within a unified model architecture, rather than relying only on a separate image-captioning system. Its native output is text and code; it does not directly generate images, audio, video, or speech.
- Text input: Supported.
- Image input: Supported, including visual questions, captions, recognition, documents, charts, and image grounding.
- Text output: Supported.
- Code output: Supported.
- Image, audio, video, music, speech, and embedding output: Not identified as native Scout outputs in the supplied specifications.
Meta reports that Scout was pretrained with prompts containing up to 48 images and was tested successfully during post-training with up to eight images. These figures describe reported training and evaluation conditions, not necessarily a universal practical limit for every serving implementation.
Image understanding is useful for more than describing photographs. The documented use cases include visual question answering, chart interpretation, document understanding, image recognition, captioning, and grounding. For example, an application could provide a scanned report, ask Scout to locate a particular figure, and then request a textual explanation of the chart. The quality of such workflows will depend on image resolution, prompt design, preprocessing, and the serving environment.
Reasoning, coding, and tool use
Scout’s instruction-tuned version is intended for reasoning over text and images, including multi-step analysis of documents, charts, and visual questions. Meta’s model card reports results on reasoning, coding, multilingual, image-understanding, and long-context evaluations. Reported instruction-model results include 74.3 on MMLU-Pro, 57.2 on GPQA Diamond, 32.8 on LiveCodeBench, and 94.4 on DocVQA. These are provider-reported benchmark results under the associated evaluation setup; they should not be treated as a guarantee for a particular application.
The model also supports coding tasks such as code generation, explanation, transformation, and codebase exploration. Its unusually large context is especially relevant when a coding assistant needs to inspect many files or preserve more project context than a conventional context window allows. Scout does not automatically make a complete software engineering system, however. Repository indexing, execution, testing, permissions, and result verification still belong to the surrounding application.
Scout’s instruction format documents zero-shot function and tool calling. It can produce a tool-call expression when functions are supplied in the prompt, but the model does not execute external tools by itself. An application must interpret the expression, call the relevant service, and return the result to the model. Scout therefore supports tool-oriented orchestration rather than independent web browsing or autonomous tool execution.
Deployment, memory, and cost trade-offs
Meta’s reference materials state that the BF16 model can be quantized on the fly to int4 and fit on a single 80 GB H100 GPU. BF16 deployment requires substantially more memory, while FP8 and int4 quantization can reduce memory requirements with a reported small quality trade-off. The exact speed, throughput, and quality impact will depend on the inference engine, batch size, prompt length, hardware, and quantization method.
This deployment profile gives Scout a different cost structure from a hosted per-token model. The canonical Meta distribution is downloadable weights, and no single official Meta per-token input or output price is provided in the supplied research. The operator instead pays for infrastructure, storage, engineering, power, and serving operations. A self-hosted deployment can be attractive when data control, customization, or high-volume usage matters, but it also transfers operational responsibility to the organization.
The 17-billion active-parameter design may reduce per-token computation relative to a dense 109-billion-parameter model, but the total model still has substantial memory and infrastructure requirements. Quantization can make deployment more accessible, although it may introduce quality or compatibility trade-offs. Scout should therefore be evaluated on total workload cost rather than parameter count alone.
Main strengths and limitations
Strengths
- Very long context: The 10-million-token context window is well suited to large documents, codebases, and multi-document workflows.
- Native image understanding: Text and images can be processed together for visual question answering, chart analysis, document understanding, and image grounding.
- Open-weight deployment: Developers can download and operate the model through supported distribution channels instead of relying on one mandatory hosted endpoint.
- Efficient sparse architecture: Only about 17 billion of 109 billion total parameters are active for each token.
- Tool-oriented prompting: The instruction format supports function and tool-call expressions for applications that provide their own orchestration layer.
- Customization potential: The model is available for fine-tuning workflows according to the supplied model data, subject to the license and technical requirements.
Limitations
- Text-only output: Scout does not natively generate images, audio, video, or speech.
- No independent browsing or execution: Web search, external APIs, code execution, and other actions require an application layer.
- Operational complexity: Self-hosting requires suitable hardware, quantization or serving expertise, monitoring, and performance testing.
- Context is not free: Very long prompts can increase latency and memory use, even when they fit within the stated limit.
- Knowledge cutoff: The pretraining-data cutoff is August 2024. Retrieval or external tools can provide newer information at runtime, but do not update the underlying model.
- License review is necessary: The Llama 4 Community License Agreement and acceptable-use requirements should be checked for the intended deployment.
Best use cases for Llama 4 Scout
Scout is a strong candidate when an application needs several of its distinguishing properties at the same time: long context, image understanding, downloadable weights, and text or code generation. Suitable examples include:
- Summarizing and querying very large document collections.
- Analyzing reports, scanned documents, diagrams, and charts.
- Exploring large software repositories or reviewing related files together.
- Building multilingual assistants that combine text and images.
- Creating visual customer-support workflows with a separate retrieval or business-system layer.
- Generating synthetic data or transforming documents inside a controlled environment.
- Deploying customized inference where an organization wants greater control over model hosting and data flow.
When to choose this model
Choose Llama 4 Scout when the 10-million-token context window is materially useful, when image input is important, or when downloadable weights and self-managed deployment are preferable to a fixed hosted API. It is particularly compelling for long-document and codebase analysis where repeatedly selecting small excerpts would lose useful context.
A smaller or more conventional model may be more appropriate when requests are short, latency and infrastructure simplicity matter more than context size, or the application does not need image understanding. A hosted model may also be preferable when the team wants predictable API operations and does not want to manage GPUs, quantization, scaling, and monitoring. Conversely, a dedicated image, speech, video, or embedding model is a better fit when those outputs are the primary requirement.
Scout should not be selected solely because its context limit is large. Test it with the documents, images, languages, code, and latency targets that matter to the application. The most useful comparison is not just benchmark performance, but the complete trade-off among answer quality, prompt length, inference cost, deployment control, and engineering effort.
Pricing and availability
Llama 4 Scout is available as open weights through Meta and distribution partners including Hugging Face. The supplied research does not identify a canonical official hosted input or output price from Meta, so there is no verified per-token price to report. Costs depend on how the weights are obtained and served, including hardware, hosting, storage, quantization, and operational requirements.
In summary, Scout is best understood as a long-context, vision-capable open-weight model rather than a turnkey multimodal media generator or autonomous agent. Its combination of a sparse architecture, image understanding, tool-call formatting, and a 10-million-token context window makes it especially relevant to organizations building specialized assistants and analysis systems around their own infrastructure.

