What is Llama 3.2 90B Vision Instruct?
Llama 3.2 90B Vision Instruct is Meta's instruction-tuned, open-weight model for understanding images alongside text. The official model card lists 88.8 billion parameters, although the model is commonly described as a 90-billion-parameter model. It is the larger model in the Llama 3.2 Vision family, alongside the smaller 11B vision model.
The model accepts text and image inputs and generates text. In practical terms, a developer can provide an image with a question such as “What information does this chart show?” or “Summarize this document,” and the model will return a written answer. Its intended work includes visual question answering, image captioning, chart and graph interpretation, map analysis, document understanding, and visual content extraction.
This is not an image-generation, speech, music, or video-generation model. Its multimodal capability is focused on visual understanding, with text as the output format.
How the vision architecture works
Llama 3.2 90B Vision Instruct is based on the Llama 3.1 text model and adds a separately trained vision adapter. The adapter connects information from an image encoder to the language model through cross-attention layers. This allows visual features to influence the generated response while retaining the language model's broader text-processing abilities.
For a user, the important distinction is that the model does not simply describe an image using a separate captioning system. It can combine visual information with written instructions and answer questions about relationships, layout, visible text, charts, objects, and other elements in the supplied image. Results still depend on image quality, visual ambiguity, document layout, and the model's ability to localize the relevant information.
Verified technical specifications
| Specification | Details |
|---|---|
| Provider | Meta |
| Canonical model ID | meta-llama/Llama-3.2-90B-Vision-Instruct |
| Release date | September 25, 2024 |
| Model family | Llama 3.2 Vision |
| Parameters | Approximately 90B; the official model card lists 88.8B |
| Context length | 128,000 tokens |
| Inputs | Text and images |
| Output | Text |
| Knowledge cutoff | December 2023 |
| License | Llama 3.2 Community License |
The 128K context window is a documented maximum context length, but it should not be interpreted as a guarantee that every deployment will expose the full limit. Hosted services and inference frameworks may impose their own request, image, memory, or output restrictions. No maximum output-token value was verified in the supplied research, so deployments should be checked individually.
Supported inputs and outputs
The model supports text and images as inputs and text as output. This makes it suitable for tasks where visual information must be converted into an explanation, answer, classification, caption, or extracted text. Examples include asking for the key figures in a business chart, identifying the sections of a photographed form, describing an image for accessibility, or answering questions about a map.
It does not natively produce images, audio, video, music, embeddings, or speech. If the required result is an edited image, a generated video, or spoken audio, a model designed for that output type would be more appropriate. Llama 3.2 90B Vision Instruct can potentially serve as one component in a larger workflow, but those other media-generation capabilities should not be attributed to this model itself.
Reasoning, coding, and tool support
The model is designed for visual reasoning: it can use visual details together with textual instructions to answer questions or explain what an image contains. Its strongest reasoning use cases are grounded in the supplied visual material, such as interpreting a graph, locating information in a document, or comparing visible elements. It does not have current knowledge beyond its December 2023 training cutoff unless an application supplies external information.
Coding capability is available in the broader language-model behavior and is useful for explaining code, generating code, or helping build an application around image analysis. The supplied comparative assessment rates coding at 7 out of 10, but that is an editorial estimate rather than a Meta-published benchmark. Coding quality should therefore be tested against the programming languages, libraries, and codebase used in a particular project.
Meta's vision prompt documentation describes tool-calling formats for text-only prompts, including code-interpreter and search-oriented tools in the reference implementation. However, the documentation states that tool calling does not work when images are included in the prompt. This means developers should not treat the model as an image-grounded autonomous browsing or action system. Tool use is better understood as a reference-runtime or integration feature, with important restrictions for multimodal requests.
Deployment options and pricing
Meta distributes Llama 3.2 90B Vision Instruct as downloadable weights through its Llama distribution channels and an official gated Hugging Face repository. The repository documents use with tools including Transformers, vLLM, and SGLang. Some serving frameworks can expose an OpenAI-compatible interface, but that interface belongs to the serving framework or hosting provider rather than representing a native Meta API identity.
The model is large and resource-intensive. Its weight files are distributed in multiple large shards, so practical deployment generally requires substantial accelerator memory, quantization, tensor parallelism, or a hosted inference service. A smaller quantized deployment may reduce memory requirements, but the supplied research does not verify a particular hardware configuration or performance level.
No universal first-party Meta-hosted API price was verified. There is therefore no reliable standard input or output price to list for the model itself. Costs depend on whether the model is run on owned hardware, through a cloud GPU service, or through a third-party inference provider. Any provider-specific price should be checked separately because it may include infrastructure, storage, queueing, or platform charges.
Main strengths and limitations
The main strength of Llama 3.2 90B Vision Instruct is the combination of a large language model with image understanding and open-weight distribution. Developers can inspect, adapt, fine-tune, and deploy the model through the Llama ecosystem rather than relying exclusively on a closed, fixed vendor endpoint. The large context window is also useful for workflows that combine lengthy instructions or documents with visual material.
- Strong fit for visual analysis: The model is designed for image questions, captions, charts, maps, and documents.
- Open-weight deployment: Downloadable weights enable self-hosted, partner-hosted, and customized deployments subject to the license.
- Large context: The documented 128K-token context supports long textual prompts and document-oriented workflows.
- Fine-tuning potential: Meta identifies ecosystem tools such as torchtune for adapting the vision models.
- Broad language-model behavior: The model can also support text-based explanation, analysis, and coding tasks.
Its limitations are equally important. The model is expensive and demanding to run compared with smaller vision-language models. It has no verified universal API price, no current-events knowledge beyond December 2023, and no native non-text generation. Visual answers can be wrong when an image is unclear, information is small or poorly laid out, or the scene is ambiguous. Like other foundation models, it can produce biased, unsafe, overconfident, or factually incorrect responses.
Tool calling has a specific multimodal restriction: the documented tool formats are for text-only prompts and do not work with images in the prompt. Applications needing image analysis and external search or actions may need to separate the visual-analysis step from the tool-use step and add their own orchestration, validation, and safety controls.
Best use cases
- Visual question answering over photographs, screenshots, diagrams, and illustrations
- Chart, graph, map, and dashboard interpretation
- Document understanding and extraction from supplied images
- Image captioning and accessibility-oriented descriptions
- Multimodal research and custom fine-tuning
- Self-hosted or partner-hosted applications that require open model weights
- Text-and-image workflows where a large context window is useful
For production document processing, developers should evaluate representative samples rather than assuming that a large parameter count guarantees accurate extraction. Low-resolution scans, tables, handwriting, unusual layouts, and small text should receive specific testing.
When to choose this model
Choose Llama 3.2 90B Vision Instruct when image understanding is central to the application and you value downloadable weights, customization, or deployment control. It is particularly suitable when a team can provide the accelerator infrastructure or use a hosted Llama-compatible service, and when the application benefits from a large model for complex visual and textual questions.
A smaller vision-language model may be more appropriate when response speed, memory usage, or operating cost matters more than maximum capability. The 11B member of the same Llama 3.2 Vision family is the most directly supported sibling alternative, although the supplied research does not provide a detailed performance comparison between the two. A closed hosted vision model may be preferable when a team wants a managed endpoint, current-information integrations, predictable service operations, or less infrastructure work.
Use a different model or a multi-model pipeline when the required output is an image, audio, video, or speech; when current information is essential and retrieval cannot be added; or when image-grounded tool execution must work in a single request. For high-risk document or decision workflows, human review and application-level validation remain necessary regardless of the selected model.
Availability and license considerations
The model is available through Meta's official Llama catalog and the gated Hugging Face repository. Access requires accepting the applicable terms and license. The Llama 3.2 Community License permits use, reproduction, distribution, and modification subject to its conditions, acceptable-use requirements, attribution obligations, and additional commercial terms for products exceeding the stated monthly-active-user threshold.
Before deployment, review the current license, acceptable-use policy, hosting provider terms, and any obligations that apply to redistribution or commercial scale. The model's open-weight status provides more deployment flexibility than a closed API, but it does not remove the need to manage infrastructure, security, privacy, evaluation, and content safeguards.

