Model catalog

Meta AI Models

Browse the AI models associated with Meta AI. Compare current and historical models by family, capabilities, context window, availability and intended use.

34 models tracked
34 Total models
21 Model families
6 Model types
34 Current / accessible
All models

Meta AI model catalog

Meta AI logo
Canopy Height Maps

Canopy Height Maps v2

Forest mapping, canopy-height estimation, restoration monitoring, ecological analysis, and geospatial research using satellite imagery

Type Other
Reasoning 2/10
Speed 6/10
Image input
Status

Current; open-source research model with gated model-weight access

Input No official hosted API pricing published; intended for local or research use
Output No official hosted API pricing published; outputs are canopy-height prediction tensors or rasters
Meta AI logo
DINOv3

DINOv3

Image embeddings, dense feature extraction, image retrieval, classification, segmentation, depth estimation, object discovery, video tracking pipelines, and geospatial computer vision

Type Multimodal
Reasoning 0/10
Speed 7/10
Image input Fine-tuning
Status

Current; downloadable open-weight research model suite

View model →
Meta AI logo
Llama 3.1

Llama 3.1 8B

Local inference, fine-tuning, private deployment, text generation, retrieval-augmented generation, research, and cost-sensitive applications

Type General Purpose
Context 131K
Reasoning 6/10
Speed 8/10
Fine-tuning Streaming
Status

Current open-weight static model; downloadable and usable through compatible local or hosted inference deployments

View model →

High-quality open-weight research, multilingual applications, coding, reasoning, synthetic-data generation, model distillation and self-hosted deployments with substantial infrastructure

Type General Purpose
Context 131K
Reasoning 9/10
Speed 3/10
Tool use Fine-tuning Streaming
Status

Available as an open-weight model; older generation superseded by newer Llama releases

View model →
Meta AI logo
Llama 3.2

Llama 3.2 1B

Private local inference, mobile and edge assistants, summarization, rewriting, retrieval-supported generation, and lightweight multilingual applications

Type Lightweight
Context 128K
Reasoning 3/10
Speed 9/10
Tool use Fine-tuning Streaming
Status

Available downloadable open-weight model; static checkpoint

Input No official Meta-hosted API price; downloadable weights are available under the Llama 3.2 Community License
Output No official Meta-hosted API price; deployment and inference costs depend on hardware or third-party hosting
View model →
Meta AI logo
Llama 3.2

Llama 3.2 3B

Private local inference, multilingual text generation, edge applications, model adaptation, and fine-tuning

Type Lightweight
Context 128K
Reasoning 5/10
Speed 8/10
Tool use Fine-tuning
Status

Available; static pretrained open-weight model

Input No official Meta-hosted API price; downloadable weights
Output No official Meta-hosted API price; downloadable weights
View model →

Visual question answering, image reasoning, chart and document understanding, image captioning, multimodal research, and self-hosted or partner-hosted AI applications.

Type Multimodal
Context 128K
Reasoning 8/10
Speed 4/10
Multimodal Image input Tool use
Status

Available open-weight model; static model trained on an offline dataset

View model →

Self-hosted visual question answering, image captioning, document analysis, visual reasoning, and multimodal assistants

Type Multimodal
Context 128K
Reasoning 7/10
Speed 6/10
Multimodal Image input Tool use
Status

Available open-weight static model

View model →

Self-hosted or hosted multilingual chat, coding assistance, long-context text generation, tool calling, synthetic data, and applications requiring open model weights

Type General Purpose
Context 128K
Reasoning 8/10
Speed 6/10
Tool use Fine-tuning
Status

Available open-weight model; static model trained on an offline dataset

Input No universal Meta-hosted per-token price; downloadable weights and third-party hosting options are available
Output No universal Meta-hosted per-token price; downloadable weights and third-party hosting options are available
View model →

Open-weight multimodal assistants, image understanding, visual question answering, coding, multilingual applications, creative writing and long-context text processing.

Type Multimodal
Context 1M
Reasoning 8/10
Speed 7/10
Multimodal Image input Tool use
Status

Available; open-weight static checkpoint

View model →

Long-context document and code analysis, visual question answering, multimodal assistants, multilingual applications, self-hosted inference, and customized deployments

Type Multimodal
Context 10M
Reasoning 7/10
Speed 8/10
Multimodal Image input Tool use
Status

Available; open-weight model

View model →
Meta AI logo
Llama Guard

Llama Guard 4

Text and image moderation for prompts and generated responses in generative AI systems

Type Other
Context 8K
Reasoning 2/10
Speed 6/10
Multimodal Image input
Status

Current and available as an open-weight model

View model →
Meta AI logo
Llama Guard 3

Llama Guard 3-1B

Low-cost prompt and response safety classification, local moderation, mobile and edge deployments, and customizable LLM guardrails

Type Other
Context 131K
Reasoning 3/10
Speed 9/10
Fine-tuning
Status

Current open-weight model; downloadable subject to access approval

View model →
Meta AI logo
Llama Guard 3

Llama Guard 3-8B

Self-hosted multilingual input and output moderation for LLM applications

Type Other
Context 131K
Reasoning 3/10
Speed 6/10
Fine-tuning
Status

Available open-weight safety classifier

View model →

Safety classification of mixed text-and-image prompts and text responses in multimodal LLM systems

Type Other
Context 128K
Reasoning 4/10
Speed 5/10
Multimodal Image input
Status

Available open-weight multimodal safety model; Meta's current model repositories continue to list it.

View model →
Meta AI logo
Llama Prompt Guard 2

Llama Prompt Guard 2 22M

Low-latency detection of prompt injections and jailbreak attempts in LLM applications, agents, retrieved documents, and other untrusted text

Type Other
Context 512
Reasoning 1/10
Speed 9/10
Fine-tuning
Status

Current open-weight safety classifier

View model →
Meta AI logo
Llama Prompt Guard 2

Llama Prompt Guard 2 86M

Multilingual prompt-injection detection, jailbreak screening, agent security, and filtering untrusted text before it reaches an LLM

Type Other
Context 512
Reasoning 1/10
Speed 8/10
Fine-tuning
Status

Current open-weight model; gated download access

View model →
Meta AI logo
Muse Glimmer

Muse Glimmer

Local agents, long-running tool workflows, coding assistants, multimodal document and screenshot understanding, private on-device inference, and model customization

Type Reasoning
Context 131K
Reasoning 8/10
Speed 7/10
Multimodal Image input Tool use
Status

Current open-weight model; self-hosted and available through selected third-party hosted inference providers

Input No official Meta-hosted API token price; open weights can be downloaded and self-hosted without per-token charges
Output No official Meta-hosted API token price; third-party hosted inference pricing varies by provider
View model →
Meta AI logo
Muse Image

Muse Image 1.0

Text-to-image generation, precise image editing, multi-image composition, anchored visual series, product imagery, creative assets, and grounded visual content.

Type Multimodal
Reasoning 8/10
Speed 7/10
Multimodal Image input Media output
Status

Current and available through Meta Model API; also available in selected Meta AI consumer experiences

Output $0.01 per successfully generated image
View model →
Meta AI logo
Muse Spark

Muse Spark 1.1

Agentic workflows, coding agents, computer-use automation, multimodal document and media analysis, long-context reasoning, tool orchestration, and web-grounded applications.

Type Reasoning
Context 1.05M
Reasoning 8/10
Speed 8/10
Multimodal Image input Audio input
Status

Current but superseded by Muse Spark 1.2 and Muse Spark 1.3; available on the Meta Model API Standard tier in public preview for US developers.

Input $1.25 per 1 million input tokens; cached input $0.15 per 1 million tokens
Output $4.25 per 1 million output tokens
View model →
Meta AI logo
Muse Spark

Muse Spark 1.2

Long-horizon coding agents, repository-scale software engineering, multimodal code generation, debugging, refactoring, and tool-driven workflows

Type Coding
Context 1.05M
Reasoning 8/10
Speed 7/10
Multimodal Image input Audio input
Status

Available; previous version, with Muse Spark 1.3 recommended for new work

Input $1.25 per 1 million input tokens; $0.15 per 1 million cached input tokens
Output $4.25 per 1 million output tokens
View model →
Meta AI logo
Muse Spark

Muse Spark 1.3

Long-horizon coding agents, software engineering, browser and computer-use workflows, tool orchestration, large repositories, document analysis, and multimodal reasoning

Type Multimodal
Context 1.05M
Reasoning 9/10
Speed 8/10
Multimodal Image input Audio input
Status

Current; available through Meta Model API and Muse Code

Input $1.25 per 1M input tokens; $0.15 per 1M cached input tokens
Output $4.25 per 1M output tokens
View model →

Real-time speech-to-text, live captions, voice agents, meeting transcription, call intelligence, dictation, and speaker-aware transcription

Type Other
Reasoning 1/10
Speed 9/10
Audio input Streaming
Status

Current and available through Meta Model API

Input $3.00 per 1,000 minutes ($0.18 per hour)
View model →
Meta AI logo
Omnilingual wav2vec 2.0

Omnilingual wav2vec 2.0

Multilingual speech representation learning, audio embeddings, low-resource language research, and custom downstream speech systems

Type Other
Reasoning 1/10
Speed 5/10
Audio input Fine-tuning
Status

Current; open-source research model family

Input Free to download and self-host; no official hosted API price published
Output Free to download and self-host; no official hosted API price published
View model →

Cross-modal audio-video-text retrieval, audiovisual embeddings, sound-event understanding, media indexing, and multimodal perception systems.

Type Multimodal
Reasoning 2/10
Speed 7/10
Multimodal Audio input Video input
Status

Current and openly available

Input No official hosted API pricing; open checkpoints are available for self-managed use.
Output No official hosted API pricing; the model returns embeddings rather than generated media.
View model →
Meta AI logo
SAM 3D

SAM 3D Body

Single-image 3D human mesh recovery, pose and shape estimation, AR/VR, robotics perception, and computer-vision research

Type Other
Reasoning 2/10
Speed 4/10
Image input Media output
Status

Current research release; downloadable checkpoints with gated Hugging Face access

Input No official hosted API pricing published
Output No official hosted API pricing published

Single-image reconstruction of textured 3D objects from natural scenes, 3D computer-vision research, Gaussian-splat workflows, and rapid asset prototyping

Type Other
Reasoning 2/10
Speed 3/10
Image input Media output
Status

Available research release; gated model checkpoints

View model →
Meta AI logo
SAM Audio

SAM Audio

Prompted audio separation, speech and noise isolation, instrument and vocal extraction, audiovisual sound segmentation, and audio-editing research

Type Multimodal
Reasoning 2/10
Speed 7/10
Multimodal Audio input Video input
Status

Current open research release; downloadable checkpoints and public demo available

Input No hosted API pricing published; downloadable checkpoints available under the SAM License
Output No hosted API pricing published; model outputs separated target and residual audio
View model →

Reference-free evaluation and benchmarking of text-guided audio-separation outputs

Type Other
Reasoning 2/10
Speed 4/10
Multimodal Audio input
Status

Current gated research release

View model →
Meta AI logo
Seamless

SeamlessExpressive

Noncommercial research on expressive multilingual speech-to-speech translation, prosody transfer and voice-style preservation

Type Multimodal
Reasoning 2/10
Speed 3/10
Audio input Media output
Status

Available as a gated research release

Meta AI logo
Seamless

SeamlessStreaming

Real-time multilingual speech recognition, simultaneous translation, speech-to-text translation, and speech-to-speech translation

Type Multimodal
Reasoning 1/10
Speed 8/10
Audio input Media output Streaming
Status

Open-source research model; publicly available

Multilingual automatic speech recognition, speech-to-text translation, text translation, text-to-speech translation, and speech-to-speech translation

Type Multimodal
Reasoning 1/10
Speed 5/10
Multimodal Audio input Media output
Status

Available open-weight research model; noncommercial research use

Input Not applicable; no official hosted API pricing published
Output Not applicable; no official hosted API pricing published
View model →
Meta AI logo
Segment Anything

SAM 3.1

Open-vocabulary object detection, pixel-level image segmentation, and multi-object video tracking

Type Other
Reasoning 2/10
Speed 9/10
Multimodal Image input Video input
Status

Current and available; hosted through Meta Model API and available as released research checkpoints

Input $2.50 per 1,000 images; $0.20 per 1,000 video frames
Output Included in the image and video segmentation pricing; no separate output-token price
View model →

Computational neuroscience, fMRI response prediction, brain encoding, multisensory research, and in-silico experiment design

Type Multimodal
Reasoning 0/10
Speed 0/10
Multimodal Image input Audio input
Status

Open-weight research release

View model →