Segment Anything

SAM 3.1

by Meta AI · Current and available; hosted through Meta Model API and available as released research checkpoints

Meta’s SAM 3.1 detects, segments, and tracks objects in images and video using text and visual prompts. Object Multiplex improves efficiency for dense multi-object tracking, while hosted pricing is based on submitted images and video frames. The model is specialized for visual perception rather than conversation, coding, or media generation.

Image generation Reasoning Coding
SAM 3.1 is Meta’s current Segment Anything model for promptable visual perception. It accepts text and visual prompts, returns pixel-level segmentation masks and bounding boxes, and tracks object identities through video. The March 27, 2026 update introduced Object Multiplex, a shared-memory approach that substantially improves multi-object tracking efficiency.
Outputs

What SAM 3.1 can produce

Image generation
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Fine-tuning Structured output Multimodal output
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
9/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Segment Anything
Model type Other
Context window tokens
Maximum output tokens
Release date 2026-03-27
Status Current and available; hosted through Meta Model API and available as released research checkpoints
Knowledge cutoff notes

No authoritative model-specific knowledge-cutoff date was identified. SAM 3.1 is a perception model whose operation is based on supplied image and video inputs rather than a documented language-model training cutoff.

Model notes

SAM 3.1 is a specialized computer-vision model rather than a text-generation model. It introduces Object Multiplex, which jointly processes multiple tracked objects using shared memory. Meta reports approximately a sevenfold speedup at 128 objects on one H100 compared with the November 2025 SAM 3 release. The official research implementation describes an 848-million-parameter SAM 3 architecture and provides SAM 3.1 checkpoints through the associated model repository. Hosted Meta Model API pricing is measured per image and video frame. Image output refers to pixel-level segmentation masks and related visual perception results, not synthesized images. The official repository supports fine-tuning workflows, but hosted API fine-tuning availability is not established here.

Cost

Model pricing

Input $2.50 per 1,000 images; $0.20 per 1,000 video frames
Output Included in the image and video segmentation pricing; no separate output-token price
Model guide

SAM 3.1: Meta’s Open-Vocabulary Model for Fast Multi-Object Video Tracking

SAM 3.1 is Meta’s computer-vision model for detecting, segmenting, and tracking objects in images and video. Its Object Multiplex update jointly processes multiple tracked objects to improve video efficiency and throughput.

What SAM 3.1 does

SAM 3.1 is Meta’s specialized computer-vision model for finding, outlining, and tracking objects in images and video. Instead of producing a conversational answer or generating a new picture, it analyzes visual input and returns structured perception results such as object locations, pixel-level masks, confidence scores, and tracked identities.

The model supports open-vocabulary concept segmentation. This means it is not restricted to a small, fixed list of categories such as “cat” or “car.” A user can provide a short text prompt such as “person,” “red car,” or “hardcover book,” and SAM 3.1 can attempt to locate matching objects. It can also use visual prompts, including points, bounding boxes, masks, and exemplar images or regions.

For video, the model can propagate object masks across frames while maintaining object identities. This makes it suitable for applications in which the important question is not only where an object appears in one frame, but how it moves through a sequence.

Where SAM 3.1 fits in Meta’s model lineup

SAM 3.1 belongs to Meta’s Segment Anything family and is positioned as a perception model rather than a general-purpose language model. Meta’s current AI products also use or reference other model families, including Muse-family systems for consumer assistant experiences and Llama for developer-focused open models, but those are different categories of technology. SAM 3.1 is intended to interpret visual scenes, not to conduct a general conversation, write long-form prose, or act as an all-purpose coding assistant.

Meta released SAM 3.1 on March 27, 2026. It is available through the Meta Model API and through released research checkpoints associated with Meta’s official SAM 3 repository and model repositories. The research implementation describes an 848-million-parameter SAM 3 architecture, while the SAM 3.1 release adds improvements for multi-object video processing.

Why Object Multiplex matters for video

The defining SAM 3.1 improvement is Object Multiplex. In earlier SAM 3 video processing, tracked objects were handled independently, so computation increased approximately linearly as more objects were added. A scene containing many tracked objects could therefore become increasingly expensive or slow.

Object Multiplex groups objects into shared processing buckets and reasons about them jointly using shared memory. The practical goal is to process crowded scenes more efficiently without running a completely separate computation path for every object.

Meta reports approximately a sevenfold speedup at 128 tracked objects on a single H100 GPU compared with the November 2025 SAM 3 release. This is a provider-reported result rather than an independent evaluation, and performance will vary with hardware, video characteristics, prompt type, and implementation. Nevertheless, it identifies the model’s most important use-case advantage: dense multi-object tracking is a central reason to consider SAM 3.1 over a simpler single-object segmentation workflow.

Inputs and outputs

SAM 3.1 accepts images and video, along with text or visual prompts. Visual interaction can include points, bounding boxes, masks, and exemplar-based prompts. Short noun phrases are the most natural text input; the supplied research does not establish that the model is designed for long instructions requiring complex language reasoning.

Its outputs include:

  • Object detections and bounding boxes.
  • Pixel-level segmentation masks that identify which pixels belong to each object.
  • Confidence scores associated with detections or segmentation results.
  • Video tracking results, including object identities propagated across frames.
  • Structured visual-perception data suitable for downstream computer-vision pipelines.

“Image output” in this context means masks and related visual annotations, not synthesized images. SAM 3.1 does not generate ordinary text, audio, music, or video files. It produces video analysis results when given video input, but it is not a video-generation model.

Technical specifications and unavailable limits

SpecificationVerified information
ProviderMeta
Model familySegment Anything
Model typeComputer-vision detection, segmentation, and tracking
Release dateMarch 27, 2026
Parameter informationThe official research implementation describes an 848-million-parameter SAM 3 architecture
Image inputSupported
Video inputSupported
Text promptsSupported
Visual promptsPoints, boxes, masks, and exemplars are supported
Context lengthNot published in the supplied documentation
Maximum output tokensNot applicable or not published; this is not a text-generation model
Tool and function callingNot established; the model is accessed as a perception system rather than a tool-using assistant

The supplied research does not identify a fixed context window, maximum number of output tokens, streaming specification, caching feature, or batch API. Those omissions are important for developers planning an integration: SAM 3.1 should not be evaluated using the same limits normally used for chat models.

Pricing and access

The hosted Meta Model API lists pricing of $2.50 per 1,000 images and $0.20 per 1,000 video frames. This is usage-based pricing, not a recurring subscription, and the supplied information does not identify a separate output-token charge.

Image and video costs should be estimated differently. For image workloads, the relevant unit is the number of submitted images. For video, the number of processed frames can grow rapidly: a long or high-frame-rate video may contain many more billable units than a short clip. A production pipeline should therefore decide whether every frame needs processing and should measure the expected frame volume before estimating cost.

Self-hosted use is also available through released checkpoints and the official research code. The repository lists Python 3.12 or newer, PyTorch 2.7 or newer, and CUDA 12.6 or newer among its installation requirements. Checkpoint access may require authentication and approval through the associated model repository. Self-hosting can provide greater control over data and execution, but it shifts responsibility for GPU capacity, software compatibility, deployment, and performance tuning to the user.

Strengths and limitations

Where SAM 3.1 is strong

  • Open-vocabulary prompting: users can describe target concepts instead of relying only on a fixed class list.
  • Detailed boundaries: pixel-level masks are more useful than rectangular detections when object shape, overlap, or precise area matters.
  • Multiple prompt types: text, points, boxes, masks, and visual exemplars support both automated and interactive workflows.
  • Video identity tracking: the model can follow objects across frames rather than treating every frame as unrelated.
  • Dense-scene efficiency: Object Multiplex is specifically designed to improve throughput when many objects are tracked simultaneously.
  • Deployment flexibility: users can choose a hosted API or released research checkpoints, subject to the applicable access requirements.

What it does not do well or does not provide

  • It is not a general assistant: SAM 3.1 is not intended for conversation, long-form writing, general reasoning, or ordinary code generation.
  • It is not a media generator: it analyzes images and video but does not synthesize images, audio, or video files.
  • Prompt complexity is limited: short noun phrases are better aligned with the documented use case than long descriptions requiring sophisticated language interpretation.
  • Specialized domains may need adaptation: highly technical scientific, medical, or otherwise out-of-domain concepts may require fine-tuning or additional validation.
  • API specifications are incomplete: the supplied documentation does not establish context, token, streaming, caching, or batch limits.
  • Reported speedups are not universal guarantees: Meta’s sevenfold comparison applies to a specific multi-object benchmark setup, including 128 objects and one H100 GPU.

Reasoning, coding, and tool support

SAM 3.1 has limited reasoning in the language-model sense. It can use a visual or textual prompt to determine which regions correspond to a concept, but the supplied research does not describe extended chain-of-thought reasoning, planning, or broad visual question answering. Its useful intelligence is specialized: identifying and separating visual objects consistently across images and frames.

Coding is not a model capability. Developers can write code around SAM 3.1 using the API or research implementation, and the repository provides installation and fine-tuning workflows, but SAM 3.1 itself does not generate or execute code as a coding assistant. Tool use and function calling are likewise not established as model features.

Best use cases

SAM 3.1 is a strong candidate when a system needs promptable visual annotations rather than generated content. Practical applications include:

  • Tracking several people, vehicles, products, or other objects through video.
  • Creating masks for image or video editing workflows.
  • Counting or monitoring objects selected with text or visual examples.
  • Building interactive annotation tools in which a user supplies a point, box, mask, or example.
  • Preparing training data for downstream computer-vision systems.
  • Analyzing crowded scenes where many objects must be tracked at once.

For a small number of objects in a short clip, another segmentation pipeline may be simpler or cheaper. For synthetic media, a generative image or video model is more appropriate. For a chat interface that must explain an image, write code, or combine visual analysis with general reasoning, SAM 3.1 would normally need to be paired with another model.

When to choose SAM 3.1

Choose SAM 3.1 when precise object masks, open-vocabulary prompts, and multi-object video tracking are more important than natural-language generation. It is especially compelling for dense scenes because Object Multiplex targets the computational cost of tracking many objects simultaneously.

Choose the hosted API when predictable integration and usage-based access are more important than infrastructure control. Choose the released checkpoints when self-hosting, customization, or research experimentation justifies managing the required GPU and software environment. In either case, validate the model on the actual objects, camera conditions, lighting, and domain data that matter to the application.

The central trade-off is specialization versus breadth. SAM 3.1 is not a broad AI assistant, but its focused output—object locations, masks, scores, and identities—can be substantially more useful than a general model when the task is visual segmentation and tracking. Its advertised pricing and Object Multiplex improvement make it particularly relevant for workloads that process many image regions or video objects, while its missing language-generation and tool-use capabilities make it unsuitable as a standalone general-purpose model.


Answers to Frequently Asked Questions

What is SAM 3.1 used for?
SAM 3.1 is a computer-vision model for detecting, segmenting, and tracking objects in images and video. It produces object locations, pixel-level masks, confidence scores, and identities that can be used in annotation, monitoring, editing, and downstream vision pipelines.
How does SAM 3.1 track multiple objects in video?
SAM 3.1 uses Object Multiplex to group tracked objects into shared processing buckets and reason about them jointly. This is designed to improve efficiency in crowded scenes, and Meta reports an approximately sevenfold speedup at 128 tracked objects on one H100 GPU compared with the November 2025 SAM 3 release.
How much does SAM 3.1 cost through the hosted API?
The hosted Meta Model API lists pricing of $2.50 per 1,000 images and $0.20 per 1,000 video frames. Video costs can increase quickly for long or high-frame-rate footage because each processed frame contributes to usage.
What types of prompts and inputs does SAM 3.1 support?
SAM 3.1 accepts images and video and supports text prompts such as “person” or “red car.” It also supports visual prompts including points, bounding boxes, masks, and exemplar images or regions.
Is SAM 3.1 a general-purpose AI assistant or video-generation model?
No. SAM 3.1 is a specialized perception model, not a conversational assistant, coding model, or media generator. It analyzes images and video to produce detections, masks, scores, and tracking results, but it does not generate ordinary text, audio, images, or video files.


Sources 5
Provider

About Meta AI