SAM Audio

SAM Audio

by Meta AI · Current open research release; downloadable checkpoints and public demo available

Meta’s SAM Audio is a specialized foundation model for separating requested sounds from complex audio or audiovisual mixtures. Text descriptions, video selections, and marked time spans can guide extraction, with target and residual audio tracks returned as output. The open research release includes multiple local checkpoints, but no hosted API pricing is published. It is best suited to prompted audio editing and research, not general text generation or fully unprompted source separation.

Audio Reasoning Coding
SAM Audio extends Meta’s Segment Anything concept into audio separation. Instead of applying only a fixed vocal, speech, or music filter, it lets a user identify the sound to extract with a text description, a selected visual source in video, or a marked time span. The released family includes small, base, large, and target-correctness or visual-prompting variants, with downloadable checkpoints available through Meta’s research release.
Outputs

What SAM Audio can produce

Audio
Inputs

What it can understand

Text Audio Video Multimodal input
Capabilities

Supported features

Multimodal output
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family SAM Audio
Model type Multimodal
Context window tokens
Maximum output tokens
Release date 2025-12-16
Status Current open research release; downloadable checkpoints and public demo available
Knowledge cutoff notes

Meta does not publish a conventional textual knowledge cutoff for SAM Audio. It is an audio-separation model rather than a general knowledge model.

Model notes

SAM Audio is a family of Meta audio-separation checkpoints rather than a conventional chat or text-generation model. Officially released variants include sam-audio-small, sam-audio-base, sam-audio-large, and small/base/large -tv variants. The model accepts audio mixtures and can use text, visual, or temporal-span prompts; audio itself is not supported as a prompt. It generates target and residual audio stems. The repository recommends Python 3.11 or newer and a CUDA-compatible GPU. Hugging Face checkpoint access requires an access request and authentication. Meta’s research materials describe a flow-matching diffusion Transformer in a DAC-VAE latent space and report approximately 0.7 real-time factor in evaluation. Editorial scores reflect its specialization in audio separation and are not vendor benchmarks.

Cost

Model pricing

Input No hosted API pricing published; downloadable checkpoints available under the SAM License
Output No hosted API pricing published; model outputs separated target and residual audio
Model guide

SAM Audio: Meta’s Prompted Model for Separating Sounds from Complex Recordings

SAM Audio is Meta’s multimodal foundation model for isolating a requested sound from mixed audio or audiovisual recordings. It accepts text descriptions, visual cues, and marked time spans as prompts, then produces target and residual audio tracks. The open research release is aimed at speech isolation, noise and sound-effect extraction, music and instrument separation, and multimodal audio-editing research rather than general-purpose text generation or unprompted source separation.

What is SAM Audio?

SAM Audio is a Meta foundation model for prompted audio separation. Its job is to isolate a target sound from a mixture rather than to generate a written answer. A mixture might contain speech, music, instruments, environmental noise, or several of these at once. The user provides a prompt identifying the desired source, and the model returns separated audio.

This makes SAM Audio different from a conventional noise reducer or a fixed vocal-removal filter. A fixed processor generally performs one predefined operation. SAM Audio is designed to follow a description such as “dog barking,” “the person speaking,” “singing voice,” or “electric guitar,” subject to the quality and distinctiveness of the recording. It can also return a residual track containing the material left after the target is extracted.

Meta released SAM Audio as an open research project on December 16, 2025. The release includes research materials, a public demonstration, source code, and downloadable checkpoints. It is a specialized audio model, not a chat model, general-purpose language model, or hosted assistant.

How SAM Audio identifies a target sound

SAM Audio supports three principal prompting modes. Text prompting uses a natural-language description of the sound to isolate. For example, a user could request a “female speaking voice,” “drums,” or “background traffic.” This is useful when the target can be described clearly but is not easy to select directly in the recording.

Visual prompting applies when the input includes video. The user can identify a sounding person or object in the visual scene, allowing the model to associate the selected source with its audio. This can help when several people or objects are present and the visual context provides information that audio alone does not make obvious.

Span prompting uses a marked time range in which the desired sound is present. A short interval can indicate the source to extract even when a verbal description would be ambiguous. For example, a user might mark a section where one instrument plays alone before asking the model to separate that instrument from the full mixture.

The prompting modes can be combined. Text can describe the target, a visual selection can identify its source, and a temporal span can indicate when it should be active. This combination is one of the model’s main distinctions: separation is guided by user-provided evidence instead of being entirely unprompted.

Architecture and audio outputs

According to Meta’s research materials, SAM Audio uses a flow-matching diffusion Transformer operating in a DAC-VAE latent space. In practical terms, the model works in a compressed representation of audio and uses a generative denoising process to produce the separated result. These implementation details matter mainly to researchers and engineers evaluating the model locally; ordinary users interact with the model through an audio mixture and one or more prompts.

The model is supported by Perception Encoder Audio-Visual, or PE-AV, which supplies audio-visual representations for multimodal prompting. This helps the system connect information from the soundtrack, video frames, text instructions, and temporal locations.

SAM Audio produces audio rather than text, images, or video. Its typical result includes a target stem and a residual stem. The target is the requested sound, while the residual contains other material that was not assigned to the target. This two-part output is useful for dialogue cleanup, editing, and analysis because it preserves access to both the extracted source and the remainder of the recording.

Checkpoints, release status, and access

The official release provides small, base, and large checkpoints, along with corresponding -tv variants. The documented checkpoint names include sam-audio-small, sam-audio-base, sam-audio-large, and the small, base, and large visual-prompting or target-correctness variants.

The checkpoints are downloadable rather than offered through a published hosted API with per-minute or per-request pricing. Access to the Hugging Face checkpoints requires an access request and authentication. The repository recommends Python 3.11 or newer and a CUDA-compatible GPU for local use. The research materials are released under Meta’s SAM License, so teams should review that license before incorporating the model or its checkpoints into a product.

There is no verified context-window limit or maximum-token output limit because SAM Audio is not a text-generation model. The relevant practical limits are determined by the local implementation, available GPU memory, audio and video inputs, and the chosen checkpoint. The supplied research does not specify a single maximum recording duration or a universal hardware requirement beyond the repository’s recommendation of a CUDA-compatible GPU.

Speed, model sizes, and trade-offs

Meta reports a real-time factor of approximately 0.7 across model scales ranging from 500 million to 3 billion parameters in its evaluation discussion. A real-time factor below 1 can indicate faster-than-real-time processing under the tested conditions. This is a provider-reported evaluation claim, not a guarantee for every computer, recording, prompt type, or software configuration.

The checkpoint sizes create a practical quality-versus-speed trade-off. A smaller checkpoint is generally the more suitable starting point when local hardware or throughput matters, while a larger checkpoint may be preferable when separation quality is more important. The supplied research does not provide a universal ranking of every checkpoint for every task, so users should validate their chosen variant on representative recordings.

Editorial assessments classify SAM Audio as highly specialized: its usefulness and cost efficiency are strongest when the task is prompted audio separation, not when a team needs text reasoning, code generation, web search, or general-purpose multimodal assistance. Because there is no hosted API price in the supplied materials, the effective cost depends on local hardware, processing time, storage, and operational support.

Main strengths and limitations

  • Flexible target descriptions: Users can identify speech, instruments, music, or environmental sounds with natural-language text instead of relying only on fixed source categories.
  • Multimodal guidance: Video selections and temporal spans add information that can disambiguate a target in a crowded recording.
  • Target and residual tracks: The output is useful for both extraction and subsequent editing because the remainder of the mixture is retained.
  • Multiple checkpoints: Small, base, large, and specialized variants let local users make a speed and resource trade-off.
  • Open research availability: Code and downloadable checkpoints support experimentation without depending on an unpublished hosted service price.

The model also has important boundaries. SAM Audio does not use audio itself as a prompting modality. An audio mixture is the material to process, but the research description does not support supplying a separate audio example as the prompt. The model is also not intended to separate every source in a mixture without guidance. It is better understood as prompted extraction than as a fully automatic inventory of all audible components.

Highly similar overlapping sources remain difficult. Examples include one singer within a chorus or one instrument inside a dense orchestral recording. Visual prompting may help when the video clearly identifies the source, but it cannot guarantee clean separation when sources overlap heavily or the visual and audio scenes do not correspond.

What is SAM Audio useful for?

SAM Audio is a strong candidate for workflows that begin with a specific sound-selection request. Dialogue cleanup is one example: a user can target a speaker or speech segment, then use the separated track for editing or accessibility work. Sound designers can extract a particular environmental effect from a field recording. Music producers and researchers can investigate vocals, instruments, or other components in a mixed recording.

Other plausible uses include audiovisual sound segmentation, assistance tools that make speech or important events more audible, dataset preparation, and experiments in multimodal audio editing. The model is especially relevant when the target is not known in advance as one of a few fixed categories and must instead be described or selected in context.

When to choose SAM Audio

Choose SAM Audio when you need locally run, prompt-controlled separation and can provide a useful text, visual, or temporal cue. It is a good fit for researchers, audio-tool builders, editors, and developers who need to isolate different types of sounds from varied recordings rather than apply one fixed vocal or noise filter. Its downloadable checkpoints are also useful when a workflow requires local processing instead of a hosted service.

A smaller checkpoint may be the more practical choice for quick experiments or constrained hardware, while a larger checkpoint may be worth testing when the target is difficult and quality matters more than processing cost. Meta’s reported approximately 0.7 real-time factor suggests that faster-than-real-time processing is possible in the reported evaluation setting, but actual performance must be measured on the intended hardware.

Another type of option may be more appropriate when the task is unprompted full-mixture separation, when the target sources are highly similar and overlapping, or when a simple fixed operation is sufficient. A general-purpose multimodal model would be more appropriate for text reasoning, coding, or tool use, since SAM Audio does not provide those functions. A hosted audio service may also be preferable for teams that do not want to manage CUDA hardware, checkpoint access, and local inference, although no official hosted pricing is supplied here for comparison.

Capability summary

CapabilitySAM Audio support
Primary purposePrompted separation of target and residual audio
InputsAudio mixtures, text prompts, visual cues from video, and temporal-span prompts
OutputsSeparated target audio and residual audio
Text generationNot supported
Image or video generationNot supported
Reasoning and codingNot a general reasoning or coding model
Tool or function callingNo documented support
Hosted API pricingNo published pricing in the supplied research
Local availabilityOfficial code and downloadable checkpoints; Hugging Face access requires authentication

Overall, SAM Audio’s value comes from the control it gives users over sound selection. It should be evaluated as a specialized audio-separation system, with results judged on target fidelity, residual quality, processing speed, and performance on the specific mixtures that matter to the intended workflow.


Answers to Frequently Asked Questions

What are SAM Audio’s main limitations?
SAM Audio is a specialized prompted separation model, not a chat, coding, or text-generation system. It does not use a separate audio example as a prompt, is not intended to automatically separate every source without guidance, and may struggle when highly similar sounds overlap heavily, such as one singer in a chorus or one instrument in a dense orchestral recording.
Is SAM Audio available for local use, and what hardware does it require?
Yes. Meta released official source code and downloadable checkpoints, including small, base, and large variants. Hugging Face access requires authentication and an access request, while the repository recommends Python 3.11 or newer and a CUDA-compatible GPU for local inference.
What audio outputs does SAM Audio produce?
SAM Audio typically produces two outputs: a target stem containing the requested sound and a residual stem containing the remaining material that was not assigned to the target.
What is SAM Audio designed to do?
SAM Audio is Meta’s foundation model for prompted audio separation. It isolates a target sound from a complex mixture containing speech, music, instruments, environmental noise, or other sources, and returns a target audio stem plus a residual track.
How can users prompt SAM Audio to separate a sound?
SAM Audio supports text prompts, visual prompts from video, and temporal-span prompts. Users can describe a sound such as “dog barking” or “electric guitar,” select a sounding person or object in video, mark a time range where the target is present, or combine these prompting methods.


Sources 5
Provider

About Meta AI