What is SAM Audio?
SAM Audio is a Meta foundation model for prompted audio separation. Its job is to isolate a target sound from a mixture rather than to generate a written answer. A mixture might contain speech, music, instruments, environmental noise, or several of these at once. The user provides a prompt identifying the desired source, and the model returns separated audio.
This makes SAM Audio different from a conventional noise reducer or a fixed vocal-removal filter. A fixed processor generally performs one predefined operation. SAM Audio is designed to follow a description such as “dog barking,” “the person speaking,” “singing voice,” or “electric guitar,” subject to the quality and distinctiveness of the recording. It can also return a residual track containing the material left after the target is extracted.
Meta released SAM Audio as an open research project on December 16, 2025. The release includes research materials, a public demonstration, source code, and downloadable checkpoints. It is a specialized audio model, not a chat model, general-purpose language model, or hosted assistant.
How SAM Audio identifies a target sound
SAM Audio supports three principal prompting modes. Text prompting uses a natural-language description of the sound to isolate. For example, a user could request a “female speaking voice,” “drums,” or “background traffic.” This is useful when the target can be described clearly but is not easy to select directly in the recording.
Visual prompting applies when the input includes video. The user can identify a sounding person or object in the visual scene, allowing the model to associate the selected source with its audio. This can help when several people or objects are present and the visual context provides information that audio alone does not make obvious.
Span prompting uses a marked time range in which the desired sound is present. A short interval can indicate the source to extract even when a verbal description would be ambiguous. For example, a user might mark a section where one instrument plays alone before asking the model to separate that instrument from the full mixture.
The prompting modes can be combined. Text can describe the target, a visual selection can identify its source, and a temporal span can indicate when it should be active. This combination is one of the model’s main distinctions: separation is guided by user-provided evidence instead of being entirely unprompted.
Architecture and audio outputs
According to Meta’s research materials, SAM Audio uses a flow-matching diffusion Transformer operating in a DAC-VAE latent space. In practical terms, the model works in a compressed representation of audio and uses a generative denoising process to produce the separated result. These implementation details matter mainly to researchers and engineers evaluating the model locally; ordinary users interact with the model through an audio mixture and one or more prompts.
The model is supported by Perception Encoder Audio-Visual, or PE-AV, which supplies audio-visual representations for multimodal prompting. This helps the system connect information from the soundtrack, video frames, text instructions, and temporal locations.
SAM Audio produces audio rather than text, images, or video. Its typical result includes a target stem and a residual stem. The target is the requested sound, while the residual contains other material that was not assigned to the target. This two-part output is useful for dialogue cleanup, editing, and analysis because it preserves access to both the extracted source and the remainder of the recording.
Checkpoints, release status, and access
The official release provides small, base, and large checkpoints, along with corresponding -tv variants. The documented checkpoint names include sam-audio-small, sam-audio-base, sam-audio-large, and the small, base, and large visual-prompting or target-correctness variants.
The checkpoints are downloadable rather than offered through a published hosted API with per-minute or per-request pricing. Access to the Hugging Face checkpoints requires an access request and authentication. The repository recommends Python 3.11 or newer and a CUDA-compatible GPU for local use. The research materials are released under Meta’s SAM License, so teams should review that license before incorporating the model or its checkpoints into a product.
There is no verified context-window limit or maximum-token output limit because SAM Audio is not a text-generation model. The relevant practical limits are determined by the local implementation, available GPU memory, audio and video inputs, and the chosen checkpoint. The supplied research does not specify a single maximum recording duration or a universal hardware requirement beyond the repository’s recommendation of a CUDA-compatible GPU.
Speed, model sizes, and trade-offs
Meta reports a real-time factor of approximately 0.7 across model scales ranging from 500 million to 3 billion parameters in its evaluation discussion. A real-time factor below 1 can indicate faster-than-real-time processing under the tested conditions. This is a provider-reported evaluation claim, not a guarantee for every computer, recording, prompt type, or software configuration.
The checkpoint sizes create a practical quality-versus-speed trade-off. A smaller checkpoint is generally the more suitable starting point when local hardware or throughput matters, while a larger checkpoint may be preferable when separation quality is more important. The supplied research does not provide a universal ranking of every checkpoint for every task, so users should validate their chosen variant on representative recordings.
Editorial assessments classify SAM Audio as highly specialized: its usefulness and cost efficiency are strongest when the task is prompted audio separation, not when a team needs text reasoning, code generation, web search, or general-purpose multimodal assistance. Because there is no hosted API price in the supplied materials, the effective cost depends on local hardware, processing time, storage, and operational support.
Main strengths and limitations
- Flexible target descriptions: Users can identify speech, instruments, music, or environmental sounds with natural-language text instead of relying only on fixed source categories.
- Multimodal guidance: Video selections and temporal spans add information that can disambiguate a target in a crowded recording.
- Target and residual tracks: The output is useful for both extraction and subsequent editing because the remainder of the mixture is retained.
- Multiple checkpoints: Small, base, large, and specialized variants let local users make a speed and resource trade-off.
- Open research availability: Code and downloadable checkpoints support experimentation without depending on an unpublished hosted service price.
The model also has important boundaries. SAM Audio does not use audio itself as a prompting modality. An audio mixture is the material to process, but the research description does not support supplying a separate audio example as the prompt. The model is also not intended to separate every source in a mixture without guidance. It is better understood as prompted extraction than as a fully automatic inventory of all audible components.
Highly similar overlapping sources remain difficult. Examples include one singer within a chorus or one instrument inside a dense orchestral recording. Visual prompting may help when the video clearly identifies the source, but it cannot guarantee clean separation when sources overlap heavily or the visual and audio scenes do not correspond.
What is SAM Audio useful for?
SAM Audio is a strong candidate for workflows that begin with a specific sound-selection request. Dialogue cleanup is one example: a user can target a speaker or speech segment, then use the separated track for editing or accessibility work. Sound designers can extract a particular environmental effect from a field recording. Music producers and researchers can investigate vocals, instruments, or other components in a mixed recording.
Other plausible uses include audiovisual sound segmentation, assistance tools that make speech or important events more audible, dataset preparation, and experiments in multimodal audio editing. The model is especially relevant when the target is not known in advance as one of a few fixed categories and must instead be described or selected in context.
When to choose SAM Audio
Choose SAM Audio when you need locally run, prompt-controlled separation and can provide a useful text, visual, or temporal cue. It is a good fit for researchers, audio-tool builders, editors, and developers who need to isolate different types of sounds from varied recordings rather than apply one fixed vocal or noise filter. Its downloadable checkpoints are also useful when a workflow requires local processing instead of a hosted service.
A smaller checkpoint may be the more practical choice for quick experiments or constrained hardware, while a larger checkpoint may be worth testing when the target is difficult and quality matters more than processing cost. Meta’s reported approximately 0.7 real-time factor suggests that faster-than-real-time processing is possible in the reported evaluation setting, but actual performance must be measured on the intended hardware.
Another type of option may be more appropriate when the task is unprompted full-mixture separation, when the target sources are highly similar and overlapping, or when a simple fixed operation is sufficient. A general-purpose multimodal model would be more appropriate for text reasoning, coding, or tool use, since SAM Audio does not provide those functions. A hosted audio service may also be preferable for teams that do not want to manage CUDA hardware, checkpoint access, and local inference, although no official hosted pricing is supplied here for comparison.
Capability summary
| Capability | SAM Audio support |
|---|---|
| Primary purpose | Prompted separation of target and residual audio |
| Inputs | Audio mixtures, text prompts, visual cues from video, and temporal-span prompts |
| Outputs | Separated target audio and residual audio |
| Text generation | Not supported |
| Image or video generation | Not supported |
| Reasoning and coding | Not a general reasoning or coding model |
| Tool or function calling | No documented support |
| Hosted API pricing | No published pricing in the supplied research |
| Local availability | Official code and downloadable checkpoints; Hugging Face access requires authentication |
Overall, SAM Audio’s value comes from the control it gives users over sound selection. It should be evaluated as a specialized audio-separation system, with results judged on target fidelity, residual quality, processing speed, and performance on the specific mixtures that matter to the intended workflow.

