What SAM Audio Judge does
SAM Audio Judge is a specialized audio-evaluation model from Meta AI. Its job is to assess the quality of an audio-separation result rather than generate or edit audio. A typical evaluation provides three inputs: the original mixed recording, the separated audio produced by another system, and a text description of the target sound, such as “the lead vocal,” “the drums,” or “a dog barking.”
The model then returns continuous scores describing how well the separated output meets the target. This makes it useful when a clean, isolated reference track does not exist. For many real-world recordings, there is no reliable ground-truth stem against which to compare a system, so conventional reference-based metrics are difficult to apply.
SAM Audio Judge belongs to the broader SAM Audio research release but is a distinct model from SAM Audio, which performs controllable audio separation. SAM Audio Judge can evaluate outputs from SAM Audio as well as outputs from other separation systems.
What the evaluation scores mean
The official materials describe four primary evaluation dimensions:
- Overall quality: a combined assessment of the separated result.
- Recall: how much of the intended target sound was successfully captured.
- Precision: how effectively unwanted sounds were excluded from the output.
- Faithfulness: how closely the extracted sound corresponds to the target in the original mixture.
These measures separate problems that can otherwise be hidden by a single score. For example, an output may contain much of the requested speech but also include substantial background music. That result could have relatively strong recall but weaker precision. Conversely, an extremely clean output that removes most of the target might score better on isolation while performing poorly on recall.
Meta AI describes the evaluation framework as perceptually oriented. Its training and evaluation process used human ratings collected with a five-point annotation scale. The resulting model scores should still be treated as model-based estimates, not as definitive replacements for listening tests or human judgments.
How SAM Audio Judge is used
SAM Audio Judge is normally placed after an audio-separation pipeline. First, another model or system creates a candidate separated track. The evaluator then receives the original mixture, that candidate output, and the text prompt identifying the intended source.
This arrangement supports several practical workflows. Researchers can compare different separation checkpoints, prompts, or inference settings using the same evaluator. They can also assess multiple systems on recordings for which no isolated reference stem is available. Batch processing is supported by the official implementation, alongside processing individual examples.
The model is therefore closer to an automated research judge than to a consumer audio tool. It does not clean a recording, remove noise on demand, export a better stem, or decide which source to separate without an upstream separation system.
Strengths and best use cases
The main strength of SAM Audio Judge is its reference-free evaluation design. In controlled datasets, researchers can often compare an output with a known target recording. In practical recordings, however, the original performance may not have been captured as separate tracks. SAM Audio Judge provides a way to estimate quality using the mixture and a semantic description of the target instead.
- Benchmarking separation systems: compare models when clean reference stems are unavailable.
- Testing prompts: measure whether different text descriptions lead to better extraction of the requested sound.
- Comparing checkpoints: evaluate changes between model versions or inference configurations.
- Analyzing target capture and interference: distinguish missing target content from unwanted material left in the output.
- Evaluating varied audio: assess speech, music, instruments, and general sound effects within a common framework.
It is particularly appropriate when the evaluation question is not simply “does this waveform match a reference?” but “did the system extract the sound described by the user, while preserving the relevant content and reducing interference?”
Inputs, outputs, and modalities
The model accepts audio and text-related evaluation information. Its required audio inputs are the original mixture and a separated result. The text input identifies the intended target sound. The output is numerical evaluation information: overall quality, recall, precision, and faithfulness scores.
| Aspect | Verified information |
|---|---|
| Primary purpose | Reference-free evaluation of audio-separation outputs |
| Audio inputs | Original mixture and separated audio |
| Text input | Description of the target sound |
| Outputs | Overall quality, recall, precision, and faithfulness scores |
| Direct audio generation | Not supported; the model evaluates audio rather than producing a separated track |
| Batch processing | Supported by the official implementation |
No authoritative context-window limit, maximum output-token limit, hosted inference endpoint, or official API pricing is provided in the supplied materials. Because this is an evaluator whose output consists of scores rather than a conversational response, language-model concepts such as a general-purpose context window or token-generation ceiling are not the central operating limits documented for this release.
Access, pricing, and release status
The official checkpoint is available through Meta AI's gated Hugging Face repository at facebook/sam-audio-judge. Access approval and authentication are required. The model card identifies it as an approximately 2-billion-parameter checkpoint released under the SAM License.
SAM Audio Judge is a current gated research release in the supplied information, with a listed release date of December 16, 2025. No official hosted API pricing was found. This means users should not assume that a standard per-request commercial endpoint is available. Running the model requires a compatible research software environment, including Python and PyTorch components described by the official implementation.
The absence of published hosted pricing is itself an important practical limitation. Organizations must account for access approval, local or managed computing resources, storage, and engineering work rather than evaluating the model solely through an advertised API rate.
Limitations and trade-offs
SAM Audio Judge is narrow by design. It does not generate separated audio, perform speech synthesis, transcribe recordings, provide general conversation, or operate as a real-time audio processor. A user seeking an immediately cleaned vocal, isolated instrument, or edited sound file needs a separation or audio-editing system instead.
Its scores also require careful interpretation. A model-based evaluator may not agree with every human listener, particularly on ambiguous mixtures, unusual sound sources, or cases where the text description is underspecified. The scores are best used consistently across experiments rather than treated as absolute proof that one output is universally better.
The checkpoint is also gated and research-oriented. That creates more friction than an openly available hosted service or a lightweight local utility. The approximately 2-billion-parameter size may also make it less suitable for low-resource or latency-sensitive deployments, although the supplied materials do not publish formal hardware, throughput, or latency requirements.
Editorial assessments supplied with the model data rate its speed as relatively favorable and its cost as moderate on a comparative internal scale, but these are editorial scores, not provider-published benchmarks or prices. They should not be read as guaranteed runtime or operating-cost figures.
When to choose SAM Audio Judge
Choose SAM Audio Judge when the primary need is systematic evaluation of text-guided audio separation and clean reference tracks are unavailable. It is a strong fit for research teams comparing separation systems, prompts, or checkpoints and needing more informative diagnostics than a single undifferentiated quality score.
Another separation model or a conventional audio tool may be more appropriate when the goal is to create the separated track itself. A reference-based metric may also be preferable when high-quality isolated stems are available and reproducible signal-level comparison is more important than reference-free perceptual assessment. For production applications that require a simple hosted endpoint, predictable pricing, or real-time processing, the gated research checkpoint may be less practical than a commercial audio service or a smaller task-specific system.
In short, SAM Audio Judge fills a focused position in the SAM Audio lineup: it is the evaluation component for judging separation quality, not the separation engine. Its value comes from making comparative assessment possible when conventional reference recordings are missing, while its limitations come from its research access model, specialized inputs, and reliance on model-estimated rather than definitive human judgments.

