What is Perception Encoder Audiovisual?
Perception Encoder Audiovisual, usually abbreviated PE-AV, is an open multimodal encoder from Meta. Its primary job is to transform different kinds of content into embeddings: numerical vector representations that capture useful relationships between inputs. Audio, video, audio-video clips, and text can therefore be compared even when they do not share the same original format.
For example, a text description could be matched against an audio recording, a video could be retrieved using a written query, or a sound could be associated with the visual content that produced it. This makes PE-AV a perception and representation model, not a conversational language model or a generative media model.
Meta describes PE-AV as part of its Perception Encoder research. The official release focuses on audiovisual correspondence learning, using contrastive learning to bring related audio, visual, and textual representations closer together in embedding space while separating less-related examples.
Where PE-AV fits in Meta’s model catalog
PE-AV occupies a specialized position in Meta’s publicly released Perception Models work. It is intended to supply multimodal representations to other systems rather than act as a complete end-user application. Meta also uses PE-AV as a component in SAM Audio, where audiovisual information helps connect visible people or objects with corresponding sounds. SAM Audio performs the generative audio-separation task; PE-AV supplies perception features that help the larger system understand relationships between modalities.
The release also includes PE-A-Frame, a related frame-level model designed for fine-grained alignment between audio events, video frames, and text. PE-A-Frame is not the same model as the main PE-AV family. It is relevant when the goal is to identify which particular video frame corresponds to an audio event, while PE-AV is the broader audiovisual encoder described on this page.
Supported inputs and outputs
PE-AV accepts four broad input types:
- Audio: speech, music, and general sound effects can be represented as audio embeddings.
- Video: video frames can be encoded into visual representations.
- Audio-video: temporally related sound and visual information can be processed together.
- Text: written descriptions and queries can be mapped into the same general representation space.
The resulting embeddings support audio-only, video-only, audio-video, audio-text, video-text, and audiovisual-text comparisons. In practical terms, a media index could use PE-AV to search for clips containing a described sound, locate videos related to a spoken query, or rank media according to how well its audio and visuals match a caption.
The model’s direct output is a vector embedding. It does not natively return a written answer, image, audio file, video, speech waveform, or music track. Any labels, captions, search results, or generated media must be produced by a separate downstream component.
Checkpoints and video-processing choices
Meta released six principal PE-AV checkpoints. They are organized into three approximate parameter sizes, each with a standard and fixed-frame version:
| Checkpoint | Video behavior | Approximate size |
|---|---|---|
| pe-av-small | Variable-length or all-frame processing | 0.8B parameters |
| pe-av-base | Variable-length or all-frame processing | Approximately 1B parameters |
| pe-av-large | Variable-length or all-frame processing | Approximately 2B parameters |
| pe-av-small-16-frame | Exactly 16 evenly sampled frames | 0.8B parameters |
| pe-av-base-16-frame | Exactly 16 evenly sampled frames | Approximately 1B parameters |
| pe-av-large-16-frame | Exactly 16 evenly sampled frames | Approximately 2B parameters |
The standard variants are intended for variable-length or broader temporal coverage, while the 16-frame variants use a fixed sampling strategy. The fixed-frame versions can make resource planning and throughput easier, but they may provide less temporal detail when an important event occurs between sampled frames. The all-frame variants can preserve more of a clip’s temporal information, at the cost of greater processing and memory requirements.
The parameter counts and checkpoint behavior above are specifications described in Meta’s model materials. They should not be interpreted as a guarantee of identical speed across hardware, input lengths, or implementation settings.
Training and reported performance
Meta trained PE-AV with large-scale multimodal contrastive learning. The research describes an audiovisual data engine built around more than 100 million audio-video pairs, covering speech, music, and general sound effects. This broad training scope is important because a system trained only on speech, for example, would not necessarily be suitable for general media search or environmental sound understanding.
Meta reports strong results across audio, video, and audiovisual retrieval benchmarks. Those are provider-reported research results rather than a guarantee for every application. Real-world performance can vary with recording quality, domain mismatch, language, visual content, temporal sampling, and the quality of the downstream retrieval or classification system.
PE-AV’s contrastive design makes it especially useful when the desired operation is comparison. It is less appropriate when the requirement is to explain a clip in fluent language, follow a multi-step conversation, or generate a new media asset.
What can PE-AV be used for?
PE-AV is a good fit for applications that need to measure relationships among sound, imagery, and text. Common examples include:
- Cross-modal search: retrieve audio, video, or audiovisual clips using a text query or another media example.
- Media indexing: create searchable representations for large audio and video collections.
- Sound-event discovery: find clips associated with events such as speech, music, or general sound effects.
- Audiovisual ranking: rank candidate clips by how well their sound, imagery, and captions correspond.
- Speech and audio retrieval: locate recordings related to text descriptions or other audio examples.
- Downstream perception systems: provide embeddings to classifiers, evaluators, captioning systems, or multimodal pipelines.
In a practical search system, an application could encode its media library once, store the resulting vectors, encode a user’s text or audio query, and then retrieve nearby vectors. PE-AV does not itself define the final user interface or ranking policy; those parts must be implemented around the encoder.
Capabilities, limits, and deployment considerations
Because PE-AV is an embedding encoder, it has no conventional generated-text context window or maximum output-token limit. The supplied official materials do not specify a context length or token-generation budget for the model. Its output is an embedding rather than a variable-length response.
The model does not provide native reasoning in the sense of a deliberative chat model, and it does not offer built-in coding assistance, tool calling, function calling, web search, JSON generation, or streaming text responses. These are not missing configuration switches in a hosted conversational API; they are outside the model’s primary role.
PE-AV also does not provide direct image, audio, video, music, or speech generation. If an application needs those outputs, it must combine the encoder with other models. Similarly, PE-AV should not be selected as a standalone replacement for a language model when the task requires dialogue, summarization, instruction following, or natural-language explanations.
Meta’s official release materials describe open checkpoints and implementation rather than a hosted token-based inference service or consumer subscription interface. There is no official hosted API price supplied for PE-AV. Deployment therefore requires managing the released model code, checkpoints, hardware, and surrounding inference pipeline. The largest approximately 2B-parameter variants are likely to demand more resources than the smaller checkpoints, particularly when processing video with broad temporal coverage; exact runtime depends on the deployment environment and is not specified in the supplied research.
The official model card states that the checkpoints are released under the Apache 2.0 license. Users should still review the applicable repository and model-card terms before deploying the model in a particular commercial or regulated workflow.
When to choose PE-AV
Choose PE-AV when the central problem is aligning or retrieving information across audio, video, and text. It is particularly suitable for a self-managed media-search system, an audiovisual ranking pipeline, a sound-event discovery tool, or a larger perception system that needs reusable embeddings rather than generated responses.
The family also gives deployment flexibility. A small or 16-frame checkpoint may be preferable when throughput, predictable video sampling, and lower resource requirements matter more than maximum temporal coverage. A larger or all-frame variant may be more appropriate when the application depends on richer audiovisual detail and can support the additional computation.
Choose another type of model when the main requirement is generation or interaction. A conversational language model is more appropriate for dialogue and text generation; a dedicated speech or audio model may be better for transcription or audio synthesis; and a generative separation model such as the larger SAM Audio system is more appropriate when the goal is to isolate or create an audio output. PE-AV can support those systems as a perception component, but it is not itself the complete application.
Bottom line
Perception Encoder Audiovisual is a specialized, open Meta model family for turning audio, video, audio-video content, and text into comparable embeddings. Its value is in cross-modal perception: finding relationships among media types, supporting retrieval, and supplying representations to downstream audiovisual systems. Its main trade-off is equally clear: it offers no hosted generation interface and does not replace a conversational or generative model. For organizations able to manage open checkpoints and build the surrounding pipeline, PE-AV provides a focused foundation for audiovisual search and understanding.

