Perception Encoder

Perception Encoder Audiovisual

by Meta AI · Current and openly available

Meta’s Perception Encoder Audiovisual family produces aligned embeddings for audio, video, audio-video, and text. Its six open checkpoints support cross-modal retrieval, sound-event discovery, media indexing, and audiovisual systems such as SAM Audio, but it does not generate text or media itself.

Embeddings Reasoning Coding
Perception Encoder Audiovisual (PE-AV) is a Meta-developed model family for comparing audio, video, audio-video content, and text in a shared embedding space. Rather than generating prose or media, it produces vector representations that downstream systems can use for retrieval, ranking, classification, audiovisual search, and sound understanding. Six principal checkpoints are available across small, base, and large sizes, with standard and 16-frame video variants.
Outputs

What Perception Encoder Audiovisual can produce

Embeddings
Inputs

What it can understand

Text Audio Video Multimodal input
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Perception Encoder
Model type Multimodal
Context window tokens
Maximum output tokens
Release date 2025-12-16
Status Current and openly available
Knowledge cutoff notes

The official PE-AV materials describe training data and release information but do not state a conventional textual knowledge cutoff. This is an embedding encoder rather than a knowledge-based language model.

Model notes

Perception Encoder Audiovisual is a family rather than one parameter configuration. Meta released six principal checkpoints: pe-av-small, pe-av-base, pe-av-large, and corresponding -16-frame variants. The standard variants support variable-length video input, while -16-frame variants sample exactly 16 evenly spaced frames. The model encodes audio, video, audio-video, and text into aligned embedding spaces and is used as a perception component in SAM Audio. PE-A-Frame is a related but distinct frame-level model for fine-grained audio-frame alignment. Official model cards state that the checkpoints are released under Apache 2.0. Editorial scores reflect an embedding encoder and should not be compared directly with generative language models.

Cost

Model pricing

Input No official hosted API pricing; open checkpoints are available for self-managed use.
Output No official hosted API pricing; the model returns embeddings rather than generated media.
Model guide

Perception Encoder Audiovisual (PE-AV): Meta’s Open Encoder for Audio-Video-Text Retrieval

Perception Encoder Audiovisual (PE-AV) is Meta’s open family of audiovisual embedding encoders. It maps audio, video, combined audio-video inputs, and text into aligned vector spaces for cross-modal retrieval, sound-event understanding, media indexing, and multimodal perception systems.

What is Perception Encoder Audiovisual?

Perception Encoder Audiovisual, usually abbreviated PE-AV, is an open multimodal encoder from Meta. Its primary job is to transform different kinds of content into embeddings: numerical vector representations that capture useful relationships between inputs. Audio, video, audio-video clips, and text can therefore be compared even when they do not share the same original format.

For example, a text description could be matched against an audio recording, a video could be retrieved using a written query, or a sound could be associated with the visual content that produced it. This makes PE-AV a perception and representation model, not a conversational language model or a generative media model.

Meta describes PE-AV as part of its Perception Encoder research. The official release focuses on audiovisual correspondence learning, using contrastive learning to bring related audio, visual, and textual representations closer together in embedding space while separating less-related examples.

Where PE-AV fits in Meta’s model catalog

PE-AV occupies a specialized position in Meta’s publicly released Perception Models work. It is intended to supply multimodal representations to other systems rather than act as a complete end-user application. Meta also uses PE-AV as a component in SAM Audio, where audiovisual information helps connect visible people or objects with corresponding sounds. SAM Audio performs the generative audio-separation task; PE-AV supplies perception features that help the larger system understand relationships between modalities.

The release also includes PE-A-Frame, a related frame-level model designed for fine-grained alignment between audio events, video frames, and text. PE-A-Frame is not the same model as the main PE-AV family. It is relevant when the goal is to identify which particular video frame corresponds to an audio event, while PE-AV is the broader audiovisual encoder described on this page.

Supported inputs and outputs

PE-AV accepts four broad input types:

  • Audio: speech, music, and general sound effects can be represented as audio embeddings.
  • Video: video frames can be encoded into visual representations.
  • Audio-video: temporally related sound and visual information can be processed together.
  • Text: written descriptions and queries can be mapped into the same general representation space.

The resulting embeddings support audio-only, video-only, audio-video, audio-text, video-text, and audiovisual-text comparisons. In practical terms, a media index could use PE-AV to search for clips containing a described sound, locate videos related to a spoken query, or rank media according to how well its audio and visuals match a caption.

The model’s direct output is a vector embedding. It does not natively return a written answer, image, audio file, video, speech waveform, or music track. Any labels, captions, search results, or generated media must be produced by a separate downstream component.

Checkpoints and video-processing choices

Meta released six principal PE-AV checkpoints. They are organized into three approximate parameter sizes, each with a standard and fixed-frame version:

CheckpointVideo behaviorApproximate size
pe-av-smallVariable-length or all-frame processing0.8B parameters
pe-av-baseVariable-length or all-frame processingApproximately 1B parameters
pe-av-largeVariable-length or all-frame processingApproximately 2B parameters
pe-av-small-16-frameExactly 16 evenly sampled frames0.8B parameters
pe-av-base-16-frameExactly 16 evenly sampled framesApproximately 1B parameters
pe-av-large-16-frameExactly 16 evenly sampled framesApproximately 2B parameters

The standard variants are intended for variable-length or broader temporal coverage, while the 16-frame variants use a fixed sampling strategy. The fixed-frame versions can make resource planning and throughput easier, but they may provide less temporal detail when an important event occurs between sampled frames. The all-frame variants can preserve more of a clip’s temporal information, at the cost of greater processing and memory requirements.

The parameter counts and checkpoint behavior above are specifications described in Meta’s model materials. They should not be interpreted as a guarantee of identical speed across hardware, input lengths, or implementation settings.

Training and reported performance

Meta trained PE-AV with large-scale multimodal contrastive learning. The research describes an audiovisual data engine built around more than 100 million audio-video pairs, covering speech, music, and general sound effects. This broad training scope is important because a system trained only on speech, for example, would not necessarily be suitable for general media search or environmental sound understanding.

Meta reports strong results across audio, video, and audiovisual retrieval benchmarks. Those are provider-reported research results rather than a guarantee for every application. Real-world performance can vary with recording quality, domain mismatch, language, visual content, temporal sampling, and the quality of the downstream retrieval or classification system.

PE-AV’s contrastive design makes it especially useful when the desired operation is comparison. It is less appropriate when the requirement is to explain a clip in fluent language, follow a multi-step conversation, or generate a new media asset.

What can PE-AV be used for?

PE-AV is a good fit for applications that need to measure relationships among sound, imagery, and text. Common examples include:

  • Cross-modal search: retrieve audio, video, or audiovisual clips using a text query or another media example.
  • Media indexing: create searchable representations for large audio and video collections.
  • Sound-event discovery: find clips associated with events such as speech, music, or general sound effects.
  • Audiovisual ranking: rank candidate clips by how well their sound, imagery, and captions correspond.
  • Speech and audio retrieval: locate recordings related to text descriptions or other audio examples.
  • Downstream perception systems: provide embeddings to classifiers, evaluators, captioning systems, or multimodal pipelines.

In a practical search system, an application could encode its media library once, store the resulting vectors, encode a user’s text or audio query, and then retrieve nearby vectors. PE-AV does not itself define the final user interface or ranking policy; those parts must be implemented around the encoder.

Capabilities, limits, and deployment considerations

Because PE-AV is an embedding encoder, it has no conventional generated-text context window or maximum output-token limit. The supplied official materials do not specify a context length or token-generation budget for the model. Its output is an embedding rather than a variable-length response.

The model does not provide native reasoning in the sense of a deliberative chat model, and it does not offer built-in coding assistance, tool calling, function calling, web search, JSON generation, or streaming text responses. These are not missing configuration switches in a hosted conversational API; they are outside the model’s primary role.

PE-AV also does not provide direct image, audio, video, music, or speech generation. If an application needs those outputs, it must combine the encoder with other models. Similarly, PE-AV should not be selected as a standalone replacement for a language model when the task requires dialogue, summarization, instruction following, or natural-language explanations.

Meta’s official release materials describe open checkpoints and implementation rather than a hosted token-based inference service or consumer subscription interface. There is no official hosted API price supplied for PE-AV. Deployment therefore requires managing the released model code, checkpoints, hardware, and surrounding inference pipeline. The largest approximately 2B-parameter variants are likely to demand more resources than the smaller checkpoints, particularly when processing video with broad temporal coverage; exact runtime depends on the deployment environment and is not specified in the supplied research.

The official model card states that the checkpoints are released under the Apache 2.0 license. Users should still review the applicable repository and model-card terms before deploying the model in a particular commercial or regulated workflow.

When to choose PE-AV

Choose PE-AV when the central problem is aligning or retrieving information across audio, video, and text. It is particularly suitable for a self-managed media-search system, an audiovisual ranking pipeline, a sound-event discovery tool, or a larger perception system that needs reusable embeddings rather than generated responses.

The family also gives deployment flexibility. A small or 16-frame checkpoint may be preferable when throughput, predictable video sampling, and lower resource requirements matter more than maximum temporal coverage. A larger or all-frame variant may be more appropriate when the application depends on richer audiovisual detail and can support the additional computation.

Choose another type of model when the main requirement is generation or interaction. A conversational language model is more appropriate for dialogue and text generation; a dedicated speech or audio model may be better for transcription or audio synthesis; and a generative separation model such as the larger SAM Audio system is more appropriate when the goal is to isolate or create an audio output. PE-AV can support those systems as a perception component, but it is not itself the complete application.

Bottom line

Perception Encoder Audiovisual is a specialized, open Meta model family for turning audio, video, audio-video content, and text into comparable embeddings. Its value is in cross-modal perception: finding relationships among media types, supporting retrieval, and supplying representations to downstream audiovisual systems. Its main trade-off is equally clear: it offers no hosted generation interface and does not replace a conversational or generative model. For organizations able to manage open checkpoints and build the surrounding pipeline, PE-AV provides a focused foundation for audiovisual search and understanding.


Answers to Frequently Asked Questions

How does PE-AV differ from PE-A-Frame and SAM Audio?
PE-AV is a broad audiovisual encoder for creating comparable audio, video, and text embeddings. PE-A-Frame is designed for fine-grained alignment between specific audio events and video frames. SAM Audio performs generative audio separation, while PE-AV supplies perception features that can help SAM Audio connect sounds with visible sources.
Which PE-AV checkpoints are available?
Meta released six principal checkpoints: pe-av-small, pe-av-base, pe-av-large, pe-av-small-16-frame, pe-av-base-16-frame, and pe-av-large-16-frame. They range from approximately 0.8B to 2B parameters. The 16-frame variants process exactly 16 evenly sampled video frames, while the standard variants support variable-length or broader all-frame processing.
Does PE-AV generate text, audio, or video?
No. PE-AV produces vector embeddings rather than written answers or generated media. Applications that need captions, explanations, audio separation, speech synthesis, or video generation must combine PE-AV with separate downstream models.
What is Perception Encoder Audiovisual (PE-AV)?
Perception Encoder Audiovisual (PE-AV) is an open multimodal encoder from Meta that converts audio, video, audio-video clips, and text into numerical embeddings. These embeddings allow different media types to be compared for cross-modal retrieval and audiovisual understanding.
What can PE-AV be used for?
PE-AV can support cross-modal search, media indexing, sound-event discovery, audiovisual ranking, audio and speech retrieval, and downstream perception systems. For example, it can help retrieve videos using text queries or find clips that match a sound or audiovisual example.


Sources 6
Provider

About Meta AI