TRIBE

TRIBE v2

by Meta AI · Open-weight research release

Meta’s TRIBE v2 predicts human brain responses to video, audio, and language using a unified multimodal architecture trained on more than 1,000 hours of fMRI data from 720 subjects. The open-weight research model supports zero-shot prediction, subject-specific fine-tuning, and in-silico neuroscience experiments.

Reasoning Coding
TRIBE v2 is a multimodal brain-encoding model from Meta that maps video, audio, and text stimuli to predicted human fMRI responses. It was designed as a computational research tool for studying perception, language, multisensory integration, and brain organization through in-silico experiments rather than generating text, images, or audio.
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Fine-tuning
Model profile

Performance characteristics

0/10 Reasoning
0/10 Coding
0/10 Speed
0/10 Cost efficiency
Specifications

Technical details

Model family TRIBE
Model type Multimodal
Context window tokens
Maximum output tokens
Release date 2026-03-26
Status Open-weight research release
Knowledge cutoff notes

A conventional textual knowledge cutoff is not documented for this brain-response prediction model. Its behavior is determined by the pretrained multimodal encoders, fMRI training data, and released model weights.

Model notes

TRIBE v2 is a deep multimodal brain-encoding model rather than a conventional generative language model. It predicts fMRI responses to video, audio, and text and returns numerical brain-activity predictions mapped to cortical space. The public release includes code, model weights, a research paper, and an interactive demo. Meta states that the release is under a CC BY-NC license. The model can generalize zero-shot to new subjects and tasks and can be fine-tuned with limited subject-specific data. Conventional token pricing, context-window, maximum-output-token, and hosted API specifications are not applicable.

Model guide

TRIBE v2: Meta’s Open-Weight Model for Predicting Human Brain Responses

TRIBE v2 is an open-weight Meta research model that predicts human brain responses to video, audio, and language stimuli using high-resolution fMRI data. It enables computational neuroscience experiments without collecting new scans for every hypothesis.

What is TRIBE v2?

TRIBE v2 is an open-weight research model from Meta for in-silico neuroscience. In practical terms, it estimates how human brain activity may respond to naturalistic stimuli such as videos, podcasts, speech, images, and written language. Instead of producing a conversation, an image, or an audio clip, the model returns numerical predictions of fMRI brain activity mapped to cortical space.

The model is designed for researchers who want to study how the brain processes vision, audition, language, and combinations of these signals. It can be used to test hypotheses or screen potential experiments computationally before collecting additional human brain scans.

TRIBE v2 is therefore best understood as a brain-encoding model, not as a conventional generative AI assistant. Its output describes predicted neural responses; it does not represent a person’s private thoughts or decode arbitrary mental content.

Where TRIBE v2 fits in Meta’s model lineup

TRIBE v2 occupies a specialized research position in Meta’s catalog. It uses multimodal foundation-model techniques, but its target is human brain activity rather than text, images, speech, or video generation. The supplied release includes model weights, implementation code, a research paper, and an interactive demo.

That positioning makes TRIBE v2 relevant to computational neuroscience and brain-imaging research rather than ordinary chatbot, media-generation, search, or production API workloads. The release is described as open-weight and intended for research use under a CC BY-NC license.

How TRIBE v2 works

TRIBE v2 combines pretrained representations for video, audio, and language with a unified Transformer architecture. A representation is a numerical description of input content that a machine-learning system can process. The model uses these multimodal representations to estimate corresponding activity across the human cortex.

According to the supplied research, training used more than 1,000 hours of fMRI recordings from 720 subjects across multiple neuroscience datasets. The material included audiovisual media, podcasts, spoken language, videos, images, and text-related experimental conditions.

At inference time, researchers can provide video, audio, text, or combinations of these modalities. The released implementation predicts responses for an average subject on the fsaverage5 cortical mesh, which covers approximately 20,000 cortical vertices. The predictions are shifted to account for the hemodynamic delay: the measured fMRI response generally follows the original stimulus rather than occurring at exactly the same moment.

Inputs, outputs, and supported modalities

AreaTRIBE v2 support
Text inputSupported
Audio inputSupported
Video inputSupported
Multimodal inputSupported across visual, auditory, and linguistic signals
Text outputNot provided
Image, audio, or video generationNot provided
Primary outputNumerical predictions of fMRI brain responses mapped to cortical space

The model’s multimodal capability refers to its ability to process different kinds of stimulus data. It should not be confused with multimodal conversational output: TRIBE v2 does not answer questions in natural language or generate media.

What TRIBE v2 can do

  • Predict brain responses: Estimate human fMRI responses to video, audio, language, and combined stimuli.
  • Support zero-shot analysis: The supplied research describes generalization to new stimuli, subjects, languages, and experimental tasks without requiring a new full training process for every condition.
  • Enable subject adaptation: Researchers can fine-tune the model with limited subject-specific data to produce individualized predictions.
  • Model multisensory processing: Compare how visual, auditory, and linguistic inputs contribute across the cortex.
  • Support brain-network analysis: Examine latent representations associated with functional brain networks.
  • Reduce preliminary scanning needs: Use computational predictions to help prioritize experiments before collecting additional human data.

These capabilities make the model useful for investigating how different brain regions respond to naturalistic content and how information from multiple senses is integrated.

Research use cases

TRIBE v2 is primarily intended for computational neuroscience and brain-encoding research. A laboratory could use it to estimate the likely cortical response to a candidate video or speech sample, compare predicted responses across modalities, or explore whether a proposed stimulus is likely to engage a particular functional network.

The supplied research identifies several application areas:

  • Visual functional localization, such as studying where visual information is represented in the cortex.
  • Language-network analysis involving spoken or written language.
  • Auditory and semantic-processing research.
  • Investigation of multisensory integration.
  • In-silico replication of visual and language neuroscience experiments.
  • Experimental design and pre-screening before new fMRI data collection.

For example, a researcher could compare predicted responses to a video with and without its soundtrack, or examine how language-related information changes the predicted response to audiovisual content. The model does not remove the need for validation with real participants, but it can help decide which questions and stimuli deserve further study.

Main strengths and trade-offs

TRIBE v2’s central strength is specialization. General-purpose language or vision models may describe content, classify it, or generate new material, whereas TRIBE v2 is built to connect multimodal stimulus representations with predicted human brain responses. Its training scale across more than 1,000 hours of fMRI data and 720 subjects also supports research claims about broad generalization.

Another advantage is the ability to work across modalities in one framework. This is important for naturalistic neuroscience, where a person may see a video, hear speech, and process language at the same time. Subject-specific fine-tuning provides a path toward more individualized predictions when limited data from a particular participant is available.

The trade-off is that this specialization makes the model unsuitable for most everyday AI tasks. It does not provide conversational answers, generate code, perform web searches, create media, or expose conventional production-model features such as token pricing and hosted API calls. Its outputs also require neuroscience expertise to interpret correctly.

Limitations and interpretation risks

TRIBE v2 predicts brain responses; it does not directly measure them. A prediction is not a substitute for an fMRI scanner or a new experiment. Results should be evaluated against the limitations of the training datasets, the imaging methodology, and the similarity between the modeled population and the target study.

The model’s average-subject predictions may not accurately represent every individual. Fine-tuning can adapt predictions with subject-specific data, but the supplied research does not specify a universal accuracy level or guarantee for individualized results. Zero-shot generalization to new subjects, languages, and tasks is a reported capability, not a reason to skip experimental validation.

TRIBE v2 also should not be described as mind reading. Its outputs do not recover arbitrary private thoughts, establish a person’s intentions, or provide unrestricted access to mental content. They are model estimates of neural responses to defined stimuli.

Technical and commercial availability

The release includes downloadable weights, research code, a paper, and an interactive demo. It is distributed under a CC BY-NC license according to the supplied research, so prospective users should review the license and confirm that their planned use is compatible with its non-commercial terms.

No conventional hosted pricing is documented. Input and output token prices, context-window length, maximum generated tokens, caching, batch processing, and streaming are not applicable to the released research model in the way they are for commercial language-model APIs. The model returns numerical brain-response predictions rather than a token sequence.

The supplied material also does not provide a conventional speed score, cost score, context limit, maximum output-token limit, reasoning score, or coding score. These omissions are meaningful: TRIBE v2 is evaluated as a neuroscience research system, not as a general-purpose coding or reasoning assistant.

When to choose TRIBE v2

Choose TRIBE v2 when the primary question concerns how video, audio, language, or their combination may map onto human brain responses. It is particularly appropriate for researchers planning fMRI experiments, studying brain encoding, investigating functional networks, or exploring multisensory perception before collecting more data.

Another option is more appropriate when the goal is conversation, document analysis, software development, web search, image generation, speech generation, or a production application that requires a hosted API with published latency and pricing. A conventional generative model may be faster and easier for those tasks, while TRIBE v2 is the better fit when predicted cortical responses are the required output.

Bottom line

TRIBE v2 is a specialized Meta foundation model for computational neuroscience. Its distinctive value is the ability to turn multimodal stimuli into predicted human fMRI responses, including cortical predictions that can support in-silico experiments. The model is not a chatbot or media generator, and it does not replace brain imaging. Researchers should choose it for brain-encoding and multisensory research, while treating its predictions as hypotheses that require appropriate scientific validation.


Answers to Frequently Asked Questions

Is TRIBE v2 available for commercial use?
The release includes downloadable model weights, research code, a paper, and an interactive demo. It is described as being distributed under a CC BY-NC license, so users should review the license and confirm that their planned use complies with its non-commercial terms.
Does TRIBE v2 read people’s thoughts?
No. TRIBE v2 is not a mind-reading system. It predicts neural responses to defined stimuli and does not recover arbitrary private thoughts, intentions, or unrestricted mental content. Its predictions also require validation against real fMRI measurements.
How can researchers use TRIBE v2?
Researchers can use TRIBE v2 to study brain encoding, visual and language networks, auditory and semantic processing, multisensory integration, and experimental design. It can help estimate likely cortical responses and prioritize stimuli before collecting additional fMRI data.
What inputs and outputs does TRIBE v2 support?
TRIBE v2 accepts text, audio, video, and combinations of visual, auditory, and linguistic inputs. Its primary output is a numerical prediction of fMRI brain responses across the cortex; it does not generate text, images, audio, or video.
What is TRIBE v2?
TRIBE v2 is Meta’s open-weight research model for computational neuroscience. It predicts how human brain activity may respond to naturalistic stimuli such as video, audio, speech, images, and text, producing numerical fMRI response estimates mapped to cortical space.


Sources 5
Provider

About Meta AI