ABSTRACT
The AI models behind ChatGPT, modern video generators, and voice assistants were never trained on brains. They learned from text, video, and audio scraped off the internet. And yet, a new paper from Meta’s FAIR lab shows that if you take what these models have internally learned and add a small layer trained on fMRI recordings (a kind of brain scan) from 720 people, the result is something close to a working model of the human brain.
You can feed it a face, a sentence, or a clip from The Bourne Supremacy, and it predicts what a typical human brain would do in response. Accurately enough that running classic neuroscience experiments on the model reproduces decades of published findings it never saw in training.
The authors compare it to AlphaFold’s moment in structural biology. The deeper story, though, is about the AI models themselves: somewhere inside them, they’ve already learned something that looks strikingly like the brain.
Every time a new AI model is released, the first questions are about what it can do. A harder, more interesting question is what these models actually contain on the inside. The patterns they form as they process information, what researchers call their internal representations, are largely a black box. We know the models work. We don’t know, in any deep sense, what they know.
TRIBE v2, published in March 2026 by Meta’s FAIR lab, offers an unexpected way to probe that question. The team took three state-of-the-art AI models (Llama 3.2 for text, V-JEPA 2 for video, Wav2Vec-Bert for audio) and asked a simple thing: if you run a movie or a spoken story through these models first, can you use what the models “see” to predict the brain activity of a real person watching or listening to the same thing? The answer is yes, well enough to produce what amounts to a digital brain you can run experiments on.
You design a stimulus, click a button, and get back a full-brain fMRI response. No participants. No scanner. The demo at aidemos.atmeta.com/tribev2 is fun to play with, but the interesting part is what the result says about AI itself.
THE IDEA: TRAIN ONE MODEL ON EVERYTHING
Instead of collecting a new controlled dataset, the authors aggregated eight existing ones. People watching Friends, The Bourne Supremacy, Hidden Figures. People listening to podcasts, audiobooks, and narrated stories. In total, over 1,100 hours of fMRI recordings across 720 participants.
They then built a model with three “eyes”: one for video, one for audio, one for text. Each eye is a frozen, off-the-shelf AI model (Video-JEPA 2, Wav2Vec-Bert, Llama 3.2). Their internal representations feed into a transformer (the same kind of neural network that powers ChatGPT), which predicts brain activity at nearly 29,000 locations across the cortex and deep brain structures. Only the transformer itself, plus a small layer that adapts to each individual person, is trained on the brain data. The vision, audio, and language models stay frozen. In other words, TRIBE v2 isn’t really learning the brain from scratch. It’s translating what frontier AI models already know into the language of brain activity.
Variants of this idea go back almost a decade, with earlier work from labs like Huth, Gallant, Jain, and Caucheteux. What’s new in TRIBE v2 is scale and reusability: three kinds of input combined in one model, trained across 720 subjects and a thousand hours of data, designed to generalize to new people and new tasks without retraining.
And the implication deserves a second look. These AI models learned from text, video, and sound, not from brain data. Yet their internal representations, once translated, line up with the functional organization of the human brain. That fact, more than any single result in the paper, is what makes this interesting. It could be a coincidence of scale. Or it could mean that the statistical structure of the world constrains any system that tries to model it, biological or artificial, in roughly similar ways. Either way, these AI models have learned something, somewhere in their weights, that looks like what the brain learned.
THE MODEL PASSES A TEST IT NEVER STUDIED FOR
After training TRIBE v2 on naturalistic stimuli (movies, podcasts, stories), the authors ran classic neuroscience experiments on it. Not on humans. On the model.
They flashed images of faces, places, and written words. They contrasted sentences versus word lists, complex versus simple sentences, and sentences about emotional versus physical pain. These are the bread-and-butter “functional localizers” that neuroscientists have used for decades to find specialized brain regions.
TRIBE v2 reproduced the findings. Faces activated the fusiform face area. Places activated the parahippocampal place area. Written words activated the visual word-form area. Language contrasts lit up the expected language regions, including Broca’s area and the temporo-parietal junction. The emotional-pain contrast recovered the regions classically associated with inferring other people’s mental states.v
None of these lab-controlled paradigms were in the training data. The model had never seen a flashing face on a gray background. It had watched Friends. And yet, asked to predict what a brain would do in a lab setting, it produced the maps neuroscientists have been publishing since the 1990s.
When the authors used a standard statistical technique to pull out the main patterns inside the model’s internal representations, the top five patterns lined up with five of the most famous functional networks in the brain: the auditory cortex, the language network, the motion network, the default mode network (active when the mind wanders), and the visual system. Nobody told the model those networks existed. They emerged from training.
WHY THIS IS A BIG DEAL
Biology had its moment in 2020 with AlphaFold, the AI system that predicts the 3D shapes of proteins. The authors of TRIBE v2 argue that something similar might be starting in neuroscience, and the case is reasonable. A foundation model of the brain (a large general-purpose model you can adapt to many questions) is a new kind of instrument: not a better microscope, more like a wind tunnel. A place to try things out in simulation before committing to the real experiment with real people.
There’s also a deeper point, and it’s squarely an AI question. Scaling laws in AI (the rule that more data and more parameters reliably produce better models) turn out to have a mirror image when you use AI to model the brain. More fMRI hours yield better brain prediction, following a clean mathematical curve called a power law. The same kind of curve that governs how well a language model predicts the next word also governs how well a model predicts the brain. That’s either a profound coincidence or a clue that intelligence, biological or artificial, might be shaped more by the statistical structure of the world than by the specific machinery that processes it.
The same kind of curve that governs how well a language model predicts the next word also governs how well a model predicts the brain. That’s either a profound coincidence or a clue that intelligence, biological or artificial, might be shaped more by the statistical structure of the world than by the specific machinery that processes it.
WHAT THIS UNLOCKS
None of the following are around the corner. None of them are science fiction either. They’re the kind of applications that become plausible the moment a model like TRIBE v2 exists.
Personalized content and education. TRIBE v2 puede adaptarse a un individuo con tan solo una hora adicional de datos cerebrales propios. A short calibration session could give you a personal brain model that predicts how you specifically respond to a piece of writing, a lecture, a therapy intervention. Educators would want this. So would anyone who makes media. So, probably, would advertisers, which is one reason this technology needs a public conversation sooner rather than later.
Designing better brain-computer interfaces. Brain-computer interfaces (devices that let people control computers directly with their thoughts) don’t just have to decode the brain. They also have to show the user something on screen or through sound, and that feedback is only as useful as the designer’s intuition about how the user will perceive it. A model that predicts brain responses to stimuli lets you stress-test those designs before running a study: which visual layout engages the right circuits, which sound is most noticeable, which interface reduces mental effort instead of increasing it.
And because Meta has released both the code and the trained model publicly, none of this requires waiting for Meta. Anyone with a powerful computer and an idea can start building.
WHAT IT CAN´T DO
A digital brain sounds like science fiction, so it’s worth being clear about what TRIBE v2 is not. It’s not thinking, it’s not conscious, and it doesn’t produce behavior. It’s a pattern-matching model: given a stimulus, predict a response. And because its target is the fMRI signal, it inherits fMRI’s limits. Real neural activity happens on the scale of milliseconds. fMRI averages activity over seconds. TRIBE v2 can tell you which brain regions activate when you hear a word, but not the millisecond-scale dance of neurons that actually produces the activation.
It also covers only three senses: vision, hearing, and language. Touch, smell, balance, and the sense of your own internal body state are all missing, as the authors flag. And it’s a snapshot of one kind of brain: healthy adults, mostly Western, in a scanner, passively perceiving. The model is a passive observer. The brain is not.
None of this is fatal, though. These are the boundaries of a first version of a new kind of tool. AlphaFold’s first version didn’t predict dynamics either.
WHERE IT GOES
Two obvious next steps pull in different directions. One is to keep scaling: more data, more subjects, more tasks, more senses. The pattern we’ve seen with other large AI models suggests that brute force will keep paying off for a while.
The other is to build models that don’t just predict brain activity but actually simulate it. Models that try to capture how thoughts form, evolve, and interact over time, from the inside out. This is closer to one of the projects we are working on at INAB, and I’ll write about it in a future post. The short version: predicting the brain and simulating the brain ask different questions, and a mature computational neuroscience will need both.
Either way, the tools are finally catching up with the ambition. The question is no longer “can we build a unified model of human cognition?” It’s “what kind of unified model do we want, and what will we do with it?” The best time to start paying attention is now, while the wind tunnel is still being built.
Fuente: d’Ascoli, S., Rapin, J., Benchetrit, Y., Brookes, T., Begany, K., Raugel, J., Banville, H., & King, J.-R. (2026). A foundation model of vision, audition, and language for in-silico neuroscience. Meta FAIR.
Demo: [aidemos.atmeta.com/tribev2](https://aidemos.atmeta.com/tribev2).
Código: [github.com/facebookresearch/tribev2](https://github.com/facebookresearch/tribev2).
Pesos: [huggingface.co/facebook/tribev2](https://huggingface.co/facebook/tribev2).
FURTHER READING
- AlphaFold (structural biology parallel). Jumper, J. et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature, 596, 583–589.
- Voxelwise encoding with language models (Huth & Gallant lab). Huth, A. G., de Heer, W. A., Griffiths, T. L., Theunissen, F. E., & Gallant, J. L. (2016). Natural speech reveals the semantic maps that tile human cerebral cortex. Nature, 532, 453–458.
- Brain–LLM alignment. Caucheteux, C. & King, J.-R. (2022). Brains and algorithms partially converge in natural language processing. Communications Biology, 5, 134.
- Encoding models with deep language models. Jain, S. & Huth, A. G. (2018). Incorporating context into language encoding models for fMRI. Advances in Neural Information Processing Systems, 31.
- Scaling laws in neural encoding of the brain. Antonello, R., Vaidya, A., & Huth, A. (2023). Scaling laws for language encoding models in fMRI. Advances in Neural Information Processing Systems, 36, 21895–21907.
- Scaling laws in AI. Kaplan, J. et al. (2020). Scaling laws for neural language models. arXiv:2001.08361.
- The original fusiform face area paper (one of the findings TRIBE v2 reproduces). Kanwisher, N., McDermott, J., & Chun, M. M. (1997). The fusiform face area: a module in human extrastriate cortex specialized for face perception. Journal of Neuroscience, 17(11), 4302–4311.










