Summary: A research team combined a large dance video dataset with state-of-the-art AI models and fMRI to map how the human brain perceives dance. Their work shows that higher-order, cross-modal brain regions—those that integrate sight and sound—explain neural responses to choreography far better than simple motion or audio cues alone, and that expert dancers exhibit more varied and nuanced neural response patterns than nonexperts.
By aligning functional MRI recordings with AI-derived cross-modal features, the researchers revealed how predictive AI architectures and human perception converge when interpreting movement and music together. The results point to new methods for studying artistic perception and the neural basis of aesthetic experience.
Key Facts:
- Cross-modal encoding dominates: Combined music-and-movement features explained brain responses more effectively than motion-only or sound-only descriptors.
- Expert brains respond differently: Professional and experienced dancers showed richer, more differentiated neural patterns while watching dance, indicating broader perceptual representations.
- AI mirrors human prediction: The choreography model’s next-step prediction mechanism aligned closely with how the brain integrates audiovisual information during dance perception.
Source: University of Tokyo
Dance is a cultural and sensory form of expression that engages rhythm, movement, and emotion across societies.
Researchers at the University of Tokyo, working with collaborators, used a modern dance video repository and generative AI to study how the brain represents complex, real-world performances. Their experiment focused on identifying where and how the brain binds music and movement into a unified perception of dance.
Traditional laboratory studies of dance have often used simplified movements, isolated sounds, or coarse categorical labels that limit the scope of inference. In contrast, this study relied on a large, naturalistic dataset of recordings and an AI model that captures both motion and musical structure, enabling fine-grained, cross-modal mapping between video features and localized brain activity.
The project was led by Professor Hiroshi Imamizu of the University of Tokyo and Associate Professor Yu Takagi of the Nagoya Institute of Technology. Building on advances in data-driven encoding and generative modeling, the team used AI-derived feature sets to predict fMRI responses while participants viewed hundreds of real dance clips spanning multiple styles and performers.
“We wanted to understand how the brain organizes and predicts body movement in everyday life; dance provides a rich and natural testbed,” said Imamizu. The team collaborated closely with street dance professionals and ballet practitioners to ground their conceptual and practical approach in authentic performance practices.
A key resource was a large public dataset of dance recordings that spans many street dance genres. That dataset, together with generative choreography models trained to predict movement from music, made it possible to extract high-dimensional, cross-modal descriptors that capture both motion dynamics and musical context.
To probe the relationship between these AI features and neural activity, the researchers recruited 14 participants with varied dance experience and recorded their brain activity while they watched 1,163 short dance videos. Using encoding models, they tested how well different feature sets—motion-only, audio-only, and combined cross-modal representations—predicted patterns of fMRI activation.
Results consistently showed that cross-modal features explained activity in higher-order association cortex better than unimodal features. These brain regions integrate auditory and visual information and appear central to perceiving choreography as a coherent audiovisual event rather than as separate streams of motion or sound.
The next-motion prediction mechanism found in the AI choreography model closely paralleled how these association areas represent upcoming movement and synchronize it with musical structure. In other words, prediction-based AI models and biological perception seem to use related computational principles to bind music and movement.
To connect neural patterns to subjective experience, the team consulted expert dancers to generate a set of descriptive concepts and carried out an online rating survey. They processed those ratings through an in silico brain-activity simulator derived from their encoding models. The simulator showed that different emotional and aesthetic impressions correspond to distinct, distributed neural signatures rather than a single scalar metric.
Intriguingly, the simulator predicted expert responses more accurately than those of nonexperts. Experts’ neural patterns were not only better explained by the AI-derived features but were also more diverse across individuals. This suggests that expertise expands the range of perceptual and emotional representations evoked by dance, producing a wider variety of neural states rather than a narrower, convergent profile.
“Our findings imply that experience enriches perceptual representations in art,” Imamizu explained. “By combining tightly controlled neuroimaging approaches with large, naturalistic datasets and generative AI, we can begin to unpack the complexity of aesthetic experience across senses.”
The team hopes their brain-activity simulator and cross-modal framework will be applied to creative practice—helping choreographers and artists explore new styles that resonate emotionally—and extended to other art forms where music, movement, and multisensory integration shape perception.
Key Questions Answered:
A: Expert observers exhibit more diverse and finely differentiated neural responses, reflecting broader perceptual and emotional representations shaped by experience.
A: Higher-order association areas that integrate auditory and visual inputs show stronger correspondence with cross-modal features than do regions sensitive to motion or sound alone.
A: Generative AI models that predict future movement from music produce cross-modal representations whose predictive structure parallels human audiovisual integration, allowing researchers to link model features to brain activity.
Editorial Notes:
- This article was edited by a Neuroscience News editor.
- Journal paper reviewed in full.
- Additional context added by our staff.
About this AI, neuroscience, and dance research news
Author: Rohan Mehra
Source: University of Tokyo
Contact: Rohan Mehra – University of Tokyo
Image: The image is credited to Neuroscience News
Original Research: Open access. “Cross-modal deep generative models reveal the cortical representation of dancing” by Hiroshi Imamizu et al. Nature Communications (DOI: 10.1038/s41467-025-65039-w)
Abstract
Cross-modal deep generative models reveal the cortical representation of dancing
Dance is a universal, multimodal art form that offers a window into cognition, emotion, and sensory integration. Quantitative, fine-grained descriptions of how its combined motion and music components are represented in the brain have been scarce. Here, we link features derived from a cross-modal deep generative model of dance to functional MRI responses recorded while participants watched naturalistic dance videos. We show that cross-modal features predict dance-evoked brain activity better than low-level motion or audio features alone. Using encoding models as computational simulators, we quantify how dances that evoke different emotions produce distinct neural patterns. Although expert dancers’ brain activity is more fully explained by these dance features than novice observers, experts demonstrate greater individual variability. This approach connects generative model representations with naturalistic neuroimaging to clarify how motion, music, and experience together shape aesthetic and emotional responses to dance.