Summary: Researchers discovered that ChatGPT can evaluate social interactions shown in images and videos with accuracy close to that of human raters. The AI’s ratings of social features—such as cooperation, hostility, facial expressions, and body movements—were, on average, more internally consistent than ratings from a single human observer.
By replacing thousands of human annotation hours with automated AI assessments, the team saved more than 10,000 work hours. This cost-effective, scalable approach has implications beyond basic neuroscience research: it could streamline patient monitoring in healthcare, improve consumer-response analysis in marketing, and enhance automated detection in security applications.
Key Facts
- Human-level annotation: ChatGPT produced evaluations for 138 social traits in images and videos that closely matched human judgments.
- Major efficiency gains: The AI performed in hours what took over 10,000 human work hours to complete.
- Broad applicability: Potential uses include brain imaging research, clinical monitoring, marketing evaluation, and security surveillance.
Source: University of Turku
Every day, people form rapid impressions of others’ behavior and interactions.
Modern large language models with visual capabilities—like ChatGPT and GPT-4V—are able to describe scenes and identify objects in images and video. What remained unclear was whether these models could also infer more nuanced, high-level social information from visual material.

Researchers at the Turku PET Centre in Finland tested this directly. They asked a popular language-and-vision model to rate 138 predefined social features across a collection of images and movie scenes. These features covered a wide spectrum, from individual facial expressions and gestures to interaction-level qualities like cooperation and hostility.
To validate the AI’s performance, the team compared its annotations with over 2,000 human evaluations for the same visual material. Results showed that the model’s judgments were highly similar to human ratings and, importantly, displayed less variability than ratings from a single human participant.
“Because ChatGPT’s evaluations were, on average, more consistent than a single person’s ratings, they can be considered reliable,” says Postdoctoral Researcher Severi Santavirta from the University of Turku. He adds that aggregating multiple human raters still yields the most accurate ground truth, but that AI offers a powerful alternative for large-scale annotation tasks.
How AI can accelerate neuroscience research
In the study’s second phase, the researchers used both human and AI annotations to model brain activity related to social perception using functional neuroimaging. Before mapping neural responses to social content, the visual stimuli need to be annotated for their social features—a labor-intensive step where automated models can provide major benefits.
“When we mapped brain networks of social perception using either ChatGPT’s annotations or human annotations, the resulting neural maps were strikingly similar,” Santavirta notes. This alignment demonstrates that AI-derived stimulus models can effectively stand in for human annotations when investigating how the brain processes social information.
The practical advantage is substantial: gathering human annotations required contributions from more than 2,000 participants and over 10,000 work hours, whereas ChatGPT produced comparable ratings in just a few hours. Automating this step reduces data-processing costs and speeds up the workflow for large neuroimaging studies.
Practical applications: healthcare, marketing and security
Although the research focused on mapping social perception in the brain, the implications extend to many real-world settings. Automated social evaluation from video could help clinicians and care staff monitor patient behavior and well-being more continuously and objectively. In marketing, AI could predict how audiences might respond to audiovisual content. In security, automated social-feature detection could flag abnormal or potentially risky interactions captured by surveillance systems.
“AI does not tire and can operate continuously,” Santavirta explains. “As models improve, routine monitoring of increasingly complex social situations may be handled by artificial intelligence, while human experts focus on validating the most critical findings.”
About this AI and social perception research news
Author: Tuomas Koivula
Source: University of Turku
Contact: Tuomas Koivula – University of Turku
Image: Image credit: Neuroscience News
Original Research: Open access.
“GPT-4V shows human-like social perceptual capabilities at phenomenological and neural levels” by Severi Santavirta et al. Imaging Neuroscience
Abstract
GPT-4V shows human-like social perceptual capabilities at phenomenological and neural levels
Humans rapidly extract social features from the people and interactions around them, using these cues to navigate social situations. Recent large language models with visual understanding capabilities have demonstrated detailed scene and object recognition, prompting the question of whether they can also infer complex social information and whether their internal feature representations align with human perceptual structure.
The study collected GPT-4V annotations for 138 social features across 468 images and 234 video clips sourced from social movie scenes, and compared them with 2,254 human annotations. Results show that GPT-4V can annotate individual social features at levels comparable to humans and that the overall structure of feature correlations generated by the model resembles human social-perceptual structure.
Using annotations from both humans and GPT-4V, researchers modeled hemodynamic responses in 97 participants viewing socioemotional movie clips. The GPT-4V–based stimulus models identified the same social-perceptual brain networks as models built from human annotations, demonstrating that AI-derived labels can reveal neural patterns of social perception similar to those revealed by human labels.
These human-like annotation capabilities of advanced language-and-vision models could enable new applications across healthcare, business, and scientific research, and open promising directions for both psychological and neuroscientific studies.