New Study: AI Surpasses Humans in Emotional Intelligence

Summary: A new study evaluated whether artificial intelligence can display emotional intelligence by putting six generative AIs, including ChatGPT, through standard emotional intelligence (EI) assessments. On average, the AIs scored 82% correct, substantially higher than the 56% average reported for human participants.

The evaluated systems not only selected emotionally appropriate responses with strong accuracy but also produced new, reliable EI test items quickly. These results indicate practical potential for AI-assisted tools in emotionally sensitive areas—such as education, coaching, and conflict resolution—when those tools are used under expert supervision.

Key Facts:

  • AI Emotional IQ: Multiple generative LLMs outperformed humans on established emotional intelligence tests, averaging 82% versus 56% for humans.
  • Test Creation: ChatGPT-4 generated new EI test items that matched expert-designed assessments in realism and clarity.
  • Practical Applications: The findings highlight possibilities for AI support in coaching, classroom settings, and conflict management, provided experts oversee their use.

Source: University of Geneva

Can AI recommend suitable behaviour in emotionally charged situations?

Researchers from the University of Geneva (UNIGE) and the University of Bern (UniBE) tested six leading large language models (LLMs) on performance-based emotional intelligence measures. The LLMs included ChatGPT-4, ChatGPT-o1, Gemini 1.5 Flash, Copilot 365, Claude 3.5 Haiku, and DeepSeek V3.

The team used five widely used EI tests that present emotionally challenging scenarios and require choosing or generating the most emotionally intelligent responses. These measures assess abilities such as recognizing others’ emotions, understanding the causes and consequences of emotions, and selecting appropriate regulatory or interpersonal actions.

Example items involved workplace conflicts, social slights, and interpersonal misunderstandings. One illustrative scenario asked how Michael should respond after a colleague takes credit for his idea. Options ranged from direct confrontation to reporting to a supervisor, harbouring resentment, or retaliating. The researchers classified seeking a conversation with a superior as the most adaptive choice in that scenario, reflecting a constructive, professional response.

Across the test battery, the LLMs achieved an average accuracy of 82%, substantially higher than the 56% human average documented in the validation studies for these measures. According to the authors, this suggests that contemporary LLMs not only identify emotion-related cues but also reason about appropriate emotional behaviour in many standard scenarios.

Generating new tests quickly and reliably

In a second phase, the researchers asked ChatGPT-4 to produce entirely new items for each EI test. Those AI-generated tests were administered to more than 400 human participants across five studies. The ChatGPT-generated items proved comparable to the original items in terms of difficulty, clarity, and realism—despite the original tests having taken years of expert development.

Although some small differences emerged on secondary measures (such as diversity of item content or certain psychometric correlations), these differences were minor and statistically limited: none exceeded a medium effect size. Original and AI-generated versions also correlated strongly with one another, suggesting the new items tapped similar emotional knowledge and reasoning.

The ability to both solve and create EI test items supports the view that LLMs possess knowledge about human emotions and can apply that knowledge to generate plausible, context-appropriate responses. This capability opens the door for AI-supported resources that aid training, assessment, and practice in emotional skills—so long as professionals guide and validate their use.

About this AI and Emotional IQ research news

Author: Antoine Guenot
Source: University of Geneva
Contact: Antoine Guenot – University of Geneva
Image: The image is credited to Neuroscience News

Original Research: Open access. “Large language models are proficient in solving and creating emotional intelligence tests” by Marcello Mortillaro et al., published in Communications Psychology.


Abstract

Large language models are proficient in solving and creating emotional intelligence tests

Large language models (LLMs) are increasingly expert across many domains, but their capacity for emotional intelligence has been uncertain. This research evaluated whether LLMs can both solve and generate performance-based emotional intelligence tests.

Six LLMs—ChatGPT-4, ChatGPT-o1, Gemini 1.5 Flash, Copilot 365, Claude 3.5 Haiku, and DeepSeek V3—participated in five standard EI assessments. They outperformed human averages reported in validation studies, with mean accuracy near 81–82% versus 56% for humans. In a follow-up phase, ChatGPT-4 generated new test items that were then administered to human participants (total N = 467) alongside the originals.

Overall, AI-generated tests displayed comparable difficulty and were rated similarly in clarity and realism. Differences on secondary psychometric metrics were small and fell below thresholds for medium effect sizes. Original and AI-generated tests showed a meaningful positive correlation, indicating consistency between formats.

These findings suggest that modern LLMs can produce and apply knowledge about emotions and emotional regulation in ways that resemble human emotional reasoning. With appropriate expert oversight, such capabilities may enhance tools for education, coaching, and conflict management without replacing human judgement.