Can ChatGPT Predict Human Personality Test Results?

Summary: A new study shows a novel approach using ChatGPT (GPT-4) to generate, validate, and predict population-level responses to personality assessment questionnaires derived from any source text. The research team tested this method by creating questionnaires from two very different texts: the clinical Diagnostic and Statistical Manual of Mental Disorders (DSM-5) and a popular astrology book.

Questionnaires produced by the large language model (LLM) captured meaningful psychological signal from both sources. The DSM-5–based survey showed strong internal consistency and patterns that aligned closely with the established Big Five Inventory (BFI). GPT-4 also forecasted how people would answer the questionnaires and predicted the pattern of correlations among items before any human data were collected.

These results suggest that modern LLMs contain implicit, expert-level models of human personality learned from language exposure, offering a rapid and scalable tool for psychometric development and automated survey evaluation.

Key Facts

  • Natural language personality embeddings: GPT-4 appears to have internalized structural relationships among personality traits from its broad language training, without targeted psychological fine-tuning.
  • DSM-5 vs. astrology — comparative validity: The DSM-5–derived questionnaire demonstrated strong internal coherence comparable to the BFI, while the astrology-derived instrument showed weak internal consistency, mirroring real-world psychometric expectations.
  • Psychological signal preserved in non-scientific texts: Despite astrology’s lack of scientific grounding, the LLM extracted trait-relevant language that still predicted outcomes such as depression, anxiety, and well-being at levels similar to the BFI.
  • Zero-shot population response prediction: GPT-4 accurately forecasted mean responses and inter-item correlation matrices for both questionnaires before data collection, functioning as an automated evaluator of its own outputs.
  • Cultural and linguistic limitations: The authors caution that predictive accuracy likely depends on the model’s English-dominant, Western-focused training data and that performance in other languages and cultural contexts remains under investigation.

Source: Cell Press

Published August 6 in the Cell Press journal iScience: Scientists describe a reproducible method for using ChatGPT to generate personality questionnaires from any textual corpus and to assess those questionnaires’ likely psychometric properties before human administration.

To demonstrate the approach, the researchers had GPT-4 produce two distinct surveys: one based on the personality-disorders section of the DSM-5 and another based on a mainstream astrology book. The DSM-5 excerpts reflect decades of clinical research and are widely used in psychiatric diagnosis. The astrology text was chosen deliberately as an example of richly descriptive but non-scientific source material.

From the DSM-5 excerpts, GPT-4 generated items phrased as statements that respondents rated on a 1-to-5 agreement scale (1 = strongly disagree, 5 = strongly agree). Examples inspired by the paranoid personality disorder section included items like “I often suspect others’ motives” or “I find it easy to trust people” (reverse-scored). The astrology-based questionnaire produced comparable first-person statements derived from zodiac descriptions in the source text.

The two LLM-generated surveys were administered to 600 adults alongside the Big Five Inventory (BFI), the current gold standard for broad personality measurement. Responses were analyzed for internal consistency, inter-item correlations, and the ability to predict life outcomes such as depression, anxiety, and overall well-being.

Results showed that the DSM-5–based instrument had high internal consistency across meaningful personality clusters: items that should correlate in real-world data did so. Those correlations and scale structures resembled patterns observed in the BFI, supporting the DSM-based survey’s construct validity. In contrast, the astrology-based survey displayed weak coherence among its hypothesized trait groupings, indicating that zodiac-based trait groupings do not map onto stable psychological dimensions in the sampled population.

Nevertheless, individual items from both questionnaires carried predictive value. Even when internal consistency was low (as with the astrology-based instrument), item-level responses still correlated with mental-health and well-being outcomes at levels comparable to the BFI. This implies that LLMs can extract psychologically meaningful language from diverse texts, even if the original source’s conceptual groupings are not psychometrically sound.

Perhaps the most striking finding was GPT-4’s ability to predict population-level response statistics before any participants completed the surveys. For both the DSM- and astrology-derived questionnaires, the model produced accurate forecasts of mean item responses and correlation matrices, demonstrating that LLMs can serve as rapid, zero-shot evaluators of questionnaire structure and likely performance.

Lead author Rotem Monsa (Hebrew University of Jerusalem) notes that because personality traits are reflected in everyday language, LLMs naturally absorb the underlying relationships among traits. The team emphasizes, however, that model performance may vary across languages and cultures and that further research is needed to evaluate non-Western and non-English contexts.

Overall, the study offers a practical framework for corpus-driven personality research: LLMs can generate candidate measures from any text, predict their psychometric behavior, and identify promising items for more traditional human validation. This combination of clinical insight and AI-assisted methodology could accelerate questionnaire development and expand the toolkit available to clinical and experimental psychologists.

Key Questions Answered:

Q: How does ChatGPT build personality questionnaires without being explicitly trained in psychology?

A: Because personality-related patterns are encoded in everyday language, large language models learn structural relationships among traits simply by training on vast amounts of text. That implicit knowledge lets the model translate descriptive source material into psychometrically structured survey items.

Q: Why did the astrology-based questionnaire show low internal consistency while still predicting well-being?

A: Zodiac groupings do not reflect coherent personality dimensions in real populations, so internally the scale was weak. Still, many individual items contained language tied to genuine personality variation, allowing those items to predict outcomes like anxiety and well-being.

Q: What is the clinical significance of an LLM predicting survey responses before humans take them?

A: Pre-validating instruments with LLMs can save time and resources by highlighting likely strengths and weaknesses in questionnaire design before costly human trials, enabling faster iteration and refinement.

Editorial Notes:

  • This article was edited by a Neuroscience News editor.
  • The journal paper was reviewed in full by our editorial staff.
  • Additional context and explanatory material were added by staff writers.

About this AI and personality research news

Author: Jordan Greer — Cell Press
Source: Cell Press
Contact: Jordan Greer, Cell Press
Image: Image credited to Neuroscience News

Original Research: Open access. “Generating and analyzing personality questionnaires using large language models” by Rotem Monsa, Aviv Zohar, and Shahar Arzy. Published in iScience. DOI: 10.1016/j.isci.2026.116909


Abstract

Generating and analyzing personality questionnaires using large language models

The Five-Factor Model (Big Five) rests on the lexical hypothesis: personality traits are encoded in language. Large language models provide new tools to explore personality directly from text. This study presents an LLM-based pipeline to generate and validate personality questionnaires from textual corpora.

Using the DSM-5 personality-disorders section and a widely read astrology book, the authors generated two questionnaires and administered them, alongside the Big Five Inventory (BFI), to a sample of 600 adults. Internal consistency was high for the BFI and the DSM-based instrument but low for the astrology-based instrument. The LLM accurately predicted population-level response patterns for both questionnaires. Individual items from both generated instruments predicted diverse life outcomes at levels comparable to the BFI.

These findings indicate that LLMs can construct viable personality measures from text and anticipate their response patterns, offering a scalable framework for corpus-based personality research and rapid psychometric evaluation.