ChatGPT Predicts Human Personality Test Results

Summary: A new study demonstrated a novel method using ChatGPT (GPT-4) to construct, validate, and predict population-level responses for personality assessment questionnaires derived from any source text. The team tested the model by generating questionnaires from two contrasting texts: the clinical Diagnostic and Statistical Manual of Mental Disorders (DSM-5) and an unscientific astrology textbook.
The LLM-generated questionnaires successfully extracted meaningful psychological signals from both texts, with the DSM-5-derived survey demonstrating high internal consistency matching the gold-standard Big Five Inventory (BFI).GPT-4 accurately predicted how human participants would answer the survey items and anticipated item correlation patterns prior to human administration.
These findings indicate that LLMs have embedded, expert-level models of human population psychology learned naturally from language training data, offering a transformative tool for rapid psychometric development and automated survey evaluation.
Key Facts
- Natural Language Personality Embeddings: GPT-4 demonstrated an innate, embedded understanding of human personality structures acquired implicitly from general internet training data without explicit psychological domain fine-tuning.
- DSM-5 vs. Astrology Comparative Validity: The DSM-5-sourced questionnaire exhibited high internal consistency matching the Big Five Inventory, whereas the astrology-sourced questionnaire showed poor internal coherence among zodiac-linked traits, reflecting real-world psychometric reality.
- Preserved Psychological Signal in Non-Scientific Texts: Despite the lack of scientific grounding in astrological zodiac assignments, ChatGPT extracted trait-relevant language that successfully predicted real-world outcomes such as depression, anxiety, and well-being at levels comparable to the BFI.
- Zero-Shot Population Response Prediction: ChatGPT accurately predicted human response means and inter-item correlation matrices before the 600-participant cohort completed the surveys, functioning as an automated evaluator of its own psychometric output.
- Cultural & Linguistic Boundaries: The authors caution that current predictive accuracy is heavily tied to Western, English-dominant training data, with ongoing research testing performance across diverse non-Occidental languages and cultures.
Source: Cell Press
Publishing August 6th in the Cell Press journal iScience, scientists outline how they developed a method for generating personality assessment questionnaires with ChatGPT from any source text.
To test the method, they applied it to both the DSM-5 and, as a deliberately unconventional example, an astrology textbook. Not only could ChatGPT be used to create and validate these questionnaires, but it could also accurately predict population-level responses before the surveys were administered.
ChatGPT and other publicly accessible large language models (LLMs) are trained on the internet by compiling trillions of human language data inputs from websites and social media. Because of this training, researchers have considered whether LLMs have an expert-level understanding of human language and, by extension, personality built into their algorithms.
“Given that personality traits are reflected in language, LLMs may have learned the structure of human personality as a natural byproduct of their training,” says lead author Rotem Monsa from the Hebrew University of Jerusalem. “So, while they were not taught specifically psychology or personality theories, these are already embedded in the language that LLMs learn from.”
To test how well LLMs can naturally assess human personality, the researchers used GPT-4 to generate two personality assessment questionnaires. They first used excerpts from the DSM-5 (a well-known text used by clinicians to diagnose mental disorders) as source text, with the assumption that personality traits are localized on the personality disorder continuum.
As a control, another questionnaire used an astrology textbook. For the former, ChatGPT generated a questionnaire with personality statements based on descriptions of personality disorders from the DSM-5. For example, based on the paranoid personality disorder section, the questionnaire asked participants to rank (1 = strongly disagree to 5 = strongly agree) how much they agree with statements like “often suspects others’ motives” or “finds it easy to trust people.”
The astrology questionnaire generated similar statements, but these were based on the source text’s assignment of personality traits to the astrological zodiac signs.
“We wanted to choose texts that describe human personality in very rich detail but also sit on opposite ends of a spectrum in terms of scientific grounding,” says Monsa. “The DSM-5 was refined through decades of clinical research and is the standard diagnostic manual in clinical psychiatry, known worldwide. The astrology text is very culturally based but not scientifically validated.”
After the LLM-based questionnaires were generated, they were given to 600 participants alongside the Big Five personality questionnaire (BFI), the most validated personality questionnaire to date, to assess their utility.
Results from the DSM-5-sourced questionnaire showed high internal consistency within personality clusters; this means traits that typically correlate together in the real world (like avoidance and dependency) also correlated in the participants’ responses. Importantly, these results mirrored those from the BFI, a finding that helps validate the strength of the questionnaire in measuring real life patterns of human psychology.
As predicted, the astrology questionnaire, by contrast, showed a weak internal consistency across traits. “Our data suggests that the astrological elements don’t reflect coherent psychological dimensions. Personality traits that were together in, for example, the fire elements don’t actually go together in real population,” Monsa says.
Despite this limitation, both questionnaires could predict life outcomes like depression, anxiety, and well-being from their responses at levels comparable to the BFI. This suggests that even though the assignment of personality traits to zodiac signs is not scientifically grounded, the personality-relevant content that LLMs extract from these texts retains meaningful psychological signal.
The most surprising result, however, was that ChatGPT could predict how participants would respond to the questionnaires before they were taken. For both questionnaires, ChatGPT anticipated the mean responses and correlations between questions with high real-world accuracy, suggesting that LLMs have an innate understanding of personality dynamics at a population level.
“The fact that LLMs can predict human response patterns before seeing any human data suggests that these models have observed something generally meaningful about human psychology,” says Monsa. “These results tell us that LLMs are not just a good content generation tool but also function as an informed evaluator of their own output.”
The authors note, however, that the results may not hold if these methods are repeated in different languages and cultures. “We would assume that in other languages and cultures, the results will be not as strong as we saw here, because LLMs were trained mostly on English texts in occidental cultures,” says Rotem. “We think it’s really interesting, and several members of our lab are currently testing it using LLMs in other languages.”
Taken together, the results demonstrate that LLMs can help researchers rapidly create personality-based questionnaires from any source material and directly test their validity for clinical use. Moreover, the combination of clinical and psychological insights with the power of AI-based tools is a transformative phase in the development of clinical and experimental psychology that may profoundly affect the field.
Key Questions Answered:
A: Because human personality traits are encoded directly into everyday language, large language models absorb the underlying structural relationships of human psychology as a natural byproduct of training on trillions of words from human text. This enables the model to translate complex source materials into psychometrically structured survey items.
A: The astrology questionnaire failed internal consistency because traits grouped under zodiac signs do not actually co-occur in human populations. However, because the individual survey items generated by ChatGPT still captured real, personality-relevant language, human responses to those specific items retained enough psychological signal to correlate with outcomes like anxiety and well-being.
A: The ability of LLMs to forecast population-level response means and item correlations allows researchers to pre-validate and refine psychometric questionnaires prior to costly, time-consuming human trials. This positions LLMs not just as content generators, but as rapid, automated evaluators for clinical and psychological research.
Editorial Notes:
- This article was edited by a Neuroscience News editor.
- Journal paper reviewed in full.
- Additional context added by our staff.
About this AI and personality research news
Author: Jordan Greer
Source: Cell Press
Contact: Jordan Greer – Cell Press
Image: The image is credited to Neuroscience News
Original Research: Open access.
“Generating and analyzing personality questionnaires using large language models” by Rotem Monsa, Aviv Zohar, Shahar Arzy. iScience
DOI:10.1016/j.isci.2026.116909
Abstract
Generating and analyzing personality questionnaires using large language models
The Five-Factor Model (“Big Five”) is based on the lexical hypothesis that personality traits are encoded in language. Large language models (LLMs) offer new ways to explore personality through text. We developed an LLM-based method to generate and validate personality questionnaires from textual sources.
Using the DSM-5 personality disorders section and a popular astrology book, we generated two questionnaires and administered them, alongside the Big-Five Inventory (BFI), to 600 adults. Internal consistency was high for the BFI and the DSM-based questionnaire but low for the astrology-based one.
The LLM predicted human response patterns for both generated questionnaires. Individual items from both the DSM- and astrology-based questionnaires predicted diverse life outcomes at levels comparable to the BFI.
These findings show that LLMs can construct personality measures and anticipate response patterns, offering a scalable framework for corpus-based personality research.



