LLMs Can Generate Personality Tests and Predict Answers, Israeli Researchers Find
AI That Knows You Better Than You Know Yourself?
Artificial intelligence models are no longer just answering questions—they are designing them and predicting how humans will respond. In a new study, Israeli researchers have demonstrated that large language models (LLMs) such as OpenAI’s ChatGPT and Google’s Gemini can be harnessed to create personality-test items and to accurately forecast how people will answer, a capability that could reshape psychometrics, automated survey design and behavioral profiling.
The findings, reported by The Jerusalem Post, raise fundamental questions about how AI might be applied in hiring, marketing, education and mental-health screening, while also renewing debate over the reliability and ethics of letting opaque algorithms interpret human psychology.
From Answering Machines to Test Designers
Popular chatbots have long impressed users with their conversational fluency, but the new research goes a step further. The study claims that LLMs can generate novel, valid personality-test questions and, given a description of a hypothetical person or a target demographic, predict the distribution of responses. This suggests that AI could serve as a low-cost engine for creating and pre-testing psychometric instruments without always requiring human panels.
The Jerusalem Post report does not specify the exact personality framework used, but such studies typically rely on well-established models like the Big Five (Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism) because they provide a stable vocabulary for comparison. The researchers are said to have evaluated the accuracy of AI-generated items by comparing model predictions against human response data, a validation step crucial for any real-world application.
What the Findings Actually Show
It is important to distinguish exactly what the research demonstrates. Rather than a product launch, this is a methodological paper exploring three distinct capabilities:
- Item generation: LLMs can write questions that mimic the style and content of traditional personality inventories.
- Response prediction: Given demographic or behavioral prompts, the model can estimate how a person—or group of people—would answer those items.
- Inference from text: In principle, LLMs already analyze linguistic patterns to infer personality traits, but the new study focuses more on the generative and predictive powers applied to formal testing.
This separation matters because each capability carries different risks and requires different validation. Generating questions is relatively uncontroversial; predicting answers touches on the model’s internal world-model of human behavior, which may mirror stereotypes or biases present in training data.
Potential Use Cases and the Double-Edged Sword
If reliable, AI-generated personality assessments could speed up recruitment screening, personalize marketing campaigns, and support large-scale academic surveys. In mental health, automated tools could flag individuals for follow-up, though clinicians caution that such systems should augment, not replace, professional judgment. The technology’s reach, however, immediately triggers red flags.
Privacy advocates warn that personality predictions derived from text or behavioral traces could be used to manipulate consumers or unfairly filter job candidates. The American Psychological Association has long stressed that psychological tests must meet rigorous standards for validity, fairness, and transparency—standards that today’s LLMs are not designed to satisfy. As the APA’s guidelines note, testing should be evidence-based and free of cultural bias, a high bar for models trained on vast, uncurated internet data.
Safeguards, Limits and the Path Ahead
Despite the promise, the reliability of LLM-driven personality tests remains fiercely contested. Critics point out that even state-of-the-art models can produce plausible-sounding but nonsensical or biased outputs. Without independent replication and extensive human trials, any practical use is premature. The Israeli team’s work likely underscores the need for transparency, auditable datasets and clear disclosure when AI is involved in generating or interpreting assessments.
Both OpenAI and Google have published research on AI safety and model behavior—see OpenAI Research and Google Gemini AI—but neither company has announced clinical-grade personality-testing tools. The study from Israel thus fits into a broader scientific conversation: exploring the frontier where artificial intelligence meets human complexity, and urging society to prepare guardrails before the technology outpaces regulation.
“These models are not just analyzing text; they are beginning to model the people who produce it,” an observer of the research told The Jerusalem Post. “That’s a profound shift—and one we must handle with care.”
As governments and professional bodies worldwide grapple with AI governance, the Israeli findings serve as a timely reminder that the capacity to simulate human judgment does not automatically confer the wisdom to deploy it safely.




