Using sets of LLM personas as stand-ins for real human populations is a key challenge in LLM-based social simulation. Recent methods use LLMs to generate personas from real social media histories or via structured sampling, and align them to a reference distribution such as Big Five psychometrics. But imagine a synthetic population whose Big Five personality profile is statistically indistinguishable from that of a real survey population. Ask these personas for their opinions on a policy, or observe how they actually behave in a simulated decision, and the match may fall apart: aligning a persona set to one reference distribution guarantees nothing about the others, and recent evidence shows that a persona's self-reported traits often dissociate from its behaviour. Instead of aligning the persona set to a single psychometric distribution, we aim to align it to multiple reference distributions simultaneously — spanning psychometric, opinion, and behavioural responses — and study whether jointly aligned persona sets generalize better to distributions and simulation tasks they were never aligned on. The team is co-led by Dr. Aditya Joshi, a Senior Lecturer in Natural Language Processing (NLP), and Haokai Zhao, a PhD student, in the UNSW-NLP research group.
The ideal student will have strong programming skills in Python. A good grounding in statistics would be highly regarded.
Computer Science and Engineering
LLM | Natural language processing
No
- Research environment
- Expected outcomes
- Supervisory team
- Reference material/links
The student will be a part of the natural language processing (NLP) research group consisting of postdocs, software engineers and PhD students. The student will have access to typical computing facilities at UNSW.
- Reproduction of existing persona synthesis pipelines.
- A benchmark of existing persona sets on survey and behavioural response distributions.
- Develop new methods for synthesizing persona sets that generalize to unseen survey and behavioral distributions.
- Well-documented code with accompanying documentation.
- A report in a form suitable for a research paper.
- Population-Aligned Persona Generation for LLM-based Social Simulation
- LLM Generated Persona is a Promise with a Catch
- Will Scaling Improve Social Simulation with LLMs?
- Scaling Synthetic Data Creation with 1,000,000,000 Personas
- The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
- Synthia: Scalable Grounded Persona Generation from Social Media Data