An AI behavioural twin built from a two-hour interview predicted a large share of how a real person would answer questions they had never been directly asked. In the latest June 2026 revision of a study involving 1,052 US participants, interview-based agents recovered 83% of each person’s own two-week consistency on held-out General Social Survey questions. Agents given both interviews and surveys reached 86%.
That does not mean researchers created 1,052 conscious digital people. They created large-language-model agents grounded in detailed self-reports, then tested whether those agents could reproduce parts of the participants’ measured attitudes, personality and behaviour. The distinction matters. This is not a mind upload. It is something narrower, but still unsettling: a machine-built approximation of how a particular person tends to answer and decide.
How an AI Behavioural Twin Was Built From Two Hours of Conversation
The research team recruited a diverse sample of 1,052 adults in the United States and collected rich information about them. The interview component followed the American Voices Project approach, which combines structured data with open-ended questions about people’s everyday lives. Rather than asking only for age, income or political identity, interviewers explored participants’ experiences, beliefs and circumstances in natural conversation.
The resulting agents were evaluated on 150 core questions from the General Social Survey, a long-running US research programme that measures attitudes, behaviours and social characteristics. The key comparison was not whether an agent achieved perfect agreement with its human counterpart. Humans do not give identical answers every time either. Participants repeated the evaluation roughly two weeks later, giving the researchers a test-retest baseline for each person.
On the raw GSS questions, interview-grounded agents matched participants 65.67% of the time, while the people themselves were 79.53% consistent across the two test sessions. Normalising the agent result against that human baseline produced the 83% figure. Survey-only agents reached 82%, combined interview-and-survey agents 86%, and agents built from demographics alone 74%, according to the current paper.
This framing is important because an AI behavioural twin is not being asked to guess what an abstract “average voter” or “typical 35-year-old” might say. It is attempting to approximate one named participant from information that person supplied. That moves the technology away from broad demographic personas and towards individual-level simulation.
What This Means for You: Your Story May Reveal More Than Your Demographics
Most people are familiar with AI systems that categorise users by visible attributes or past clicks. This research points to a different possibility. A long conversation can contain enough interconnected clues for a model to build a more individualised representation of someone’s preferences and outlook than a short demographic profile can provide.
The result does not show that any chatbot can reliably predict everything you will do after chatting to you for two hours. An AI behavioural twin in this study was built through a carefully designed research protocol, controlled evaluations and information supplied specifically for the experiment. But the findings suggest that apparently ordinary autobiographical conversation can become structured predictive material when processed by a capable language model.
That matters as AI systems become more agent-like. LiveAIWire has previously examined why technology companies are racing to build agentic AI, where systems act across tools and workflows rather than simply answer prompts. A behavioural simulation is a different kind of agent, but combining the two ideas raises a longer-term question: what happens when software can both approximate your preferences and act on your behalf?
The Personality Results Were Stronger Than the Economic Ones
The agents were also tested on the Big Five personality framework. Interview-only agents achieved a normalised correlation of 0.80, with a raw correlation of 0.78 against the participants’ answers. Participants’ own two-week test-retest correlation was 0.95. Demographic-only agents scored lower, while agents given surveys or interviews generally performed better, the researchers reported.
Economic decision-making was a tougher test. In behavioural games, the interview agents produced a normalised correlation of 0.66. The researchers found no statistically significant overall difference between the different agent construction methods on those games. That is a useful boundary: richer personal information helped substantially in several evaluations, but it did not turn the agents into universally superior predictors of individual decisions.
This is one reason the phrase “AI copy” needs care. A useful copy of some survey responses is not a reliable duplicate of a human being. Real decisions depend on context, incentives, changing moods, new information and situations that may never have appeared in the source interview.
The Most Surprising Test Removed 80% of the Interview
In an exploratory analysis, the researchers randomly removed 80% of each interview transcript. The remaining material represented roughly 24 minutes out of the original two-hour conversation. Even then, the agents reached a normalised score of 0.79 on the GSS evaluation and 0.73 on the personality evaluation, according to the paper.
It would be wrong to conclude that a purpose-built 24-minute interview is therefore almost as effective as a two-hour one. The retained material was sampled from across the full conversation, meaning those fragments could contain information elicited only because the longer interview had explored a wide range of topics. The result is better read as evidence that the full transcript contained considerable redundancy and that useful behavioural signals were distributed throughout it.
Even with that caution, the finding sharpens the privacy question. If predictive value can survive after much of a rich conversation is removed, the important issue may not simply be how much text exists. It may be what kinds of personal relationships, beliefs and experiences the text reveals.
The AI Was Not Simply Repeating What People Had Told It
The researchers took steps to reduce direct answer copying. Questions that duplicated information already asked during data collection were excluded from the relevant evaluation. The paper also describes cases where the agent had to infer an answer from related facts. For example, if a participant had said they were a full-time student but never directly answered an employment question, the model could infer that they were less likely to have a supervisor at work.
That distinction is central to why this research is interesting. Database retrieval is familiar: tell a system your birthday and it can repeat your birthday. Behavioural simulation is more consequential when a system combines fragments of your life story and produces an answer you never explicitly gave.
The earlier wave of research on generative agents showed that language-model characters could store memories, reflect and plan inside a simulated environment. Those 2023 agents were fictional. The newer work asks a more personal question: can similar techniques be grounded in data from real people strongly enough to reproduce measurable aspects of those people’s responses?
Where the Digital Copy Starts to Break Down
The current paper is a preprint, and its authors set out several limitations. The sample was designed to be diverse and approximately representative through quotas, but it was somewhat more educated, more female and more Democratic than the US population, with regional differences as well. The work was conducted in English and tested one main model family rather than establishing that the findings generalise across every modern AI system.
There is another statistical problem. Matching individuals reasonably well does not automatically mean a synthetic population will preserve every relationship found in the real population. The authors warn that agents could reproduce individual answers while still distorting correlations between variables. Only five experimental replications were available for one part of the evaluation, which also limits strong conclusions about whether synthetic agents can estimate population-level treatment effects.
The paper therefore supports a specific claim, not a science-fiction one. Detailed self-reports can help language-model agents simulate some measured attributes of individuals better than simple demographic prompts. It does not establish that those agents are complete digital replicas, that they will predict unfamiliar high-stakes choices, or that a simulated population can replace real human research.
The Privacy Problem Starts Before the AI Says a Word
The research team treated the interviews as sensitive data. Their consent process explicitly warned participants that qualitative interviews can be difficult to anonymise because stories, political views and personal history can identify people even after names are removed. Participants were told that model capabilities could improve and potentially infer more from the data in future, and the project offered a withdrawal process for research use.
That is a notable model for responsible research because the risk is not limited to whether a system leaks a literal fact. A behavioural model may reveal patterns inferred from combinations of facts. The distinction echoes wider questions LiveAIWire has covered around the right to be forgotten in AI systems and what happens after an AI chat is deleted.
For ordinary users, the practical lesson is not to panic about every long conversation with an AI. This experiment is not evidence that consumer chatbots are secretly constructing research-grade replicas of everyone who talks to them. The more defensible conclusion is that conversational data can carry behavioural information beyond the obvious facts it contains. For an AI behavioural twin, that makes retention, consent, secondary use and access controls especially important.
Who Owns an AI Version of Your Behaviour?
There is no single answer supplied by this study, because ownership and commercial rights were not what the researchers tested. But the technical capability exposes a future policy problem. If an organisation can build a model that approximates how you respond, is that model merely an analysis of data, or does it become a new kind of personal representation?
We already recognise the intuitive difference between possessing someone’s photograph and creating a convincing impersonation of them. LiveAIWire has explored that boundary through AI deepfakes in everyday life. Behavioural simulation adds another layer because the valuable output may not look or sound like you at all. Its value lies in predicting what “you” might say.
Commercial digital twins raise similar questions. In fashion, for example, LiveAIWire has examined AI models and the licensing of digital likenesses. A behavioural twin could eventually make likeness rights look straightforward by comparison. Faces and voices are observable. A statistical approximation of someone’s preferences is harder to define, inspect and contest.
An AI Behavioural Twin Is Not You, and That Distinction Matters
The most important conclusion is also the least sensational. These agents did not recreate people. They recreated measurable slices of people well enough to perform surprisingly strongly on certain tests. The best-performing combined agents recovered 86% of participants’ own short-term consistency on held-out GSS questions, while performance varied across personality measures and economic games.
Yet that narrower achievement may be more consequential than the fantasy of a perfect digital clone. Businesses, researchers and policymakers do not necessarily need an AI that “is” you. In some settings, a system that is merely good enough to anticipate your likely answer could already be useful for testing messages, exploring scenarios or personalising services.
The same limitation creates the danger. A simulation that is accurate often enough can acquire authority before it deserves it. If its answer is treated as a substitute for a real person’s view, errors cease to be abstract benchmark misses. They become misrepresentations of somebody who may never know the synthetic version of them was consulted.
The study’s real achievement is therefore not that AI can copy a human mind. It is that two hours of structured conversation can be transformed into a surprisingly capable behavioural proxy. The next question is no longer simply whether AI can learn about us. It is how much permission a system should need before it is allowed to start answering as us.
About the Author
Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity, and the social impact of emerging technology. He publishes daily at LiveAIWire.com.
