AI Ethics & Privacy

AI Agents Reflected Their Owners and Leaked Personal Details

ai agents digital twins leaked personal details
AI agents can reflect their owners closely enough to reveal personal details, raising new questions about privacy and digital identity.

AI agents can carry traces of their owners into public conversations in ways that create a new AI agent privacy problem. A preprint analysing 10,659 matched human-agent pairs on Moltbook found systematic similarities in topics, values, emotional patterns and writing style, while 34.6% of agents were flagged as revealing at least one piece of owner-related personal information that was not in their public configuration.

AI agent privacy begins with behavioural transfer

The researchers paired autonomous agents on Moltbook with the Twitter or X accounts of the humans publicly linked as their owners. They then compared behaviour across 43 features spanning topics, values, affect and linguistic style. The central question was whether an agent simply sounded like a generic product of its model, or whether it systematically reflected the particular person associated with it.

The work was also presented through Wharton Human-AI Research, which highlighted the owner-disclosure finding. For wider regulatory context, the UK’s Information Commissioner’s Office warns that agentic systems can create new privacy risks when they combine, retain and act on personal information across tools and services.

The study found that matched owner-agent pairs were more similar than random pairs across many of those features. The transfer was not confined to agents whose public biographies explicitly described their owners. It persisted among agents without those configurations, and pairs that aligned on one behavioural dimension often aligned on others.

That does not prove an agent is a psychological copy of its owner. Similarity can arise through prompts, memory, files, examples, repeated interactions or the wider computer environment the agent can access. The paper argues that accumulated context is consistent with the observed pattern, but its observational design cannot identify one universal mechanism.

More than a third of agents triggered a disclosure flag

The privacy analysis examined 44,588 agent posts in the final matched sample. Using a validated LLM-as-judge pipeline, the researchers identified 6,220 posts containing high-confidence owner-referential information across six categories. At agent level, 3,685 of 10,659 agents, or 34.6%, had at least one detected disclosure.

Occupational information was the most common category, appearing among 27.3% of all agents. Location-related information appeared for 12.1%, relational information for 6.1%, behavioural details for 4.6%, financial information for 4.0% and health information for 1.0%. Those percentages can overlap because one agent could disclose more than one kind of information.

The classification is important context. Researchers did not manually verify every public post as a legally defined privacy breach. They used an automated judging pipeline, restricted the headline results to high-confidence classifications and excluded information already present in the agent’s configured biography. The result therefore indicates detected owner-referential disclosure under the study’s method, not a court finding that 34.6% of agents leaked protected data.

Stronger behavioural transfer was associated with more disclosure

The authors also tested whether agents that more closely resembled their owners were more likely to surface personal information. In the full sample, a one-standard-deviation increase in the holistic behavioural-transfer score was associated with a 1.32 percentage-point higher probability of at least one disclosure post after controls.

The association remained positive across several robustness checks and became larger in some subsets with richer posting histories. That is evidence of a relationship, not proof of causation. More active owners and agents create more observable behaviour, and more active agents also have more opportunities to mention something personal. The authors explicitly avoid claiming that behavioural transfer itself caused the disclosure.

That caution matters because the mechanism is easy to dramatise. An agent need not be secretly profiling its owner before a privacy risk appears. If it has useful context about the person and is encouraged to act autonomously in public, ordinary generation can turn that context into an unintended statement.

Moltbook makes the privacy problem unusually visible

Moltbook is a social network built for autonomous AI agents, making it a useful environment for observing agents interacting in public. It is also an unusual platform, so the findings should not be assumed to describe every workplace assistant, coding agent or consumer chatbot.

LiveAIWire has already covered another Moltbook-derived experiment in which targeting a small share of agents contributed to wider polarisation in a simulated community. That work focused on how information stored in agent memory could propagate through a network. The new privacy study turns the lens back towards the owner and asks what personal context can travel outward.

The two results share a broader point. Persistent memory makes agents more useful because it lets them carry information across tasks and conversations. The same persistence means information can appear in a context different from the one in which it was originally supplied. Memory is therefore both a capability and a boundary-management problem.

A useful agent may know exactly the things you would rather it not publish

An agent that books travel, manages communications or organises work may need to know preferences, relationships, schedules and professional details. Some systems may also have access to email, documents or connected applications. The practical value comes from context, but context creates something worth protecting.

Britain’s National Cyber Security Centre has warned that AI agents need restricted permissions, monitoring and reliable ways to stop autonomous activity. The Moltbook findings add another reason for least-privilege design. An agent should not receive broad personal context merely because a connection is technically available.

The privacy control also has to operate at output time. A system may legitimately know an address or medical appointment while still being prohibited from mentioning it in a public post. That requires separating what an agent can access from what it is allowed to disclose, with additional checks when the destination is public or external.

Behavioural twins show why context can become surprisingly predictive

Previous research has shown that relatively rich personal data can create agents that reproduce parts of an individual’s measured behaviour. LiveAIWire reported on 1,052 interview-based behavioural twins that predicted a large share of their participants’ later survey responses. Those systems were deliberately built to imitate people under controlled research conditions.

The Moltbook study is different because the transfer appears in ordinary public agent behaviour rather than a benchmark explicitly designed for simulation. That makes the privacy question more concrete. A system does not have to recreate a person perfectly to reveal patterns that help others infer occupation, location, relationships or preferences.

Inference risk grows when several weak signals can be combined. A single phrase may reveal little. Repeated vocabulary, posting times, recurring topics and references to places can collectively narrow down who a person is or what they do. Privacy protection therefore cannot focus only on obvious secrets such as account numbers.

The detection pipeline is itself an AI system

There is an additional layer of uncertainty because the study uses an LLM to classify disclosures. The authors report validation and simulation-based robustness checks, but automated labels can produce false positives and false negatives. Their supplementary analysis estimates classification error and tests whether the main transfer-disclosure association survives plausible relabelling noise.

That is a reasonable research approach at this scale, but it is another reason not to turn the percentages into claims about confirmed individual harm. A post classed as occupational disclosure might reveal a job detail the owner does not consider private, while a subtle disclosure could evade the classifier. The study measures a systematic signal across a platform rather than adjudicating each case.

The paper is also a preprint dated 21 April 2026. It has not yet been through journal peer review. Its large dataset and robustness checks make it worth attention, but independent replication on other agent platforms would materially strengthen the conclusion.

Privacy controls need to follow the agent beyond the chat window

Traditional chatbot privacy advice focuses on what a user types into a conversation. Agentic systems make the boundary more complicated. They can act later, use persistent context and communicate with services or audiences the user is not watching in real time.

That means a sensible privacy architecture should treat destination as a permission. Posting publicly, sending an external email and writing into a private local note are not equivalent actions even if the content originates from the same memory. The more public the destination, the stronger the disclosure check should be.

Organisations also need logs that make it possible to reconstruct what an agent accessed and why a particular statement was produced. Without that trail, a privacy incident can be difficult to investigate because the relevant context may be spread across prompts, files, memories and external tools.

LiveAIWire’s reporting on AI assistance hiding the boundary between tool output and human capability describes a different measurement problem, but the governance lesson is similar. When AI becomes embedded in ordinary activity, organisations need to preserve visibility into which parts came from the person and which came from the system.

The agent can become a public shadow of its owner

The strongest interpretation supported by the Moltbook study is not that agents become digital clones. It is that they can absorb enough owner-specific context to produce measurably owner-like behaviour, and that stronger resemblance is associated with greater risk of owner-referential disclosure.

That creates a privacy category that ordinary account settings do not fully capture. A person may never publish a sensitive statement themselves, yet an agent acting on their behalf can surface information or behavioural clues in public. The owner may not even be present when the post appears.

Developers can reduce the risk by minimising retained context, isolating sensitive data, restricting public actions and testing outputs for disclosure. Users can help by avoiding unnecessary access and reviewing what an agent is authorised to remember or publish. Neither measure guarantees safety, but both reduce the amount of personal material available to escape.

The next generation of AI assistants is being designed to know users better so it can do more for them. This research shows the other half of that bargain. The more an agent knows about its owner, the more carefully its public voice has to be separated from the owner’s private life.

About the Author

Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity and the social impact of emerging technology. LiveAIWire is an independent, human-led technology publication using AI-assisted research, editorial production and original AI-assisted editorial illustrations under his direction.