Three AI chatbots sent readable conversation text to an outside analytics service during a controlled privacy test. Researchers at the University of California, Davis reported that Genspark, SeaArt and ChatOn transmitted prompts and parts of responses to Microsoft Clarity. Their April 2026 preprint examined 20 chatbot websites, using a sensitive question in fresh browser sessions. It is a snapshot of observed behaviour, not proof that those services operate identically today.
The finding raises a practical question that an accurate answer cannot resolve: who else receives information while a chatbot answers you? The company operating the model and the companies measuring activity on its website are not necessarily the same. A private-feeling exchange can therefore deserve scrutiny even when the assistant behaves helpfully and says nothing obviously alarming.
AI chatbots have a website around the conversation
Microsoft’s Clarity documentation explains that session recordings reconstruct what a visitor sees and does using page information and interaction events. They are not conventional video recordings. For a website owner, this can help explain where people get stuck, what they click and how they navigate. On a conversational website, however, the page being reconstructed may contain the exchange itself.
Consider the difference between watching somebody navigate a hotel homepage and seeing the message they type about travelling after a bereavement. Both are website interactions. Only one necessarily exposes the reason behind the visit. The sensitivity belongs to the content and context, rather than to whether a product describes its measurement system as analytics.
This is an architectural issue as much as an AI issue. A model could be carefully designed to refuse inappropriate requests while the surrounding interface still sends information elsewhere. Conversely, a website could minimise tracking while its model produces unreliable answers. Privacy and answer quality require separate questions because success on one does not establish success on the other.
What was measured, and what was not
The researchers used the prompt “pregnancy test near me” and inspected network traffic. Seventeen services shared some information with third parties; that does not mean seventeen shared conversation text. The three plaintext findings are narrower. In supported private-chat modes, the researchers observed no third-party content or identity exposure in the tested sessions. These distinctions are recorded in the study’s methods and results.
A transmitted conversation identifier is also different from a readable conversation. An identifier can distinguish one exchange from another without revealing its words. Treating every outgoing request as equivalent would blur the very distinctions that make a privacy investigation useful. The question should be what information moved, to which recipient, under which conditions.
Nor does observing a transmission establish everything that happened afterwards. Storage duration, staff access, onward use and deletion are separate matters. A responsible reading should resist turning a browser measurement into an unsupported claim that somebody sold a particular user’s secrets or that a human employee read them. Those stronger allegations require their own evidence.
What this means before you share something sensitive
For users, a useful starting point is to separate information needed for an answer from details that merely make the account more personal. A question about wording a difficult email may not require the recipient’s full name, address or employer. A request to explain a document may work with a short, carefully selected extract instead of the complete file.
This is not a promise that removing names makes a story anonymous. A rare combination of circumstances can still identify somebody. LiveAIWire’s coverage of personal information inferred from apparently anonymous text examines that different problem. Here, the additional concern is whether information reaches another service before the user has considered that possibility.
Private or temporary modes should be evaluated on their actual description. Does the mode concern the visible conversation history, personalisation, training, retention or third-party tracking? These are distinct questions. A reassuring answer about one should not be silently extended to all the others. The relevant setting is the one that addresses the particular exposure you want to avoid.
The same distinction applies afterwards. Removing an exchange from a sidebar is an action within a product’s interface. It is not, by itself, evidence about every previous transmission associated with that exchange. Readers considering that separate issue can see LiveAIWire’s explanation of what deleting an AI chat does and does not establish.
Masking requires attention to the whole interface
Microsoft’s masking guidance provides three modes. Strict masks all content; Balanced, the default, masks content classified as sensitive, including numbers and email addresses; Relaxed leaves content unmasked apart from input boxes and dropdowns. Administrators can also mask specific page elements. Microsoft states that masked content is not uploaded to Clarity.
That documentation highlights why a generic statement that sensitive information is masked needs closer examination. A sentence about a family argument need not contain an email address or a number. From a user’s perspective it may still be deeply private. The relevant test is whether the actual conversation area is protected, rather than whether familiar identifier formats receive special treatment.
An input box and a displayed message should also be considered separately when reviewing a design. The same words can move from a typing field into the conversation history after submission. An implementation review should follow the information through those states, rather than inspect one screen and assume protection follows the text everywhere.
For a hypothetical chatbot operator, a sensible test would include a fresh exchange, an older conversation, an edited message, an error screen and a private session. Those are proposed checks, not additional findings from the Davis study. They illustrate how to turn a broad privacy claim into questions that an implementation team can actually answer.
A privacy review should follow the data
A small business choosing an AI service can ask who operates each part of the product, what information each part receives and why. The model provider, hosting service, payment processor and analytics supplier may serve different purposes. Listing company names is only the beginning; the useful explanation connects each recipient to a defined category of information.
The business should also distinguish its public website from any workspace handling confidential material. Measuring how visitors find a contact page is a different decision from measuring the contents of a client discussion. Applying one convenient configuration everywhere may save administration, but convenience is not evidence that the configuration suits every type of page.
An informative supplier response would describe the relevant controls and the evidence supporting them. An unhelpful response would repeat that privacy matters without explaining the data flow. Buyers need enough detail to judge the service against their own use case, rather than a general assurance that another organisation has thought about security.
This is particularly relevant when staff use an assistant for material about other people. The person operating the keyboard may accept a trade-off for themselves, yet have less reason to expose a customer’s dispute or a colleague’s personal circumstances. A useful internal rule should identify what may be submitted and where, rather than leave each employee to improvise.
Why this is not a permanent league table
The study provides grounds for specific follow-up questions, not a permanent ranking of safe and unsafe brands. Website code, settings and suppliers can change. Independent replication would help establish whether a reported exposure persists, while a documented fix would materially alter the practical assessment. Publication date matters when a finding concerns configurable software.
Equally, a service absent from a particular finding should not be presented as comprehensively private. A test of browser transmissions cannot certify every feature, account type, mobile application or backend practice. Readers should be wary of both sweeping condemnation and sweeping reassurance when neither follows from the measured conditions.
The enduring lesson is to judge the whole route taken by a conversation. A helpful chatbot answer tells you something about the response. It tells you very little about the surrounding website. Before treating an assistant as a place for confidential discussion, the important question is whether the privacy arrangement matches the intimacy of the conversation.
For providers, that means making an answer about privacy as concrete as an answer about features. Explain which information is collected, distinguish necessary processing from optional measurement, and show how the controls apply to the conversation itself. Users should not need to infer a confidential service from a confidential-looking screen.
About the Author
Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity and the social impact of emerging technology. LiveAIWire is an independent, human-led technology publication using AI-assisted research, editorial production and original AI-assisted editorial illustrations under his direction.
