Consumers who ask an AI what to buy may receive a different recommendation when they repeat the same request. A new audit of AI-generated product recommendations examined 2,528 real commercial-advice queries and 1,536 product responses from ChatGPT, Google Gemini and Google Search AI Overviews. The researchers found substantial variation across repeated outputs and major differences in the sources shown to users.
AI product recommendations can look more definite than the evidence behind them
The study focuses on a growing use of general-purpose AI: asking a chatbot to choose a phone, appliance, subscription or other product. Unlike a conventional search results page, a chatbot can collapse many possible sources into one fluent recommendation. That can feel like a considered judgement even when the underlying output is unstable.
In the audit, ChatGPT used first-person preference language in 79% of product-recommending responses, compared with 7% for Gemini and 2% for AI Overviews. The researchers also found that recommended products often changed when the same query was repeated. That does not prove any individual recommendation was wrong, but it weakens the idea that one answer necessarily represents a stable ranking.
The language matters because statements such as “I would buy” or “my pick” can imply a coherent preference. A model does not shop, own products or experience long-term satisfaction. It generates a response from the information and constraints available in that interaction. A confident personal tone can therefore make statistical output feel more settled than it is.
The source lists changed as well as the products
The researchers compared the domains displayed by different systems for the same query. ChatGPT and Gemini interfaces shared only 5.4% of source domains on average, and 76.7% of comparisons had no domain in common. The corresponding APIs also differed from consumer-facing interfaces, with mean domain overlaps of 12.0% for ChatGPT and 14.8% for Gemini.
Those figures do not show that one system had the correct sources and another did not. They show that auditors cannot assume the API represents what consumers actually see, or that two leading assistants are drawing attention to the same evidence. Interface design, retrieval layers and product-specific behaviour can all shape the answer.
LiveAIWire has already reported that AI review summaries changed stated purchase intent in a controlled experiment. A separate study found that robot retail assistants reduced early exits. The new audit addresses a different point in the buying journey: the apparent independence and repeatability of the product recommendation itself.
Advertising makes auditability more important
The paper is motivated partly by the expansion of advertising around AI products. The presence of advertising does not establish that a recommendation has been bought or biased. It does make transparency more important because commercial advice is valuable precisely when users believe the system is helping them compare options rather than quietly optimising for a seller.
Independent audits therefore need to test the interface people actually use, repeat the same query and record the sources shown at the time. A single screenshot can demonstrate what happened once, but it cannot establish a stable product policy. Repeated sampling is especially important when models and retrieval systems are updated frequently.
Consumers can apply the same principle informally. A recommendation is stronger when the reasons remain coherent across rephrased questions, the cited sources actually support the claimed advantages and the user can identify what trade-offs would change the result. If the top pick changes with no explanation, that instability is information in itself.
A shopping assistant should explain the decision boundary
Useful product advice is rarely about identifying one universally best item. The right choice depends on price, reliability, ecosystem, accessibility, repairability, size, performance and personal priorities. A system that gives a single confident answer without showing which criteria drove it can hide those trade-offs.
A better assistant would make the decision conditional. It could say which product is strongest for a particular priority, what evidence supports that conclusion and which alternative becomes preferable if the user values something else. That makes the recommendation easier to challenge and less dependent on the authority of a fluent sentence.
Source provenance should also be visible. Manufacturer specifications can establish dimensions or battery capacity, while independent testing may be better for durability or real-world performance. Retail pages reveal price and availability but have a commercial interest in conversion. Treating every source domain as equivalent would obscure those different roles.
The audit does not prove hidden manipulation
The study does not establish that advertising payments caused a particular recommendation, that one platform deliberately favoured a seller, or that AI shopping advice is generally worse than human advice. It audits observable responses and source patterns across selected systems at a point in time.
Its strongest result is therefore about reliability and inspection. If the same shopping question can produce different products and very different source sets, consumers and regulators should be cautious about treating a single answer as a transparent ranking. The recommendation is an output of a changing system, not a timeless verdict.
That does not make AI shopping advice useless. It makes it something to interrogate. The best use may be to surface options, expose trade-offs and gather evidence, while keeping the final purchase decision anchored to criteria the buyer can actually see and verify.
Repeatability should become part of consumer AI auditing
Traditional product testing assumes that a reviewer can document the version of a device and repeat a benchmark. Generative advice is harder because the same interface can answer differently across sessions. Auditors therefore need repeated prompts, timestamps and enough samples to separate normal variation from persistent preference patterns. Without that design, a strong conclusion may rest on an output that another user never sees.
The paper also shows why API audits alone are insufficient. Researchers often prefer APIs because they are easier to automate, but consumer interfaces can add retrieval, source selection, system instructions and commercial layers that are absent from the developer endpoint. An audit of the API may be technically reproducible while measuring a different product from the one shoppers actually use.
Users can ask the system to expose what would change its answer
A practical way to test a recommendation is to vary the decision criteria. Ask what becomes preferable if price matters most, if repairability matters most or if the user already owns products in a particular ecosystem. Stable reasoning should produce understandable changes rather than a sequence of unexplained winners. That does not eliminate bias, but it makes the recommendation less opaque.
Shoppers should also inspect the cited evidence rather than treating citation count as authority. Ten links can all repeat a manufacturer claim, while one independent test may provide the more relevant evidence. The new audit’s source-domain differences show why source identity matters as much as the fluency of the final answer.
As AI assistants move closer to transactions, the separation between advice and sales will become more important. A system may eventually compare, recommend and purchase in one interaction. At that point, disclosure of sponsorship, ranking criteria and commercial relationships becomes part of consumer protection rather than a cosmetic feature of the interface.
Regulators and consumer organisations may eventually need standard audit protocols for these systems. A useful protocol would record the exact interface, account state, geography, time, prompt, number of repetitions and displayed sources. That level of detail sounds technical, but without it two audits can appear to disagree simply because they tested different versions of a rapidly changing product.
The paper therefore points towards a broader principle for AI-mediated commerce: recommendation quality cannot be judged from one attractive answer. Consumers need repeatability, inspectable evidence and disclosure of commercial incentives. Platforms need evaluation that follows the real user experience rather than a simplified laboratory endpoint. The more purchasing power an assistant receives, the more important those controls become.
The same caution applies when an assistant appears decisive. A recommendation can feel more useful precisely because it removes ambiguity, but shopping decisions often involve trade-offs that cannot be collapsed into one universal winner. A buyer may value price, durability, privacy, compatibility or after-sales support differently. A trustworthy assistant should expose those trade-offs instead of presenting a transient ranking as though the product itself had an objective position.
About the Author
Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity and the social impact of emerging technology. LiveAIWire is an independent, human-led technology publication using AI-assisted research, editorial production and original AI-assisted editorial illustrations under his direction.
