AI & Society

Would You Buy It? AI Review Summaries Changed the Answer

Guardian-style illustration of a shopper considering a smart speaker as a small AI presents a review summary beside a swirl of conflicting reviews.
AI-generated review summaries can change whether people decide to buy a product, even when the underlying customer feedback is mixed.

AI review summaries can look like a neutral convenience, but a controlled study found they changed what people said they would buy. Participants chose a manufacturer in 83.7 per cent of cases when they read positively reframed summaries, compared with 52.3 per cent when they read the original neutral or negative reviews. They were also willing to pay 4.5 per cent more. The experiment measured stated choices rather than real checkout behaviour, but it showed that compression can quietly become persuasion.

The finding comes from a peer-reviewed paper presented at IJCNLP-AACL 2025. Its authors examined how large language models alter source material through framing, misplaced emphasis and hallucination. The shopping experiment was deliberately built around unusually strong framing changes, so the result should not be treated as the average effect of every summary on every product page. It is better understood as a demonstration of what can happen when an apparently helpful synopsis removes the criticism that made the original review balanced.

How the AI Review Summaries Experiment Worked

The researchers began with 1,000 reviews from the Amazon Reviews 2023 dataset covering household, electronics and kitchen products. GPT-3.5 generated summaries. GPT-4 was then used to identify product types with the most extreme shifts from neutral or negative source text to positive summary, after which the researchers manually selected ten product pairs for the human study.

Participants saw two manufacturers for each kind of product and chose between them. One item was described either by its original review or by a positive AI summary of that review, while the comparison item retained neutral framing. Each participant saw a mixture of formats. The study recruited 72 English-speaking adults through Prolific and accepted 70 submissions after a comprehension check.

The published paper reports a statistically significant difference between conditions. The manufacturer backed by the positive summary was selected 83.7 per cent of the time, against 52.3 per cent when people saw the original review. The paper describes that as a 32 percentage point increase. It also found a 4.5 per cent rise in stated willingness to pay relative to the comparison product.

The incentives in the experiment also deserve attention. Participants received a small performance bonus for choosing the option they believed matched majority opinion. That encouraged a considered prediction about typical consumer preference, but it was still not their own money and the products were not delivered. The measured effect is therefore evidence about decision framing under controlled conditions, not a sales forecast for an online shop.

The researchers selected cases with the largest framing changes because they wanted to test a worst-case mechanism. That choice makes the causal contrast easier to see, while reducing the basis for generalising its size. A routine summary that preserves the leading criticism may have little effect. A summary chosen or tuned because it converts better could move decisions much more.

The Summary Did Not Invent a New Product

The mechanism was not a fabricated specification or a false discount. It was selective emphasis. A source review might say a vacuum cleaner picked up dirt well but had weak battery life, flimsy attachments and doubtful value. A summary could preserve the cleaning performance and attachments while dropping the warnings. Every retained sentence might sound defensible on its own, yet the overall impression would become more positive than the source.

That distinction matters because many systems and users treat summarisation as extraction rather than authorship. If a chatbot supplies a recommendation, people know they are reading a judgement. If it supplies a summary, they may assume the judgement has already been made by the reviewers and merely shortened by the model. The model gains persuasive power precisely because it appears not to be persuading.

LiveAIWire’s reporting on the recommendation algorithm shaping Hollywood documents a related transfer of influence. An interface can alter what people choose without issuing an explicit command. Ordering, omission and emphasis are enough. Review summaries bring the same dynamic to the point where a consumer is deciding whether a product is worth the money.

What the Wider Study Found About Model Bias

The shopping test was one part of a larger investigation involving five model families and several summarisation and fact-checking tasks. Across the tested models, summaries changed the source text’s sentiment in 26.42 per cent of cases. They also overemphasised information from the beginning of the source in 10.12 per cent of cases. Those aggregate figures came from automated comparisons across datasets rather than from the seventy-person consumer experiment.

The paper separately reported hallucination on 60.33 per cent of post-knowledge-cut-off fact-checking questions. That number concerns a different task with a deliberately self-updating news dataset. It should not be merged with the shopping result or used to claim that six in ten product summaries are false. The common thread is that model output can depart from the source in more than one way, but the denominators and consequences are different.

The ACL Anthology record identifies the work as a December 2025 conference paper and provides the authors, DOI and authoritative PDF. The researchers also tested eighteen mitigation methods. Results varied by model and bias type, which is a warning against assuming one prompt instruction can make summaries neutral across every system.

Why AI Review Summaries Matter More as Shopping Becomes Agentic

A biased summary is consequential when a person reads it. It becomes more consequential when the same system also filters the products, compares the prices and completes the purchase. LiveAIWire’s analysis of agentic shopping shows that this transition is already under way. The boundary between describing a product and acting on the description is shrinking.

Once an AI agent handles the whole path, a user may never see the reviews that were omitted, the competing products that were ranked lower or the uncertainty hidden by a concise answer. The commercial incentive is obvious. A retailer wants less friction and more completed transactions. A platform wants the summary to feel decisive. Yet a system optimised for conversion may produce exactly the positive compression the study found capable of changing purchase intent.

Accountability also becomes harder. If the model reframes a review, the agent selects the product and the merchant accepts the order, several companies may have contributed to the final choice. LiveAIWire’s guide to AI agent liability explains why authority, logs and confirmation controls will matter when an automated recommendation turns into an expensive action.

How to Read an AI Product Summary Without Being Led by It

The most useful defence is comparison with the source at the moment the decision matters. For an expensive, safety-critical or difficult-to-return purchase, open several full reviews and look specifically for repeated complaints, not just the features the summary chooses to mention. A good interface should let the reader move from every summary claim to the underlying passages that support it.

It also helps to ask a model for the strongest reasons not to buy the product and to separate facts from reviewer sentiment. That does not guarantee neutrality, because the same model is still choosing what counts as strong. It does, however, create a second pass that can expose a missing battery complaint, durability warning or incompatibility that a smooth positive summary discarded.

Retailers and platform designers have a higher obligation than individual shoppers. Summaries should identify the review set and date range, preserve material negative findings, disclose uncertainty and provide source access. Systems should be tested for changes in purchase intent, not merely for whether each sentence can be traced to a review. A faithful-sounding synopsis that systematically drives one commercial outcome is not neutral in practice.

The study does not show that every AI summary manipulates buyers or that the measured selections became real purchases. It does prove the answer to a buying question can change before the product, price and source reviews change at all. Only the summary changed. For consumers and regulators, that is enough to make the summary itself part of the decision system rather than a harmless layer on top.

The safest summary is therefore not the smoothest one. It is the one that preserves the reasons a reasonable buyer might say no, makes its source set visible and gives the reader an easy route back to the evidence before money changes hands.

About the Author

Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity and the social impact of emerging technology. LiveAIWire is an independent, human-led technology publication using AI-assisted research, editorial production and original AI-assisted editorial illustrations under his direction.