Why a perfect rating sells less, and what your review block is actually doing to conversion

2026-08-10

A five-star average sells worse than a 4.5. Negative information has greater diagnostic value since it separates good products from bad ones more effectively than praise does, simply because praise is something that is expected by everyone.

  • Ecommerce
  • CRO
  • PDP

The quarter the average went up and the revenue went down

A few years ago I received a catalogue which had a review problem that, in theory, appeared to be a review success. The previous team had tightened up the moderation process so that reviews which were off-topic, abusive, or clearly concerned with delivery rather than the product were rejected, as well as a reasonable number that were just unhelpful. That was sensible housekeeping and, over a period of two quarters, the average rating in the top-selling category rose from around 4.3 to around 4.8.

Conversion rate on those product pages fell. Not dramatically, but consistently, and against a category where traffic mix, price and stock availability had not moved in any way that explained it. We spent a while looking for the usual suspects before anyone suggested that the thing we had improved was the thing that had broken.

The true noise as well as the genuine noise had been taken away, together with the visible sign that the reviews had not been curated. The block had turned into a wall of praise, and a wall of praise is not convincing since that is precisely what a company would do if it were constructing one.

The shape of the curve, and how old the evidence is

The best-known number here comes from the Spiegel Research Center at Northwestern, working with PowerReviews data. Their finding, first published in 2015 and expanded in the 2017 report, is that purchase likelihood typically peaks for products rated between 4.0 and 4.7 stars and then begins to decline as the average approaches 5.0. The same research found that showing five reviews on a page rather than none lifted purchase likelihood by 270%, with the effect stronger on expensive, high-consideration items than on cheap ones.

I should be clear regarding the date in question. Twenty-fifteen is a long way back in this area, and it wouldn't be reasonable to base an argument for 2026 solely on that. The number of review volumes is now one order of magnitude greater, enforcement of fake reviews has changed, and the platforms have introduced summary features which were not there before. Therefore, the fair question is whether the more recent work supports the curve or contradicts it.

On the descriptive side, it holds up. PowerReviews consumer research continues to find that a substantial minority of shoppers are actively suspicious of a flawless five-star average, and the pattern is stable enough that it now appears in local search benchmarks as well as ecommerce ones. What newer work has done is not overturn the curve. It has explained the mechanism better, and in doing so it has made the operational advice more specific and less comfortable.

Why the critical review does the work

The main idea is diagnosticity, a concept that originated in the study of person perception and was introduced into marketing many years ago. Information is said to be diagnostic when it enables you to divide things into categories. Praise is only weakly diagnostic since almost anything is praised; therefore, the fact that a product has admirers tells you very little about whether it should be placed in the good category. On the other hand, criticism is strongly diagnostic since it is rarer and more specific, so one reliable complaint has a greater effect on your assessment than one compliment does.

That is the reason why negativity bias appears so consistently in research into reviews. When the overall tone is positive, consumers tend to rate negative reviews as more helpful than positive ones, while the effect reverses when the overall tone is negative. To put it simply, helpfulness is linked to what is lacking in the group. A one-star review on a page that contains only five-star ratings is seen as information, whereas a one-star review on a page consisting of only two-star ratings is regarded as noise.

Another aspect of the situation involves people's way of building up their confidence rather than the way they handle evidence. When a shopper reads three negative reviews concerning a fit that tends to be small, decides to buy one size larger and then goes ahead and purchases the item, they have resolved their own concern. They had not accepted any reassurance offered by the brand; instead, they came to their own conclusion based on information that the brand could not control. This kind of conclusion is much more lasting than any statement made on the page, which is also why review sections that highlight objections tend to result in lower return rates as well as higher conversion rates. The customer has already taken the flaw into account before arriving at the decision.

That is also the reason why the curve drops above 4.7 instead of level off; the shopper is not seeking a poorer product. They want the material that enables them to carry out this task, and when the page takes it away, they either switch to another option or buy without having first adjusted their expectations. In both cases, the result is undesirable.

The 2026 correction: helpful review, less credible reviewer

The paper by Junha Kim and Joseph Goodman, which is presented in a helpful manner, was published in the Journal of Consumer Research on 31 July this year. In nineteen studies involving 7,330 participants, it was found that consumers regard negative reviewers as less credible than positive ones; the explanation put forward is expectancy violation, since there is a social norm that reviews should be positive, negativity breaches that norm, and as a result the person who does so suffers a loss of credibility.

Study one in that paper is the important one for anyone running a product page. A negative review is judged more helpful, and the reviewer who wrote it is judged less credible. Those are two different judgements about the same piece of text, and they point in opposite directions. The paper also shows the effect is mitigated when prior expectations are ambiguous and reverses when prior expectations are already negative.

When you look at it in light of the Spiegel curve, the picture no longer just amounts to a simple request to tolerate negative reviews. Although the critical review gains credibility since it is diagnostic, the individual giving the review is discounted, and the extent of that discounting depends on what the shopper expected before they went there. If a brand has been advertising flawlessness for five years, it has raised the bar for expectations and therefore makes its own critics seem less believable. On the other hand, a brand with a more honest positioning has vague expectations, and so its critics are listened to.

The commercial reading is that your advertising register and your review block are coupled in a way most organisations never look at, because one sits with brand marketing and the other sits with ecommerce. If the campaign promises perfection, the review block has to work harder to be believed, and the cost of that shows up on the product page rather than in the campaign report.

Dispersion, and what the AI summary layer is doing to it

The newest variable is the AI-generated summary that now sits at the top of the review block on Amazon and a growing number of other retailers. It compresses hundreds of reviews into a few sentences, and the obvious worry is that it destroys exactly the thing the block was good for: letting the shopper see the spread.

The more interesting result is the one based on empirical evidence. The study by Wang and Wang, which was presented at HICSS in 2025, applied a difference-in-differences approach to 48,019 panel observations relating to 3,781 Amazon products from February 2023 to February 2024. Overall, the use of AI to summarise reviews increased purchasing behaviour. More importantly, rating dispersion had a positive moderating effect: summaries were more helpful when ratings were spread out. The effect was moderated by review count in a U-shaped manner, and search goods saw a greater benefit than experience goods.

That makes sense once you stop thinking of the summary as a replacement for the reviews and start thinking of it as a navigation layer. Where opinion is divided, reading the reviews is expensive, and the summary saves the shopper that cost. Where everyone agrees, the star average already told them everything and the summary adds little.

The problem is inconsistency. A study carried out in Information and Management in 2025 looks at the situation in which the AI summary and the overall rating go in opposite directions. The result is that the discrepancy harms consumer attitudes, and both trust in the platform and the availability of an explanation of how the summary was created reduce the extent of this harm. This point should concern a trading team since you don't have control over the summariser on a marketplace and you may not have control over it on your own website as well. In the case where the algorithm highlights a durability complaint for a product that has a rating of 4.6, you have issued two conflicting messages about your own product, and the customer then has to decide which one is lying.

What this changes about how the block is run

The first thing to do is to cease seeing the average rating as a target; it is in fact an outcome, and when it exceeds about 4.7 it becomes an outcome which costs you money. The moderation policy should be aimed at deleting reviews that are fraudulent, abusive or concern something other than the product, and it should not delete reviews which are just unflattering; since your rejection rate has been rising while your average rating has also been rising, these two facts are connected.

The other option is to make the dispersion obvious rather than hiding it; the use of a rating histogram, the choice to sort items by default in a way that does not merely highlight the most positive content, and the provision of a truly accessible means of reaching the one-star and two-star reviews all achieve the same result. You are not acting nobly; what you are doing is giving shoppers the material they need to convince themselves, and you are providing it on your own page rather than directing them to a search engine where they would have to look for it in a place that you can't control.

The third is to check the summary against the average on your own top hundred products, by hand, once a quarter. It takes an afternoon. Where the summary contradicts the stars, you have a data problem or a product problem, and either way you would rather know.

The fourth is the one nobody wants: look at what the brand is promising. If the marketing register is absolute and the reviews are human, the reviews will be discounted, and you will be paying for that on the product detail page every day.

What to measure

You shouldn't rely solely on conversion rate here, since an intervention which increases conversion may also increase returns if it succeeds in getting people past their own doubts rather than helping them to resolve those doubts. Revenue per visitor on the product page is a better main indicator and should be looked at together with the return rate for the same products in the next ninety days.

The interaction with the block should be treated on its own: the fraction of product page sessions which scroll until the reviews appear, open the histogram, change the sort order, and open a one-or two-star review are cheap events to trigger, and they show whether the block is being used as a decision-making tool or is being ignored merely as decoration.

The comparison worth carrying out is not having more reviews per fewer products; rather, it is the same products being shown with dispersion visible as opposed to dispersion being hidden, with the products being held over the full return cycle. It is this test that distinguishes a review programme from a review widget.