A re­search team from Pader­born Uni­ver­sity and LMU Mu­nich has pub­lished a study on the ag­greg­a­tion be­ha­viour of cus­tom­er re­views in the sci­entif­ic journ­al PLOS ONE

 |  Heinz Nixdorf InstituteBehavioral Economic Engineering and Responsible Management / Heinz Nixdorf Institute

Whether on Amazon, Booking.com or Google Reviews, star ratings are now among the most important factors influencing online shopping decisions. Platforms summarise the reviews of thousands of customers into a single metric, usually the arithmetic mean. But does this calculation actually reflect the way in which consumers perceive and process rating distributions? A new study by an interdisziplinär research team from Paderborn University and Ludwig Maximilian University of Munich now provides experimental answers to this question for the first time.

Research gap: How do customers aggregate ratings?
Previous research has primarily focused on whether the arithmetic mean is a useful indicator for purchasing decisions. By contrast, the aggregation principles that consumers actually apply when looking at rating distributions have remained largely unexplored. Do they really use the average, or do they weight certain rating categories more heavily than others, for example by focusing on extremely negative or positive ratings?

The experiment: sorting products based on ratings according to personal preferences
To answer this question, the researchers conducted a controlled laboratory experiment with 107 participants. The participants were given only the rating distributions for three products at a time and had to rank them according to their personal preferences without knowing the product names or any other characteristics. An incentive-compatible design ensured that the participants revealed their true preferences. The higher they ranked a product, the more likely they were to receive it. The ranking decisions obtained in this way were then analysed using various statistical Plackett-Luce models and compared with various theoretically derived aggregation functions.

The results: It is not only the arithmetic mean that is relevant to customers
The majority of consumers do indeed aggregate rating distributions in line with the arithmetic mean. However, more than 40 per cent of participants systematically deviate from this. A significant proportion follow a binary strategy, distinguishing only between positive (4–5 stars) and negative (1–2 stars) ratings, whilst the middle category (3 stars) receives little attention. Other groups focus predominantly on negative ratings, particularly 1-star ratings. It is also noteworthy that none of the participants used the median as an aggregation principle, even though this is considered particularly robust in the theoretical literature. These patterns proved to be stable. They persisted regardless of whether participants were provided with additional numerical information, and were not significantly influenced by individual characteristics such as age, gender or online shopping experience.

What does this mean for platforms and consumers?
The findings suggest that the common practice of summarising product reviews solely via the arithmetic mean does not provide the optimal basis for decision-making for a significant proportion of consumers. Platform operators could benefit from offering customisable aggregation options, just as specialised travel portals already allow users to filter reviews by type of trip. For example, consumers could choose whether they prefer a traditional average rating, a metric focusing on negative experiences, or a simplified positive-negative breakdown. As all the necessary rating data is already available, such adjustments would be technically feasible and could lead to more satisfied customers and better purchasing decisions.

Publication:
The study “Aggregation Processes in Customer Rating Systems — Insights from an Economic Decision Experiment” by Dirk van Straaten, Behnud Mir Djawadi, Vitalik Melnikov, Eyke Hüllermeier and René Fahr has been published in the specialist journal PLOS ONE (https://doi.org/10.1371/journal.pone.0343851).

This publication was funded by Paderborn University’s Open Access Publication Fund, as well as by the DFG as part of the Collaborative Research Centre (CRC) 901 On-the-Fly Computing.