Your default sort order is a trading decision. Most sites have never made it
2026-08-24
When I opened the sort dropdown for a client's biggest category, I saw that it was set to price, low to high. This had been the case since the site went live. No one had changed it; it was the platform's default setting and had therefore been silently determining which products were sold for four years.
- Ecommerce
- Merchandising
- CRO
- Strategy
The setting nobody chose
On almost all of the sites I have checked, this situation is described as follows: when you open the category page, you select the sort dropdown and discover that it is set to whatever the platform originally had; on some versions it is ordered by price, on others it is alphabetical, and on a third it is arranged by position, this position meaning the sequence in which someone had imported a spreadsheet in 2019. No decision was made, yet a decision was still made, and it has been in effect every day since.
What makes this worth writing about is not that the setting is wrong. It is that the setting is doing far more work than the people responsible for the P&L think it is. The default sort determines which products occupy the first screen of every category and every unfiltered search. On a mobile viewport, that is four to six products. On desktop it might be twelve. Everything else in the category exists in a formal sense, in that it is present in the database and reachable by scrolling, and in a commercial sense it barely exists at all.
Trading teams spend enormous effort on which products to buy, how to price them, and what to promote. Then the sequence in which those products actually meet the customer is left to a dropdown that nobody owns. I have never once found this on a deliberate schedule of review.
Position, not merit, is what allocates demand
The behavioural fact underneath all of this is position bias. Users examine a ranked list from the top down, and the probability that a given item is examined at all declines steeply with its rank, largely independently of how relevant that item is. This is not a soft claim from a UX blog. It is the working assumption of every serious ranking system in commerce, formalised in the cascade and position-based click models and treated as a measurement problem that has to be corrected before click data can be trusted.
Look at who publishes on it. eBay has written on estimating position bias for unbiased learning to rank in ecommerce search. Walmart's search team has published models that separate position effects from genuine click and conversion propensity in sponsored results. Rakuten has published on estimating position bias where the display order is fixed, and the data is sparse. These teams are not doing this for academic amusement. They are doing it because if you train a ranking model on raw clicks, the model learns that whatever you happened to show first is good, and then shows it first again.
That last sentence is the whole commercial problem in miniature. Position bias means your click and conversion data is contaminated by your own merchandising decisions. A product that converts at 4% from row one and a product that converts at 4% from row nine are not equally good products, and no report on your dashboard will tell you which is which unless somebody has deliberately corrected for rank. Most mid-market retailers have not, which means the bestseller list they are so confident about is partly a record of what they promoted.
So the default sort is not showing customers what sells. It is deciding what will sell, and then reporting the result back as evidence that the decision was correct.
Price, low to high: four seconds of the wrong education
The most common example of an accidental default that I come across is price ascending, mainly since it's the most easily defensible in a meeting. Since customers care about price, you should show them the cheaper options first; this seems to be a customer-friendly approach. It isn't.
In Baymard's usability testing, sites using price as the default sort produced a skewed perception of the range on offer. Users looking at Crutchfield with a price default came away with a sparse and misleading impression of the selection. In the reverse case, at Mahalo, expensive multiproduct sets sat at the top, and one participant said plainly that because the most expensive items appeared first, she would probably leave. The mechanism is straightforward. People form an impression of the category from the first few products and treat that impression as representative, because it is the only sample they have been given.
From a commercial point of view, setting prices in ascending order has a particular and costly effect: it places your least valuable stock in the most visible positions on the website. In the case of most catalogues, this involves accessories, spare parts, individual items that are usually purchased in packs, and entry-level SKUs kept in order to cover the full range rather than to provide a good margin. As a result, these products take up the click share that would have gone to the product lines around which the category was actually developed. This causes the revenue per category visit to decrease, and since the average order value drops as well, the product mix shifts towards those items which contribute the least.
There is a second effect that shows up later. The price anchor set on the first screen carries into how everything else in the category is judged. If the opening row is nine pounds, the sixty-pound hero product is not evaluated against its competitors, it is evaluated against nine pounds. You have handed the customer a reference price drawn from your least representative stock.
Alphabetical, newest, and rating: order that is not information
Alphabetical is rarer but still out there, usually on smaller catalogues and B2B sites. Baymard found it at Eden's Garden. The problem is that it is arbitrary sequencing dressed up as structure. It looks organised, so nobody questions it, and it conveys no information about what the category contains beyond the initial letters of product names, which are typically a marketing decision rather than a taxonomy.
It is reasonable to prioritise the most recent items in only a limited number of cases where freshness actually determines a customer's purchase, such as in the case of fashion drops. In all other situations, the items that happened to be loaded last by the buying team appear, even though this has no connection with anything the customer cares about and, in fact, puts well-established sellers below items that have never been tested.
Sorting by user rating as a default has its own failure: a five-star average from three ratings outranks a 4.6 from four hundred. Baymard's guidance is to use rating average and rating count together, and that guidance exists because sites that do not routinely lead their category with obscure products carrying a handful of enthusiastic reviews. That is a separate discussion, and I have written about rating dispersion before, so I will leave it there.
Best selling: the loop that eats your own assortment
This is the one that deserves the most attention, because it is the most commonly considered choice and it looks like the most commercially sensible option available. Sort by what sells. What could be safer.
Start with what it does to the customer's read of the category. Baymard observed a participant at eBags describing the bestseller-sorted laptop sleeve list as though somebody had opened a box and pulled items out one at a time. The first 20 results were dominated by a single sleeve type. At Sephora, the opening products under the default bestselling sort came from a single brand. In both cases, the category held far more variety than the customer could see, and the customer's decision about whether to keep looking was made on the visible sample.
There is also the matter of structure, and this problem worsens over time. Since sales rank determines position, position determines exposure, exposure in turn determines sales, and sales then affect the rank, the study on popularity bias in recommender systems from 2024, published in User Modeling and User-Adapted Interaction by Klimashevskaia, Jannach, Elahi and Trattner, lists this phenomenon in the literature: recommender systems which are optimised on the basis of observed interaction data give more exposure to items that are already popular and suppress the long tail, the effect strengthening itself over the cycles rather than diminishing.
To a retailer, the effects are visible in three areas of the profit and loss statement. New products take much longer to become established than their real appeal would suggest, since they start off with no history and are positioned in places where people don't notice them. Sales of the items in the lower part of the range decline, and it is in this lower section that your need to offer discounts at the end of the season is located. At the same time as the actual variety of products on offer increases, the customer's perception of the range becomes narrower, causing them to turn to competitors when it comes to any products other than your bestsellers.
The uncomfortable part is that the reporting looks excellent while this happens. Your bestsellers convert brilliantly. They convert brilliantly partly because they are the only products anyone sees.
What a diversity-weighted relevance sort actually does
The solution is not to give up on the sales rank; rather, it is to put limits on it. It is worth stating exactly what Baymard proposes: maintain an order that is close to that of the best-selling products, but make it necessary for any product type making up more than 10% of the list to appear within the first 20 products on desktop and within the first 10 on mobile, no matter what its sales rank is.
As a result, the opening screen acts not as a leaderboard but as a representative sample. If a customer visits a broad shoe category page, they will see a variety of shoes such as trainers, boots, sandals and formal shoes rather than four rows of trainers only. Baymard gives the example of Home Depot's Best Match feature for refrigerators, which displays a range of sizes, finishes and price points; Costco's Best Match for the combined Toys and Books category; and REI's Best Match for running shoes across different brands, styles and types of running. The thing these examples have in common is that the customer is able to tell what the category is about on the first screen.
Why this converts is the part worth being precise about. Customers do not audit your catalogue. They take the first few products as a sample, infer the population from it, and decide whether the category is worth their attention. When the sample is unrepresentative, they conclude the category does not contain what they want and either apply an overly restrictive filter or leave. Baymard recorded exactly this at Macy's, where a user in a broad Jackets and Blazers category saw mainly blazers, decided she needed a jacket style filter to find jackets, and filtered away most of the available jackets in the process.
That failure costs you a customer who was in the right place with the right intent. It is the most expensive kind of loss, because you already paid to acquire the session, and the product they wanted was thirty items down.
The more advanced approach substitutes the single 10% rule with a matrix, using product subtype as the main factor and then making adjustments to the order based on sales rank, rating, pageviews, and the relationships between products. This is a suitable application for machine learning and for real multivariate testing, and the weights will vary from category to category. However, the simpler version provides most of the benefits and can be implemented within a sprint, which is important if you are aiming to get it on the roadmap at all.
Name it honestly, and set it by category
Two smaller points that are easy to get wrong once you have built the thing.
The first is naming. Baymard's testers equated "Popular" with "best selling", so a diversity-weighted sort labelled Popular sets an expectation the sort does not meet, and users notice when the first results are not obviously the most bought. "Relevance" and "Featured" both work better, because they carry a useful ambiguity: they signal that the retailer has had some say in the ordering without making a specific claim you then break. That is not evasion, it is accurate labelling of a curated order.
The second is that the right default is not the same across your site. A seasonal apparel category needs a default that responds to the season, and Baymard found users disappointed by lightweight blazers dominating a jacket category during cool weather. A spares category probably should default to something close to price or part number, because the buying task is genuinely different. Treating the default sort as one global setting is how you end up optimising for the average category, which does not exist.
What to measure, and how not to fool yourself
If you do this, you should use revenue per category visit, broken down by category and by device, compared with a suitable holdout sample. The conversion rate on its own will mislead you since a diversity sort intentionally displays some products that have a lower conversion rate and achieves its return by reducing the number of category abandonments and getting a wider variety of products purchased.
A measurement that is worth establishing is sell-through according to rank decile. In the case where the first decile is the one doing all the work and the fifth decile is remaining unchanged, this is not an indication of demand but rather an exposure artefact, and it is this figure that shows you whether the change has any structural effect. The third is the proportion of category sessions in which the customer changes the sort themselves; a high rate is a complaint about your default, expressed in behaviour, not in words.
Where the evidence is firm and where it is not
I want to be straight about the sourcing here, because two different kinds of evidence are doing work in this piece.
Position bias and popularity reinforcement are on solid ground. The click model literature is mature, the correction methods are in production at eBay, Walmart and Rakuten, and the 2024 popularity bias survey is a current review rather than an isolated finding. The mechanism is not in dispute.
The qualitative findings regarding default sort types put forward by Baymard, together with the statistic that 24% of the desktop sites in their benchmark do not display category breadth in their default sort, are from an article published in May 2021 and are based on their assessment of the top-grossing US websites. This finding is still included in their present guideline suite, the examples of the websites in question are named and can be verified, but since a 2021 benchmark figure represents a particular sample at a particular time, I would not regard it as a current measure of the market in 2026. The qualitative observations concerning how users infer breadth from the first screen are the more lasting element; the percentage is relative to the circumstances.
What I have not found, and would genuinely like to see, is a large published test of diversity-weighted defaults against bestseller defaults with revenue per visit as the outcome. Retailers who run these tests do not publish them. If you have run one, I would rather hear the result than keep inferring it.