Revenue per visitor: how to judge Shopify A/B tests

Revenue per visitor combines conversion rate and order value in one number. How to calculate it, and how to use it to judge Shopify A/B tests.

Portrait of Chirag Babu

Written by

Chirag Babu

Co-founder & CEO at Cascayd

7 min read

Revenue per visitor (RPV) is your store's revenue divided by every visitor, buyers and non-buyers alike, which works out to conversion rate × average order value. It is the best single metric for judging a Shopify A/B test, because a change cannot look good by trading conversion rate against order value. It is also noisier than conversion rate, so RPV tests need more traffic, a cap on extreme orders, and guardrails for returns and margin.

Ask most store owners how their site is doing and they will quote a conversion rate. That number tells you how many visitors placed an order. It says nothing about what those orders were worth, and that gap is where a lot of A/B tests go wrong.

What revenue per visitor is

Revenue per visitor is the revenue your store earns divided by the number of visitors who arrived, including everyone who left without buying.

Revenue per visitor = revenue ÷ visitors

Revenue is orders multiplied by average order value, so RPV can also be written using two numbers most stores already track:

Revenue per visitor = conversion rate × average order value

Take a store with 50,000 visitors in a month. 2.5% of them buy, which is 1,250 orders, at an average order value of $80. Revenue is $100,000, so revenue per visitor is $2.00. The second formula gives the same answer: 0.025 × $80 = $2.00.

That $2.00 is what an average visitor is worth to the store. Every change you make to the site either raises it or lowers it.

Be consistent about what counts as a visitor and what counts as revenue. Shopify, for example, defines average order value as gross sales minus discounts, divided by orders, and leaves out post-order adjustments such as returns [Shopify]. Any sensible definition works, as long as both sides of a test use the same one.

Why conversion rate alone misleads

Conversion rate and average order value often move in opposite directions. A test that watches only one of them can call a money-losing change a winner.

Here are two common experiments on the same store, starting from a 2.5% conversion rate, an $80 average order value, and $2.00 per visitor.

A 10% off banner. Conversion rate rises to 2.8%, a 12% relative lift that looks like a clear win. But everyone now pays 10% less, and shoppers who needed a discount to buy tend to spend less, so average order value falls to $70. Revenue per visitor is 0.028 × $70 = $1.96. The store earns less per visitor than before.

A higher free shipping threshold. Raising the threshold nudges shoppers to add one more item. A few abandon, so conversion rate slips to 2.4%, and a conversion-rate test would reject the change. Meanwhile average order value climbs to $90, and revenue per visitor becomes 0.024 × $90 = $2.16, an 8% lift.

These numbers are illustrative. Judging by conversion rate alone would ship the first change and kill the second, and revenue per visitor gets both right. That is why an experiment should be judged against a single overall evaluation criterion that reflects what the business actually wants [Kohavi et al., 2020].

Why revenue per visitor is harder to measure

The better metric is also the noisier one, and that changes how you run a test.

Conversion rate is a yes or no for each visitor. Revenue per visitor is a spread of values. Most visitors contribute $0, most buyers contribute something near your average order, and a few contribute a great deal. One wholesale-sized order can add more to a group's revenue than a hundred ordinary ones. Statisticians call this a skewed distribution, and skewed metrics need more visitors to reach the same confidence as a simple conversion rate [Kohavi et al., 2014].

That has three practical consequences.

Plan for a larger sample. A test judged on revenue per visitor usually runs longer than the same test judged on conversion rate. Treat the sample size calculator result as a minimum.

Cap extreme orders. A single huge order can decide a test on its own. Cap each visitor's revenue at a high threshold, set before the test starts, so one outlier cannot swing the result. The same research recommends capping skewed metrics for this reason [Kohavi et al., 2014].

Use what you know about returning customers. Returning customers bring a spending history into the test. A technique called CUPED uses data from before the experiment to strip out part of the natural variation between shoppers, which makes a test more sensitive at the same traffic, or equally sensitive with less [Deng et al., 2013]. It helps most on stores with a large base of repeat buyers.

From a lift to incremental revenue

A lift in revenue per visitor converts straight into money.

Incremental revenue = (variant RPV − control RPV) × visitors × period

Back to the example store. If a winning test lifts revenue per visitor from $2.00 to $2.20, a 10% lift, then across 50,000 monthly visitors that is an extra $10,000 a month, or $120,000 a year, with no extra ad spend. A store with $1.5M in annual revenue that ships a 30% lift adds $450,000 a year.

Run your own numbers in the A/B testing ROI calculator before a test starts. It tells you how big a lift has to be before the work pays for itself.

Keep two cautions in mind. A lift measured in a test is an estimate with a margin of error around it, so keep watching revenue per visitor after the change goes live. And the comparison that counts is variant against control over the same weeks. Comparing this month's revenue per visitor with last month's mixes in seasons, promotions, and traffic sources, which is exactly what an A/B test removes.

Revenue is not profit

Revenue per visitor decides the test. A few guardrails make sure the win is real.

  • Returns. A change that encourages impulse buying can lift revenue and returns together. Returns arrive weeks after the order, so check net sales, which subtract them, once they have had time to come in. Shopify describes net sales as the best approximation of actual revenue for this reason [Shopify].
  • Discount cost. A variant that wins on revenue by giving away margin can still lose on profit. Price the discount in before you declare a winner.
  • Product mix. If the variant shifts shoppers toward lower-margin products, revenue can rise while profit falls. Check what people bought as well as how much they spent.

These sit alongside revenue per visitor as checks. They stop a real revenue win from hiding a real cost.

How a win gets verified

Cascayd only gets paid when an experiment makes a store money, so every result is held to the same standard before it counts as a win.

  • A 50/50 split against the live store. The variant and control run over the same period, so seasons, promotions, and traffic sources affect both sides equally.
  • Full weekly business cycles. Tests run in whole weeks so weekday and weekend shoppers are both represented, and nobody stops a test early because a chart looks good.
  • 90% confidence or higher. If the change truly made no difference, a result this strong would appear by chance less than one time in ten.
  • Incremental revenue only. The win is the difference in revenue per visitor between variant and control, multiplied by the visitors it applies to. A flat or negative test counts as zero.

The Shopify A/B testing guide covers the rest of the process, from choosing what to test to protecting the split. To find where your own store's revenue per visitor is leaking, start with a free audit.