Shopify A/B testing: how to run tests that prove revenue

How to A/B test a Shopify store: what you can test, how much traffic you need, which metric to judge, and how to avoid false wins.

Portrait of Rishi Babu

Written by

Rishi Babu

Lead marketer at Cascayd

9 min read

To A/B test a Shopify store, split visitors at random between your current theme and a changed version, judge the result by revenue per visitor, and run the test for full weeks until it reaches a sample size you set in advance. Product pages, collections, the cart, and offers are all testable on any plan. Checkout changes need Shopify Plus. Most false wins come from stopping early, a broken traffic split, or a promotion landing mid-test.

Most stores still change their site on instinct. A new hero image goes up, a free shipping banner comes down, and a week later someone checks whether sales moved. Sales always move, because traffic, promotions, and seasons move them. That week of data says nothing about whether the change itself made money.

What an A/B test on Shopify actually is

An A/B test randomly splits visitors into two groups. One group sees your current store (the control) and the other sees a changed version (the variant). Both groups shop on the same days, under the same promotions, from the same traffic sources. The only systematic difference between them is your change, so a reliable difference in results can be credited to it.

A before-and-after comparison cannot do that. Researchers who have run tens of thousands of these tests at large online businesses describe randomized controlled experiments as the most reliable way to show that a change caused an outcome, instead of merely coinciding with it [Kohavi et al., 2020].

Three conditions have to hold for the result to mean anything:

  • Random assignment. Visitors land in a group by chance, never by device, source, or time of day.
  • Sticky assignment. A returning visitor keeps seeing the same version, so nobody experiences both.
  • One decision metric, chosen in advance. You know before launch which number decides the winner.

What you can test on Shopify, and what is harder

Almost everything a shopper sees before checkout lives in your theme, which makes it testable. The areas worth testing first are usually:

  • Product pages. Image order, the first lines of the description, how variants are picked, where reviews sit, and delivery and returns details near the add to cart button.
  • Collection pages. Default sort order, filters, products per row on mobile, and whether price and ratings show on the card.
  • Cart and cart drawer. Upsells, free shipping progress bars, trust messages, and drawer versus full cart page.
  • Homepage and navigation. Where visitors land first, and how quickly they reach a product they want.
  • Offers. Free shipping thresholds, bundle pricing, and how a discount is presented.

Checkout works differently. Shopify controls it, and your access depends on your plan. Apps that customize the information, shipping, and payment steps are available only on Shopify Plus, while some customizations of the thank you and order status pages work on Basic Shopify and above [Shopify]. If you are not on Plus, put your testing effort into everything before checkout.

Price tests need extra care. Showing two prices for the same product at once can confuse returning shoppers and flood your support inbox. A safer option is to test how an offer is framed, such as a bundle or a free shipping threshold, and leave the base price alone.

Pick the metric before the variant

The metric decides what "winning" means, so choose it first.

Conversion rate is the obvious candidate, but it only counts whether an order happened. It ignores how much the order was worth, so a change that nudges shoppers toward cheaper products can lift conversion rate and still bring in less money. Average order value has the opposite blind spot. It ignores every visitor who never bought.

For most Shopify tests, the best single number is revenue per visitor: total revenue divided by every visitor in the group, buyers and non-buyers alike. It counts both how many people bought and how much they spent, so a change cannot look good by trading one for the other. The revenue per visitor guide explains how it works and why it needs more traffic than conversion rate.

Pick one primary metric, then a short list of guardrails that must not get worse, such as refund rate or page load time. Judge a finished test by whichever metric happens to look best and you will keep finding winners that are not real.

How much traffic you need

Four inputs decide how many visitors a test needs:

  1. Your baseline. Your current conversion rate or revenue per visitor.
  2. The smallest lift worth detecting. A 5% lift needs far more traffic than a 20% lift.
  3. Confidence. How sure you want to be that a winner is real.
  4. Power. How likely the test is to catch a real lift of the size you care about.

Small lifts on low baselines need a lot of traffic. A store converting at 2% that wants to detect a 10% relative lift needs roughly 80,000 visitors per group at 95% confidence and 80% power. The sample size calculator runs these numbers for your store and shows how many days that takes at your traffic.

If the answer is months, change the plan. Lower-traffic stores get better results by:

  • Testing bigger changes. A redesigned product page layout is far more likely to produce a detectable lift than a new button color.
  • Testing where traffic concentrates. Your top product pages and your cart see far more visitors than any single collection.
  • Using a closer metric to learn. Add to cart rate responds faster than revenue, though revenue should still decide what ships.

Run full weeks and do not peek

Shoppers behave differently on a Tuesday lunch break than on a Sunday evening. A test that runs Monday to Thursday only sees part of your customer base. Run tests in whole weeks, and for at least one to two full weekly cycles even when the sample arrives sooner.

The easiest way to fool yourself is peeking: checking results every day and stopping the moment the variant looks like a winner. Standard significance tests assume you look once, at a sample size set in advance. Checking repeatedly and stopping at the first good-looking result makes false positives far more likely than the stated confidence level suggests [Johari et al., 2017]. Early numbers also swing harder, because a handful of large orders can dominate a small sample.

Promotions distort tests too. A sitewide sale, a big influencer post, or Black Friday week brings in shoppers who behave unlike your usual traffic. Both groups get them equally, so the test stays fair, but the result may not hold once the promotion ends. Avoid starting or ending a test around a major campaign where you can, and log any promotions that ran during it.

Protect the split

A test is only as good as its randomization. Check the plumbing before you trust any result.

Sample ratio mismatch. A 50/50 split should end up close to 50/50. When it does not, something is sending visitors to one side more often, or failing to record them. A study of experiments across several software companies found that a sample ratio mismatch usually invalidates a test entirely, whatever the headline numbers say [Fabijan et al., 2019]. Check the split before you look at revenue.

Bots and measurement changes. Bots inflate visitor counts and dilute every metric. Shopify filters identified bot sessions out of its session reports by default, and it changed how sessions are counted in a rollout from September 21 to 23, 2026, which can shift conversion rate even when orders stay the same [Shopify]. Never compare a test period against a baseline measured under the old definitions.

Flicker and speed. If the variant loads a moment after the original page, shoppers briefly see the control before it switches. That flicker, plus any extra load time, becomes part of the variant and can sink a good idea. Check every variant on a real phone over a mobile connection before launch.

Discount codes and currencies. Make sure a code that works in one version works in the other, and that stores selling in several currencies convert revenue the same way for both groups.

Read the result and decide

Once the test reaches its planned sample and has run full weeks, read it once, against the decision rule you set before launch. There are three honest outcomes:

  • The variant wins. The lift clears your confidence threshold and no guardrail got worse. Ship it, and keep watching the metric after rollout.
  • The variant loses. You avoided shipping a change that would have cost money.
  • It is flat. The change did not move revenue by as much as the test could detect. Keep the simpler version, record what you learned, and try a bolder idea.

Flat is the most common outcome. Experience from thousands of experiments shows that changes rarely produce a big positive impact on key metrics, which is exactly why guessing is expensive [Kohavi et al., 2014].

To see what a winner is worth, turn the lift into money. A 5% revenue lift on a store doing $2M a year is $100,000 a year. The A/B testing ROI calculator does the maths for your own numbers.

Where to start

Start where traffic is high and shoppers drop out. For most stores that means top product pages, the cart, and the mobile experience. Look at where shoppers leave in your funnel report, watch a few real sessions, and read your support inbox for the questions people ask before buying. Each of those points to a hypothesis worth testing.

The CRO audit framework gives you a structured way to find those leaks, and the A/B testing best practices guide covers writing hypotheses a test can answer. To skip the legwork, Cascayd's free audit names your top three conversion leaks and the first experiments it would run on your store.