Average Order Value A/B Test Calculator

Model AOV uplift, simulate an A/B test on revenue per visitor, and size your sample before you run it. Built for Shopify stores.

Average Order Value A/B test calculator for Shopify stores

Disclaimer: Results are estimates based on your inputs. Sample size calculations assume normally distributed AOV. Always confirm statistical significance before shipping a winning variant.
Current AOV
Monthly revenue
Gross profit / order
Orders / day

8%
New AOV
Added monthly revenue
Added gross profit / mo

Most Shopify operators decide what to test the same way they decide what to cook for dinner: by feel. The problem is that average order value moves money differently than conversion rate does, and a test that looks like a winner on one metric can quietly lose on another. This calculator exists so you can answer three questions before you spend a single session of traffic: how much would a higher AOV actually add, would a specific change beat your current setup, and how long would the test take to give you a trustworthy answer.

It has three connected tabs. The AOV calculator turns your current revenue and orders into a baseline and projects an uplift scenario. The A/B test simulator pits a control against a variant and decides the winner on revenue per visitor, not order value alone. The sample size tab tells you how many visitors each variant needs before the result means anything. Work through them in order and you move from a rough idea to a sized, fundable experiment.


How to use the calculator

AOV calculator

Enter your total revenue, number of orders, the time period those cover, and your average margin. The calculator returns your current AOV alongside monthly revenue, gross profit per order, and orders per day. From there, the uplift scenario lets you drag a target AOV increase and watch the annualized revenue and gross profit impact update in real time.

The reason gross profit sits next to revenue is deliberate. A higher AOV driven by discounting or low-margin bundles can grow your topline while shrinking what you actually keep. Entering your margin means the projection shows whether an eight-dollar AOV increase improves the bottom line or just inflates the headline number.

A/B test simulator

Plug in control versus variant AOV, order count, and conversion rate. The simulator calculates revenue per visitor for each group and flags the winner. It decides on revenue per visitor rather than AOV because revenue per visitor is the metric that accounts for both levers at once: a variant can raise average order value while converting fewer people, and the combination is what determines how much money the change makes.

That is the trap a simple AOV comparison misses. A variant that lifts AOV from $62 to $72 but drops conversion rate from 3.2% to 2.6% produces a lower revenue per visitor than the control ($1.87 against $1.98), so the higher order value is actually destroying value. The simulator catches this automatically.

Sample size

The sample size tab takes your baseline AOV, your AOV standard deviation, the minimum detectable effect you care about, and your target statistical power, then returns orders per variant, total orders needed, and an estimated runtime at your daily order volume. The curve updates live so you can see exactly how chasing a smaller effect trades off against how long the test runs. Color-coded status, green through amber to red, is blunt about whether the test is feasible at your traffic.

One number here is easy to misread, so it is worth stating plainly. The output is visitors per variant, not total sessions. If the tab returns 4,200, you need 4,200 in your control group and 4,200 in your variant group before the result is reliable, for 8,400 in total.


The four inputs that decide your sample size

The sample size tab asks for four things, and each one pulls the required number of visitors in a different direction. Understanding them is the difference between a test that ships a clear answer in three weeks and one that runs for months and tells you nothing.

Minimum detectable effect

The minimum detectable effect is the smallest AOV improvement you want the test to reliably catch. It is the single biggest driver of how much traffic you need, and the relationship is not gentle: halving the effect you want to detect roughly quadruples the sample you need, because the math scales with the square of the effect size. If your current AOV is $65 and you set the minimum detectable effect to $1, the calculator may ask for hundreds of thousands of sessions, which on most stores is a test measured in months. Starting with a 5 to 10 percent lift target keeps the timeline realistic. Set the effect to match the smallest improvement that would actually change your decision, not the smallest number the calculator will accept.

AOV standard deviation

Standard deviation measures how spread out your order values are, and it feeds directly into the sample size. A store where most orders cluster around $60 needs far fewer sessions to detect a $6 lift than a store selling everything from $15 accessories to $250 bundles. The wider the spread, the more noise sits between you and a clear signal, and the only way to overcome noise is more data. This is why two stores with the same AOV can need very different sample sizes, and why the tab asks for the deviation rather than assuming it.

Statistical power

Power is the probability that your test detects a real improvement when one genuinely exists. The standard setting is 80 percent, which means that even with a true winner in front of you, the test will fail to flag it roughly one time in five. Raising power to 90 percent lowers that miss rate but noticeably increases the sample you need. Power is the dial that protects you from the quiet failure mode of A/B testing: concluding that a change did nothing when it actually worked, simply because the test was too small to see it.

Statistical significance

Where power guards against missing a real effect, significance guards against believing in a fake one. It is the calculator's way of asking whether an observed AOV difference could plausibly have happened by chance. At 95 percent confidence, the simulator only flags a winner when the gap is unlikely to be a fluke, accepting about a 5 percent risk of calling a winner that is not really there. Significance and power address opposite errors, and the sample size tab balances both when it works out how many visitors you need. One keeps you from acting on luck; the other keeps you from ignoring a genuine improvement.


A worked example: testing a bundle

Say you sell a product individually at $45 and want to know whether a three-item bundle at $75 makes you more money. This is a textbook case for the simulator. You enter the control scenario (the single product at its current AOV and conversion rate) against the variant (the bundle at its expected AOV and conversion rate), and the tool projects which one produces more revenue per visitor.

The bundle will almost certainly raise AOV, since each order is worth more. The question the simulator actually answers is whether it raises AOV enough to offset any drop in conversion rate, because a $75 bundle will convert a smaller share of visitors than a $45 product. If the revenue-per-visitor math comes out ahead, you have a business case. The sample size tab then tells you how many visitors you need to prove it, and the uplift scenario translates the projected AOV change into a monthly gross profit figure you can take to whoever signs off on the test.


From calculator to live test

A projection is only useful if you can act on it. Once the calculator confirms a test is worth running, Elevate AB Testing lets you launch it directly inside your Shopify store. Price, bundle, shipping, and checkout tests run through the native theme editor, so there is no developer needed to set up the variant or split the traffic, and no second tool to learn. The plan you build here becomes a live experiment the same day.

Ready to run this test on your store? Price and AOV tests on Shopify Plus launch in minutes with Elevate AB Testing, with no developer ticket required. Start a 14-day free trial.

Frequently asked questions

What is average order value and why test it?
Average order value is your total revenue divided by your number of orders: the average a customer spends per purchase. It is worth testing because raising it grows revenue without needing more traffic, and unlike conversion rate, it is often easier to move with changes like bundles, thresholds, and pricing. A single dollar of AOV across thousands of orders adds up faster than most operators expect.
Why does the simulator decide on revenue per visitor instead of AOV?
Because revenue per visitor combines both things that matter, order value and conversion rate, into one number. A variant can raise AOV while converting fewer people, which can leave you worse off overall. Revenue per visitor is AOV multiplied by conversion rate, so it reflects the true financial result of a test rather than just one half of it.
How many visitors do I need for an AOV test?
It depends on four things: your baseline AOV, how spread out your order values are, the smallest improvement you want to detect, and your target statistical power. The sample size tab calculates it for you. As a rule of thumb, smaller effects and more variable order values both require dramatically more traffic, and the output is the number you need per variant, not in total.
What is a good minimum detectable effect to start with?
For most stores, a 5 to 10 percent relative lift is a practical starting point. Setting it much smaller can push the required sample into the hundreds of thousands of sessions, which turns a test into a multi-month commitment. Choose the smallest improvement that would genuinely change your decision to roll the change out, and size the test to that.
Can I run these tests on Shopify without a developer?
Yes. Elevate AB Testing runs price, bundle, shipping, and checkout tests through Shopify's native theme editor, so you do not need engineering involvement to set up a variant or split traffic. The calculator helps you plan the test; Elevate AB Testing runs it.