Model AOV uplift, simulate an A/B test on revenue per visitor, and size your sample before you run it. Built for Shopify stores.
Most Shopify operators decide what to test the same way they decide what to cook for dinner: by feel. The problem is that average order value moves money differently than conversion rate does, and a test that looks like a winner on one metric can quietly lose on another. This calculator exists so you can answer three questions before you spend a single session of traffic: how much would a higher AOV actually add, would a specific change beat your current setup, and how long would the test take to give you a trustworthy answer.
It has three connected tabs. The AOV calculator turns your current revenue and orders into a baseline and projects an uplift scenario. The A/B test simulator pits a control against a variant and decides the winner on revenue per visitor, not order value alone. The sample size tab tells you how many visitors each variant needs before the result means anything. Work through them in order and you move from a rough idea to a sized, fundable experiment.
Enter your total revenue, number of orders, the time period those cover, and your average margin. The calculator returns your current AOV alongside monthly revenue, gross profit per order, and orders per day. From there, the uplift scenario lets you drag a target AOV increase and watch the annualized revenue and gross profit impact update in real time.
The reason gross profit sits next to revenue is deliberate. A higher AOV driven by discounting or low-margin bundles can grow your topline while shrinking what you actually keep. Entering your margin means the projection shows whether an eight-dollar AOV increase improves the bottom line or just inflates the headline number.
Plug in control versus variant AOV, order count, and conversion rate. The simulator calculates revenue per visitor for each group and flags the winner. It decides on revenue per visitor rather than AOV because revenue per visitor is the metric that accounts for both levers at once: a variant can raise average order value while converting fewer people, and the combination is what determines how much money the change makes.
That is the trap a simple AOV comparison misses. A variant that lifts AOV from $62 to $72 but drops conversion rate from 3.2% to 2.6% produces a lower revenue per visitor than the control ($1.87 against $1.98), so the higher order value is actually destroying value. The simulator catches this automatically.
The sample size tab takes your baseline AOV, your AOV standard deviation, the minimum detectable effect you care about, and your target statistical power, then returns orders per variant, total orders needed, and an estimated runtime at your daily order volume. The curve updates live so you can see exactly how chasing a smaller effect trades off against how long the test runs. Color-coded status, green through amber to red, is blunt about whether the test is feasible at your traffic.
One number here is easy to misread, so it is worth stating plainly. The output is visitors per variant, not total sessions. If the tab returns 4,200, you need 4,200 in your control group and 4,200 in your variant group before the result is reliable, for 8,400 in total.
The sample size tab asks for four things, and each one pulls the required number of visitors in a different direction. Understanding them is the difference between a test that ships a clear answer in three weeks and one that runs for months and tells you nothing.
The minimum detectable effect is the smallest AOV improvement you want the test to reliably catch. It is the single biggest driver of how much traffic you need, and the relationship is not gentle: halving the effect you want to detect roughly quadruples the sample you need, because the math scales with the square of the effect size. If your current AOV is $65 and you set the minimum detectable effect to $1, the calculator may ask for hundreds of thousands of sessions, which on most stores is a test measured in months. Starting with a 5 to 10 percent lift target keeps the timeline realistic. Set the effect to match the smallest improvement that would actually change your decision, not the smallest number the calculator will accept.
Standard deviation measures how spread out your order values are, and it feeds directly into the sample size. A store where most orders cluster around $60 needs far fewer sessions to detect a $6 lift than a store selling everything from $15 accessories to $250 bundles. The wider the spread, the more noise sits between you and a clear signal, and the only way to overcome noise is more data. This is why two stores with the same AOV can need very different sample sizes, and why the tab asks for the deviation rather than assuming it.
Power is the probability that your test detects a real improvement when one genuinely exists. The standard setting is 80 percent, which means that even with a true winner in front of you, the test will fail to flag it roughly one time in five. Raising power to 90 percent lowers that miss rate but noticeably increases the sample you need. Power is the dial that protects you from the quiet failure mode of A/B testing: concluding that a change did nothing when it actually worked, simply because the test was too small to see it.
Where power guards against missing a real effect, significance guards against believing in a fake one. It is the calculator's way of asking whether an observed AOV difference could plausibly have happened by chance. At 95 percent confidence, the simulator only flags a winner when the gap is unlikely to be a fluke, accepting about a 5 percent risk of calling a winner that is not really there. Significance and power address opposite errors, and the sample size tab balances both when it works out how many visitors you need. One keeps you from acting on luck; the other keeps you from ignoring a genuine improvement.
Say you sell a product individually at $45 and want to know whether a three-item bundle at $75 makes you more money. This is a textbook case for the simulator. You enter the control scenario (the single product at its current AOV and conversion rate) against the variant (the bundle at its expected AOV and conversion rate), and the tool projects which one produces more revenue per visitor.
The bundle will almost certainly raise AOV, since each order is worth more. The question the simulator actually answers is whether it raises AOV enough to offset any drop in conversion rate, because a $75 bundle will convert a smaller share of visitors than a $45 product. If the revenue-per-visitor math comes out ahead, you have a business case. The sample size tab then tells you how many visitors you need to prove it, and the uplift scenario translates the projected AOV change into a monthly gross profit figure you can take to whoever signs off on the test.
A projection is only useful if you can act on it. Once the calculator confirms a test is worth running, Elevate AB Testing lets you launch it directly inside your Shopify store. Price, bundle, shipping, and checkout tests run through the native theme editor, so there is no developer needed to set up the variant or split the traffic, and no second tool to learn. The plan you build here becomes a live experiment the same day.