Conversion rate optimisation is sold as a discipline of continuous testing. For most Shopify stores that is not achievable arithmetic, and pretending otherwise produces confident decisions made on noise.
The numbers
To detect a change with any statistical confidence you need enough traffic for the result to be distinguishable from chance. For a standard two-sided test at 95% confidence and 80% power, per variant:
- 2% conversion rate, detecting a 10% relative lift: about 81,000 sessions per variant, so roughly 161,000 sessions for one A/B test.
- 2% conversion rate, detecting a 20% lift: about 21,000 per variant, roughly 42,000 total.
- 3% conversion rate, detecting a 10% lift: about 53,000 per variant, roughly 106,000 total.
Now compare that against your actual monthly sessions. If you are doing 15,000 a month, a single test that could detect a 10% improvement would take most of a year, during which nothing else about the store may change or the test is invalid.
This is why so much CRO reporting is nonsense. Tests get called at two weeks because that is the reporting cycle, on a few thousand sessions, and the winner is whichever variant got luckier.
What "we got a 30% lift" usually means
Stopping a test early, when it happens to be up.
Conversion rates wander. Run any A/A test, where both variants are identical, and you will watch one "win" by double digits at some point in the first fortnight purely from randomness. If you stop when you like the number, you will find a winner every time, and the effect will not survive contact with the next month's revenue.
What is actually worth doing below the threshold
Plenty, and none of it is testing.
Find what is broken. Not suboptimal, broken. A checkout that fails on one mobile browser, a variant picker that does not update the price, a shipping cost that first appears on the final step. These do not need a test because a fix is not a hypothesis.
Watch session recordings instead of dashboards. Twenty recordings of people who reached the cart and left will tell you more than any test you can afford to run. You are not measuring, you are finding out what confuses people, and for that a sample of twenty is fine.
Fix the mobile experience specifically. Most of these stores get the clear majority of their traffic on mobile and are designed on a desktop by someone who checks mobile last.
Change things that are obviously right. Adding the delivery estimate to the product page, showing the returns policy before checkout, putting real product photography where a supplier stock image is. You do not need statistical proof to stop guessing about the delivery date on your own store.
What we tell clients
If you are under roughly 30,000 sessions a month, we will not sell you a testing programme, because we cannot honestly promise the results would mean anything. What we will do is the unglamorous version: audit the store on a real phone, fix what is broken, remove friction that has no defenders, and improve the pages that get the most traffic.
That work has no p-value attached and it is the work that moves the number at this size.
Once traffic supports real tests, testing becomes the right tool and we will say so. The mistake is running the ritual of testing at a scale where the ritual cannot produce knowledge.
That is the approach behind our conversion rate work, and it is why the first month usually looks like a repair job rather than an experiment.
