Conversion rate optimisation and A/B testing for online stores
Conversion rate optimisation (CRO) is the practice of getting more of your existing visitors to buy, by finding where the site loses them and fixing it with evidence. A/B testing is one tool within it: you show half your visitors the current page and half a changed version, and measure which earns more. It is the most reliable way to know whether a change worked, but only when enough people see the test.
That condition rules out more stores than the industry admits. Detecting a small improvement needs tens of thousands of visitors per version. A store with a few hundred orders a month cannot run a valid test on a button colour or a headline, and a test that ends early will tell you something false with great confidence.
So the useful question is what kind of optimisation your traffic supports. Larger stores should test continuously. Smaller ones should fix what is obviously broken, make bigger changes, and measure them honestly.
The short version
- CRO is research first and testing second: most of the value is in finding the right problem.
- An A/B test is only valid if you decide the sample size before it starts and run it to the end.
- At a 2% conversion rate, detecting a 10% relative lift needs roughly 80,000 visitors per version.
- Small stores should test big changes, or skip testing and fix known problems.
- Judge tests on revenue per visitor as well as conversion rate, because a discount can lift one and sink the other.
- Most tests do not produce a winner, and a programme that reports only wins is not being measured properly.
How do you find what to test?
Not by brainstorming. Ideas from a meeting room reflect the opinions of the people in it. Good test ideas come from evidence that visitors are struggling at a specific point.
- Follow the funnel in your analytics. Product page to basket, basket to checkout, checkout to order, split by device. Look for the step where mobile falls furthest behind desktop, or where one category underperforms the rest.
- Watch real sessions. Session recordings and heatmaps show hesitation: repeated taps on something that is not a link, scrolling up and down a product page looking for delivery costs.
- Read what customers tell you. Customer service tickets, live chat logs, reviews and returns reasons. If ten people a week ask whether an item comes with a plug, the product page has a gap.
- Read your on-site search terms. People type their unanswered questions into the search box. Our guide to site search and merchandising covers how to pull these.
- Check speed and errors. A slow or broken step is not a test candidate. It is a fix.
- Write a hypothesis. "Because we saw X, we believe changing Y will cause Z, measured by this metric." If you cannot fill in the first part, you are guessing.
Then rank the list by how many visitors the change touches, how strong the evidence is, and how much work it takes. Checkout and product pages usually come first because every buyer passes through them.
A real example of that chain: for Euronics, analysis showed shoppers adding items to the mini cart and then hesitating. The hypothesis was uncertainty about delivery. Showing a delivery date for each product in the mini cart lifted conversion rate by 18.14% and revenue by 33.15% in the test. The idea came from the data, not from a list of best practices.
How A/B testing works
Visitors are split at random. The control group sees the current version, the other group sees the variant. Because the split is random and both run at the same time, the weather, a sale or a newsletter affects both groups equally. Any remaining difference is either the change or chance.
Statistics deal with the chance part. Two numbers matter:
- Significance, usually set at 95%, limits how often you will declare a winner when there is no real difference.
- Power, usually set at 80%, is how likely the test is to detect a real difference of the size you care about.
The rules that keep a test honest are simple and often broken. Decide the sample size before you start. Run for whole weeks, at least two, so weekdays and weekends are both covered. Do not stop the moment the dashboard turns green: checking daily and stopping on the first significant result produces false winners. When it has finished, our free A/B test significance calculator will tell you whether the difference is likely to be real.
On tools: Google Optimize, the free option many stores used, closed on 30 September 2023. Shopify now has experiments built into its Rollouts feature, available on the Grow plan and above at the time of writing (October 2026), which can split traffic between theme versions. Dedicated testing platforms, and personalisation tools such as Nosto, cover the rest.
How much traffic do you need for a valid test?
More than most people expect. The smaller the improvement you want to detect, the more visitors you need, and the relationship is steep: halving the lift you are looking for roughly quadruples the sample.
These figures assume a 2% starting conversion rate, 95% significance and 80% power, using the standard formula for comparing two proportions.
| Lift you want to detect | Conversion rate moves from | Visitors needed per version | Total for a two-way test |
|---|---|---|---|
| 10% | 2.0% to 2.2% | About 80,000 | About 160,000 |
| 20% | 2.0% to 2.4% | About 21,000 | About 42,000 |
| 50% | 2.0% to 3.0% | About 3,800 | About 7,600 |
Put that against a real store. With 30,000 visitors a month and a 2% conversion rate, you take 600 orders a month. A test looking for a 10% lift would need more than five months of all your traffic. Over that long, seasons change, cookies expire and the result is unreliable anyway. A 20% lift would take about six weeks, which is workable. And a test on the basket page only counts the visitors who reach the basket, so the numbers are tighter still.
Minor changes, such as button colour, a reworded headline or a moved trust badge, rarely produce lifts of 20%. That is the whole problem. A small store testing small changes will see a string of inconclusive results, or worse, will stop early on a lucky week and ship something that did nothing.
What to do if you cannot test properly
- Fix what is plainly broken. A form that fails on one browser, a missing delivery price, a slow mobile product page. These need no test, only a fix and a check.
- Make bigger changes. A rebuilt product page template, a new checkout flow or a different offer structure can move conversion by enough to detect. Test those.
- Test higher up the funnel. Add-to-basket rate happens far more often than purchase, so it reaches significance sooner. Treat it as a guide, not proof of revenue.
- Use before-and-after comparison, carefully. Compare the same weeks year on year, account for promotions and traffic mix, and accept that this is weaker evidence than a test.
- Talk to customers. Five moderated user tests will show you problems that no dashboard will, at any traffic level.
- Test where volume is higher. Email subject lines and ad creative reach more people than a single page. Our guide to email and SMS covers the flows worth testing.
Common mistakes
- Stopping early. The most common cause of "winning" tests that make no difference once live.
- Copying another retailer's winner. Their customers, prices and traffic are not yours.
- Measuring conversion rate alone. Track revenue per visitor and, where you can, margin.
- Running overlapping tests on the same pages without enough traffic to separate their effects.
- Testing during a sale and assuming the result holds at full price.
- Ignoring device. A change can help on desktop and hurt on mobile. Lisa Angel's tested change was specific to mobile: a simplified header that raised conversion rate by 5.9% and revenue per user by 13.2%.
- Expecting every test to win. Inconclusive and losing tests are normal. A losing test that stops you shipping a harmful change has paid for itself.
When CRO is the wrong priority
If the site has very little traffic, the product is uncompetitive on price or delivery, or stock is unreliable, optimising pages will not rescue it. Sort the proposition and the traffic first. CRO multiplies what is already there.
How we'd approach it
- Check the measurement. We confirm that analytics records every funnel step and that revenue matches the store's own figures. Optimising against wrong numbers is worse than not optimising.
- Do the sums on traffic. We work out what size of effect your store can detect in four to six weeks. That decides whether you get a testing programme or a prioritised list of fixes.
- Research. Funnel analysis, recordings, search terms and customer feedback, turned into a ranked backlog of hypotheses.
- Fix the obvious without testing, and record the before and after.
- Test where it is valid, with sample size and run time agreed in advance, and report every result including the losers.
That is the shape of our optimisation work, and it sits within our wider growth and optimisation practice alongside paid media and email.
Questions we get asked
What is a good e-commerce conversion rate?
How long should an A/B test run?
Can a small store do CRO at all?
Does A/B testing harm SEO?
What does 95% significance actually mean?
Want a second opinion on your own setup? Thirty minutes with a strategist, and you leave with a clear next step.