You will be able to explain why small tests mislead and use a sample size calculator before starting a test.
Farah started her size chart test from lesson 7.1 on a Monday. By Wednesday, version B had 9 add-to-carts from 41 visitors and version A had 4 from 39. B was more than twice as good. She was ready to switch every page over that afternoon. By the end of the second week, the two versions were almost level.
Nothing had gone wrong with the test itself. Early results swing about by chance, and an early winner very often fades as more visitors arrive. This lesson explains why, in plain words, and shows how to work out how long a test needs before you start it.
Think of each visitor as a coin toss that comes up "adds to cart" some of the time. Even if both versions are exactly as good as each other, a few dozen tosses will rarely split evenly, and whichever version gets the lucky run looks like the winner.
The swings are bigger than most people expect. Suppose the true add-to-cart rate on Farah's page is 12 percent (an example figure). With only 40 visitors, the share who add to cart will usually land anywhere between about 2.5 percent and 22.5 percent, purely by chance: somewhere from 1 to 9 people. So B's 9 out of 41 and A's 4 out of 39 is exactly the kind of gap that luck alone produces, even with no real difference at all.
As the number of visitors grows, the swings shrink, and the measured rate settles closer to the true one. That is why the size of the sample decides what a test can tell you.
The second idea follows from the first. If version B is much better than A, the gap shows through the chance swings quickly. If B is only a little better, the gap is buried in the noise until you have a lot of visitors.
Here are Farah's numbers, worked out with a standard sample size formula for comparing two rates. The inputs are example figures: a current add-to-cart rate of 12 percent, about 450 visitors a week to the trousers page split between the two versions, and the two standard settings most calculators start with, a 95 percent confidence level and 80 percent power, which are explained below.
To detect a rise from 12 to 18 percent, she needs about 555 visitors per version, roughly 1,110 in total, or about two and a half weeks of traffic.
To detect a rise from 12 to 15 percent, she needs about 2,036 per version, roughly 4,070 in total, or about nine weeks.
To detect a rise from 12 to 13 percent, she needs about 17,169 per version, over 34,000 in total, or about 76 weeks. A year and a half, for one test.
The pattern is the point. Halving the size of the difference you want to detect roughly quadruples the sample you need. A small business cannot run tests for tiny differences, which is why lesson 7.1 says to test changes big enough to matter.
When a testing tool says a result is statistically significant, it is answering a narrow question. Suppose there were really no difference between A and B. How unlikely would a gap this big be, from chance alone? When the answer is very unlikely, the tool declares a winner.
Most tools use a 95 percent confidence level for that call, which means accepting a 5 percent chance of naming a winner when the versions are really the same. Power, usually set at 80 percent, works from the other direction. It is the chance that the test spots a real difference of the size you chose, if there is one. Both numbers are habits of the trade rather than rules, and stricter settings simply need a bigger sample.
The verdict has limits. It does not say how big the effect is, whether it matters to the business, or that the winner will keep winning. A confirmed result on a tiny difference may not be worth acting on. A test that names no winner does not prove the versions are equal either; it means the test could not tell them apart with the visitors it had.
Now back to Farah's Wednesday. The calculation above assumes you look at the result once, at the end. If you check every day and stop as soon as one version looks clearly ahead, you give chance many opportunities to produce a lucky streak, and one of those streaks will often cross the line. The more often you peek and are willing to stop, the more likely your "winner" is a fluke.
That is why lesson 7.1 asks you to set a stopping rule before you start, and why Farah's rule was to run the test until each version had at least 555 visitors and the test had covered three full weeks, whichever came later. She could look at the numbers along the way. She just could not stop early because she liked what she saw.
You do not need to do the maths yourself. Free sample size calculators for A/B tests are easy to find online, and many testing tools include one. They ask for the same inputs Farah used: your current conversion rate for the measure you are testing, the smallest change you want to be able to detect, and the confidence and power settings, which you can usually leave at 95 and 80 percent.
The calculator gives you the visitors or recipients needed per version. Divide the total by your weekly traffic to the page, round up to full weeks, and you know how long the test must run. If the answer is six months, that is useful too. It tells you to test something bolder, or to test somewhere with more traffic, which is the subject of lesson 7.3.
Your own homepage is the obvious first page to try this on, with GA4 supplying the current conversion rate and the weekly traffic.
Put your own traffic and conversion rate into a free sample size calculator and note how many weeks a test of your homepage would need.
Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).