You will be able to set up a fair A/B test on an email and know when the result is too small to trust.
Priya changed three things in one newsletter: a question as the subject line instead of a statement, a Tuesday send instead of a Thursday, and a button instead of a text link. The issue got more clicks than any before it. She was delighted, and then stuck. Which of the three changes worked? Should she keep all of them? She had no way to know.
Testing is how you find out what actually works for your list, rather than what someone online says works. Done carelessly, it produces confident answers that are wrong. This lesson covers how to run a fair test on an email, and how to tell when a result is too small to mean anything.
An A/B test sends two versions of an email that differ in one way to two randomly chosen parts of your list, then compares the results. Most email tools have a built-in feature for this, often called an A/B test or split test.
The important words are "one way". If version A and version B differ in subject line and send time, and B wins, you cannot say which change caused it. Priya's three-change newsletter was not a test at all, just a different email.
Good candidates for a single change are the subject line, the preview text, the send day or time, the call to action wording, the position of the main link, or plain text against a designed template. Pick the one you most want to learn about, and keep everything else identical.
Before the test goes out, write down which number decides the winner and what result would make you act.
For most email tests, the measure should be clicks. Lesson 8.1 explained why opens are unreliable: privacy features count many emails as opened that nobody read, so a subject line test judged on opens may simply reward whichever version happened to go to more Apple Mail users, while a click only happens when a person acts. If your email asks for replies, replies work too. If you can trace bookings or orders through the UTM tags from lesson 8.2, those are better still, though they are usually too few on a small list to decide a test alone.
Deciding the measure in advance stops you from picking whichever number happens to favour the version you liked.
Here is the uncomfortable part. On a small list, most differences between two versions are chance.
Suppose Priya's example list of 300 parents is split in half. Version A gets 12 clicks from 150 people, and version B gets 15 clicks from 150. B looks better, 10 percent against 8 percent. But three clicks is a tiny difference. If she sent version A twice to two random halves, she could easily see a gap that size between two identical emails, just from which parents happened to be free to read that evening.
As a working rule, a difference of a few clicks is usually chance, and the smaller your list, the bigger a gap has to be before it means anything. Some email tools show a confidence figure beside a test result. Use it if yours does, and be cautious if it does not.
For the proper treatment of sample size, see module 7 of the course Marketing analytics: GA4, attribution and testing, where lesson 7.3, Testing when you do not have much traffic, speaks directly to small lists. For email on a small list, though, the habit below will help you more than any formula.
When one version wins, do not turn it into a rule straight away. Try the same kind of change on your next two or three sends. If question-style subject lines beat statement-style ones three times out of three, you have learned something about your readers. If they win once and lose twice, the first result was probably chance.
Keep a simple test log alongside the send log from lesson 5.5, with these columns: date, what changed, version A, version B, the measure, the result for each, and what you concluded. Over a few months, the log becomes a record of what works for your list specifically, which is more useful than any general advice.
Split a very small list in two and each half may be too small for one send to teach you anything. Compare issue against issue over time instead. Alternate the two approaches across several months and read the pattern in your log, which is slower and far more honest.
Some changes need no test because the answer is already clear. A broken link, missing alt text or a fake RE: in a subject line gets fixed straight away, with no test needed.
In the activity below you will write a test plan for your next send: the one change, the measure, how you will split the list, and what result would change what you do in your following email.
Write a test plan for your next send naming the one change, the measure, the split and what result would change your next email.
Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).