You will be able to explain why randomised experiments are strong evidence for causes and when they are not possible.
Daniel's manager is not convinced by the questions from lesson 4.1. "We could argue about confounders forever," she says. "How would we actually know if the app makes people spend more?" It is a fair challenge, and it has a clear answer. Instead of looking at customers who chose the app, you decide who gets nudged towards it, and you decide by chance.
That is the idea behind a randomised experiment. It is the strongest everyday tool for telling whether one thing causes another, and you can use a simple version of it at work.
Lesson 4.1, Three other reasons two things move together, showed that observational data, where you only watch what people chose to do, is full of hidden differences. Customers who download the app may be more loyal, richer, younger or furnishing a new flat. You can try to adjust for the differences you know about, but you can never be sure you have found them all.
Random assignment gets around this. You take one group of people and split it by chance, say with a coin flip or a random number in a spreadsheet, into a group that gets the treatment and a group that does not. Because chance decided who went where, every confounder, the ones you know about and the ones you do not, ends up spread roughly evenly across the two groups. Loyal and casual customers, rich and less rich, new homeowners and renters all land on both sides in similar proportions, as long as the groups are large enough.
The two groups now differ in only one systematic way: the treatment. So if their results differ by more than chance would explain, the treatment is the most likely reason. That is why medical research relies on randomised controlled trials before a new treatment is approved.
You have taken part in many randomised experiments without knowing it. When an app or website shows half its visitors one version of a page and half another, chosen at random, and then compares what each group does, that is an A/B test. It is a randomised experiment in a product.
Here is how Daniel could test the app question. Take 8,000 recent website customers who do not have the app. Assign them at random into two groups of 4,000. Group B gets an email with a voucher for downloading the app. Group A gets the same email without the app message. After three months, compare spending in the two groups.
The numbers below are a made-up illustration. Suppose 120 customers in Group A place a large order in those three months and 150 in Group B do. That is 3 percent against 3.75 percent, a relative rise of 25 percent. Because the groups were randomised, the difference is not explained by loyalty or new flats. It could still be chance, since 30 extra customers is a small number, and a proper test would check how likely a gap that size is by luck alone. Analytics tools and online calculators can do this, and anyone who runs A/B tests regularly should learn how. What the randomisation removes is the worry that the groups were different to begin with.
At work, you can use the same logic for many questions. Does the new onboarding email reduce support calls? Send it to a random half of new customers. Does a different subject line get more replies? Split your mailing list at random. The key step is that a coin, not a person, decides who gets what.
Many causal questions cannot be tested this way. You cannot randomly assign people to smoke, to grow up in a particular neighbourhood or to go through a recession. In those cases, two kinds of evidence come closest.
The first is a natural experiment: a situation where something outside anyone's control split people into groups in a way that was close to random. A policy that applied to people born after a certain date, or a change rolled out in one area before another for administrative reasons, can create groups that differ mainly in the one thing you care about. Researchers look for these carefully, and when they find one, the evidence can be nearly as strong as an experiment.
The second is many studies pointing the same way. When different research teams, using different data and methods, in different countries, keep finding the same link, it becomes harder to explain away with any single confounder. Each study has weaknesses, but they rarely all share the same one.
People often find a causal claim more convincing when someone explains how it would work. "The app has push notifications, so customers are reminded to come back and buy." A plausible mechanism like this does make a claim more believable. It tells you the claim is not absurd and gives you something specific to check, for example whether app users who turned notifications off spend less.
But a good story about how something could work is not evidence that it does. Plenty of treatments with convincing mechanisms have failed in trials. Treat a mechanism as a reason to take a claim seriously enough to test it.
There is probably a causal claim at your workplace that everyone repeats and nobody has tested: a sales technique, a training course, a new process. Pick one of those for the activity below, where you will design the comparison that could settle it.
Describe how you would test one causal claim at your workplace with a simple A/B comparison, including how you would assign the groups.
Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).