Three other reasons two things move together

You will be able to list the alternatives to causation whenever two things are linked.

Daniel works in marketing for an online home goods retailer. His team's monthly report has a striking finding: customers who download the app spend about twice as much a year as customers who only use the website. His manager's conclusion is quick: "We need to push everyone to download the app." The budget for an app download campaign is approved the same week.

The finding may be accurate. The conclusion still does not follow. App users and big spenders go together, but that tells you nothing yet about whether the app makes people spend more. Lesson 2.3, Sunk costs and other traps in your own head, introduced post hoc reasoning, where B follows A and A gets the credit. This lesson goes further. Whenever two things are linked, there are at least four explanations other than "one causes the other", and you can check for each of them.

A third factor drives both

The first and most common alternative is a confounder: a third factor that drives both things, so they rise and fall together while neither one causes the other.

A classic example: people who have gym memberships tend to be healthier. The gym may help. But people with higher incomes are more likely to afford a membership and are also more likely to be healthier for many other reasons, such as better housing, more time to cook and less physically damaging work. Income sits behind both the membership and the health. Even if the gym did nothing, you would still see the link.

For Daniel's report, the obvious confounder is how much a customer already likes the shop. A loyal customer who buys often is more likely to bother downloading the app. Loyalty drives both the download and the spending. If that is the main story, a campaign that gets casual customers to install the app leaves them as casual as they were.

To find confounders, ask: what kind of person, team or situation would be more likely to have both of these things? Whatever you come up with is a candidate.

The effect is causing the cause

The second alternative is reverse causation, where the arrow points the other way and the thing you took for the effect is actually producing the supposed cause.

In Daniel's case, picture a customer who has just collected the keys to a new HDB flat and will spend heavily over the next few months on furniture and fittings. Because she is ordering so much, she installs the app to track deliveries and use vouchers, so her spending came first and led to the download.

Reverse causation is common in workplace data. Teams that hold more meetings have more problems, possibly because problems cause meetings. Employees who use the company's wellness programme report more stress, possibly because stressed employees seek it out. Ask: could the second thing have caused the first?

Pure chance

The third alternative is chance. If you look at enough pairs of numbers, some will line up by accident. Over any few years, the number of something in one country can track the number of something completely unrelated elsewhere, simply because both happened to rise. Collections of these absurd matches exist online and are worth a laugh.

At work, chance matters most when someone has searched through many possibilities to find a pattern. Check fifty customer characteristics against spending and a few will appear linked by luck alone, and of course the one that reaches the slide is the most interesting of them, while the forty-nine that showed nothing go unmentioned. So ask how many other things were checked before this one was found.

Small samples make chance worse. A link seen across twelve customers is far more likely to be luck than one seen across twelve thousand.

How the group was chosen

The fourth alternative is selection: the way a group was picked creates a link that does not exist across everyone.

Suppose Daniel's report only includes customers who made at least one purchase this year. Website users who browsed and bought nothing are not in the data at all. Depending on who is left out, the comparison between app users and website users can look quite different from what it would be across all visitors. This is a close cousin of survivorship bias from lesson 3.4, Survivorship bias and missing data.

Selection is easy to miss because it happens before the analysis starts, in a decision about which rows to pull. Ask who was included, who was excluded and whether the exclusion could be related to both things being compared.

What Daniel asks before the campaign

Daniel does not argue that the app is useless. He writes down the four alternatives with a question for each: were app users already heavier buyers before they downloaded it, and what did spending look like before and after the download for the same customers? How many other customer features were compared before this one stood out? And does the comparison include visitors who never bought anything at all?

The first two questions are the most useful, because they target the confounder and reverse causation directly. If customers' spending was already high before they downloaded the app, and did not change much afterwards, the campaign is unlikely to work. Lesson 4.2, Why experiments beat observation, shows how Daniel could test the idea properly, with a small trial instead of a full campaign.

Claims of the form "people who do X have better Y" are everywhere: in health articles, management books and company reports. In the activity below you will take three of them and write one confounder and one reverse explanation for each.

Take three claims of the form people who do X have better Y, and write one confounder and one reverse explanation for each.

Course

Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).