Not all evidence weighs the same

You will be able to rank pieces of evidence by how much they can support a claim.

Wei Ling runs operations for a small logistics firm in Tuas. On Tuesday her manager, Kumar, forwards a sales deck for a new warehouse scanning system. Page three has a quote from a happy client: "Our picking errors dropped overnight." Page five says the system is used by over 200 warehouses. That afternoon a friend in a WhatsApp group mentions that her company tried something similar and hated it. By Wednesday Wei Ling has three pieces of evidence pointing in different directions, and she has to tell Kumar what she thinks.

Lesson 1.1, Claim, evidence and the link nobody says out loud, showed you how to separate a claim from what it rests on. This lesson is about the next question: once you have the evidence in front of you, how much can each piece hold up?

One story is not a pattern

The client quote and the friend's complaint have something in common. Each is a single story. Stories are easy to remember and easy to repeat, which is why sales decks and group chats are full of them. They tell you that something happened at least once. They tell you very little about how often it happens.

That is the problem with the quote on page three. The vendor chose it out of every client they have. If the system worked well for one warehouse in ten, the deck would still show you that one. The friend's complaint has the same weakness pointing the other way. One bad rollout says something went wrong somewhere, but it could have been the system, the training, the timing or the company.

None of this makes stories useless. A story can point to a mechanism you had not thought of, and the friend's remark that the scanners struggled in a cold room is something Wei Ling can go and check. What a story cannot do is tell you how likely the good or bad outcome is for you.

Bigger and fairer samples carry more weight

The figure on page five is a different kind of evidence. It counts something. Counts and surveys can carry more weight than stories, but only under two conditions.

The first is size. Ten customers who all say the same thing tell you more than one, and five hundred tell you more than ten, because in a small group a couple of unusual cases can swing the whole result.

The second condition matters more: the group has to look like the people the claim is about. A survey of five hundred customers who chose to fill in a feedback form is large, but it is not representative. People who are very happy or very annoyed fill in forms. The quiet middle does not. A representative sample is one where the people counted look like the wider group the claim is about, in the ways that matter for the question.

So "used by over 200 warehouses" is better than one quote, but Wei Ling still needs to ask what kind of warehouses, how long they have used it, and how many tried it and left. A count of current users says nothing about the ones who gave up, which you will meet again in lesson 3.4, Survivorship bias and missing data.

Ask who made it and what they left out

Every piece of evidence was produced by someone. Before you weigh it, ask three plain questions. Who produced this? What do they gain if I believe it? What would they have left out?

The vendor gains a sale. That does not make the deck false, but it means the deck will show the best clients, the best results and the most flattering way to count them. If the new system was Kumar's idea, he has a stake as well, and he will notice the good news first. Wei Ling's friend has no stake in the sale, which is a point in her favour, but she may have a grudge against a project that made her weekends miserable.

The question about what was left out does the most work. A deck that reports error rates after the rollout but not before has left out the comparison you need. A survey that reports satisfaction but not how many people answered has left out the size. Missing pieces are often more telling than the ones on the page.

Ten copies of one source are still one source

The last trap is repetition. Suppose Wei Ling searches online and finds the same "errors dropped overnight" line on a trade website, a LinkedIn post and a blog, which feels like three confirmations until she notices that all three quote the same press release. That is one source, seen three times.

Evidence gets stronger when separate sources, working on their own, reach the same answer. If two warehouses Wei Ling contacts directly, with no link to the vendor, both report fewer picking errors, that agreement means something. Each one could have been wrong in its own way, and they were not.

In practice, a rough ranking from weakest to strongest looks like this. A single story or testimonial sits at the bottom. Next comes a handful of hand-picked cases, then a count or survey of a group that may not be representative. Above that is a larger, fairer sample. At the top are several independent sources that agree. A piece of evidence can move up or down this ladder depending on who produced it and what they left out.

Putting it together

Wei Ling writes back to Kumar on Wednesday afternoon. She says the deck has one hand-picked quote and a user count with no information on the users who left. The friend's story is a single case, but it raises a specific worry about cold rooms that she can check. Her suggestion is to call two warehouses of a similar size that the vendor did not choose, and to ask the vendor for error rates before and after at three sites. That reply took her ten minutes, and it gives Kumar something to act on.

Think of a decision your own team made recently, such as a new tool, a change of supplier or a hire. What evidence was on the table when you made it? In the activity below you will list it and put each piece in order of how much weight it could really carry.

List the evidence behind one decision your team made recently and rank each piece from weakest to strongest with a reason.

Course

Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).