Data mining, overfitting and transaction costs

You will be able to judge whether a backtest result is likely to hold up out of sample.

A video in Marcus's feed promised a moving average rule that "beat the market by 6% a year for twenty years". The presenter had tested every pair of moving averages from 5 days to 300 days on one index and picked the pair with the best record, 37 days and 187 days, whose backtest chart climbed beautifully. Marcus noticed that nobody would choose 37 and 187 for any reason except that they had worked best on that particular history, and that was the problem.

A backtest applies a trading rule to past prices to see how it would have done. It's the right way to check a rule, and the most common way to fool yourself. This lesson covers the three ways backtests mislead, data mining, overfitting and missing costs, and how to judge whether a result is likely to hold up, using made-up figures.

Test enough rules and some will win by luck

Suppose you test 100 trading rules that have no real predictive power at all. The usual statistical test accepts a result if it would happen by chance only one time in twenty. Run 100 useless rules through it and about five should pass on luck alone, and if those five are the only ones you report, they'll look like discoveries.

That's data mining: searching through many rules on one set of data until some look good. It doesn't require dishonesty. It happens whenever someone tries variations until one works, then forgets how many they tried. The video's 37 and 187 came from testing thousands of pairs. Lesson 10.1, What technical analysis claims and what the evidence supports, described what happened when researchers allowed for this on the rules tested in an earlier study: the results looked much weaker.

The fix starts with honest accounting. Keep a list of every rule you test, the losers along with the winner. The more you tried, the better the winner has to look before you take it seriously.

A rule tuned to the past fits the noise

Overfitting is data mining's close relative. It happens when a rule has enough settings to adjust that it can be tuned to match the quirks of one particular history. Every price series contains real patterns, if any, and random noise. A rule with many adjustable settings can fit both, and the noise won't repeat.

You can spot overfitting by its precision. Round, simple settings such as 50 and 200 days, chosen before looking at the results, are less likely to be fitted to noise than 37 and 187. A rule with five conditions, each with its own threshold, has far more room to fit noise than a rule with one.

The standard defence is out-of-sample testing. Split your history before you start. Use the first part, say the first twelve years of twenty, to design the rule and choose its settings. Then lock the rule and test it once on the last eight years, which you haven't touched. If it works only on the data used to design it, it was fitted to that data. A stricter version, called walk-forward testing, repeats this in rolling steps, always designing on the past and testing on the period that follows.

Two quieter errors belong here too. Look-ahead bias uses information that wasn't available at the time, such as acting on a day's closing price at that same close, when you could only know it afterwards. Survivorship bias, from lesson 1.7, Read a fund's performance table without being misled, creeps in when a stock data set leaves out companies that were delisted, which makes any stock-picking rule look better than it was.

Costs turn winners into losers

Most backtests in videos and forums show returns before trading costs. Real trades pay a spread, a commission and some slippage every time, as lessons 5.3 and 5.8 measured. Rules that trade often pay these costs often.

Take a made-up rule that beats buying and holding by 2 points a year before costs, and trades 24 times a year. At 0.15% per trade, a reasonable figure for a large ETF with a low commission, costs come to about 3.6% a year. The 2-point advantage becomes a loss of about 1.6 points a year against buying and holding. Most of the rules that look profitable in forum posts trade far more than this.

Small accounts suffer most, because minimum commissions take a bigger share of small trades. Here's the arithmetic for a made-up monthly-rebalanced rule on a S$20,000 account that places four orders of S$2,000 each month. That's 48 orders a year. With a made-up minimum commission of S$10, commissions come to S$480. Spread costs at a made-up 0.15% of the S$96,000 traded add S$144. The total of S$624 is about 3.1% of the account every year, before any slippage or currency conversion. A rule would need to beat buying and holding by more than 3 points a year just to break even.

What a rule worth trusting looks like

A rule deserves some trust when it passes four tests. It's simple, with few settings, chosen for a reason before you saw the results. It makes economic sense: there's a plausible reason, such as investors' slow reaction to news, why it might work. It survives realistic costs from your own broker's fee schedule. And it works on data it wasn't designed on, in more than one market or period.

Very few rules pass all four. That's useful to know, because it means most of the confident chart claims you'll meet can be set aside without much effort, and the rare one that passes deserves a careful test of your own, which lesson 10.8, Test one trading rule with costs on real price history, walks through.

For the activity, open your broker's fee schedule and get the figures you need to price a year of monthly trading in a S$20,000 account: commission per order, any minimum, platform fees and currency conversion charges.

List the costs a monthly-rebalanced rule would pay in a year on your broker and convert them to a percentage of a S$20,000 account.

Course

Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).