Costs and fills come first
Your backtest uses the spread the broker advertises. Real fills cost more, especially at the open and around news. Measure it: compare the price you wanted to the price you got, and average the gap. On one system over 164 live trades that came to 0.13R per trade, about double what was assumed.
Your backtest also fills your stop at your stop price, because that's the only number it has. Live, that price is often gone. Trailing stops are the worst for this. If the edge dies once you add real slippage and trading costs, stop testing anything else.
Your data knows things you wouldn't have known
Three ways this happens.
Indicators need history before they work. A 200 period average needs 200 candles. If your data starts the same day your test starts, the engine either skips those bars or borrows candles from before the start. Borrowing is a leak.
Price files get adjusted. A stock split cuts the price in half overnight, and on the wrong file that looks like a crash. A dip buying strategy catches every one, and every win is fake.
Stock lists only show survivors. Build your list from what's trading today and you've deleted every company that went to zero.
Too few trades, too many tries
At a 40% win rate, 70 trades tells you almost nothing. The real number could be 29% or 52% and the sample can't tell them apart. Those are two different businesses.
Losing streaks work the same way. At 40%, nine losses in a row is normal across 170 trades. Most people read that as the system dying.
And the more versions you test, the more likely the best one is luck. Try 50, keep the winner, and you probably found noise. AI tools make this much worse, because they'll test hundreds while you watch. That's overfitting.
You judged it on the wrong trades
A filter that only fires on 20 trades out of 400 still gets scored on all 400. The other 380 water it down until it looks harmless. Score it on the trades it actually changed.
Time windows do the same thing. Your 1 year sits inside your 2 year, which sits inside your 5 year. Five results, same recent trades in all of them. Use separate blocks that share no trades. Where they disagree is the real answer. That's the point of out of sample and walk forward testing.
You only saw one version of history
Your backtest is one order of trades. Shuffle the same trades a few thousand times and you get a range instead of one number, and the bad end is usually much worse than what you saw. That's monte carlo. The worst drawdown you've had is only the worst so far.
What to check first
Cost first. Then the ones that fail quietly.
| Check | Time it takes |
|---|---|
| Real cost per trade | minutes |
| Warm up test (run it twice) | minutes |
| Trade count and range | minutes |
| Shuffle the trade order | minutes |
| Separate time blocks | hours |
| Count your versions | ongoing |
None of these throw an error. Nothing breaks, nothing warns you, the curve looks fine. It just stops matching live later, when you've already sized up. You can measure that gap before you risk real money, with forward testing and by tracking live vs backtest.
