Skip to content
← Blog

Four ways a backtest lies to you

Most backtests look better than reality for reasons that have nothing to do with the strategy. Here are the four that do the most damage, with real numbers from NSE data.

4 min read

A backtest is a claim about what would have happened. Most of them are wrong in the same four ways, and all four push the result in the same direction: too good.

These are not subtle modelling disagreements. They are large, and three of them are silent.

1. Costs

The most common and the most boring.

An intraday options strategy in India pays brokerage, STT, exchange transaction charges, GST, stamp duty and SEBI turnover fees. On a strategy trading every day, these compound into something that routinely decides whether the whole idea works.

STT in particular is asymmetric — it is charged on the sell side — which means a backtest that applies a flat percentage to both legs gets it wrong in both directions.

Then there is slippage, which is not a fee and is usually ignored entirely. If your backtest fills at the closing price, it assumes you got the last printed price of the day on every leg. You did not.

The fix is unglamorous: apply every charge separately at real rates, and assume you cross the spread. If a strategy only works with zero costs, it does not work.

2. Corporate actions

This one is silent, and it is the reason most stock-basket backtests are quietly wrong.

Exchange closing prices are not adjusted for splits and bonuses. When a stock does a 1:10 split, the printed price falls by 90% overnight. The holder lost nothing — they now own ten times as many shares — but a backtest chaining closing prices reads that day as a -90% catastrophe.

We measured this across 2.67 million rows of NSE equity data. Even computing returns the careful way — using each row’s own stated previous close rather than chaining closes — 368 days still showed moves beyond ±100%, across 358 different symbols. Ninety-seven of those sat within 6% of an exact 1:N ratio: unadjusted splits, hiding in plain sight.

One of those inside a basket makes the entire result meaningless, and it looks exactly like an ordinary bad day.

The honest fix is to detect them and divide them out. A real example from recent data: HDFC Bank on 26 August 2025 shows a -50% day. That was a 1:1 bonus. Divide the ratio out and the genuine move that day was -0.88% — an ordinary session.

Our basket tool does this and shows you every correction it made, with the date and the ratio, so you can check it against the company’s own announcement.

3. Survivorship

If you build a basket from today’s most-traded stocks and test it over five years, you have selected for companies that are still here and still liquid. The ones that were delisted, suspended, or collapsed are not in your list — because you picked the list after knowing how things turned out.

This is unavoidable to some degree with a fixed universe. What you can do is be honest about it: know that your list contains survivors, and treat any result as the best case.

A related trap is renaming. Zomato became ETERNAL; Shriram Transport Finance became Shriram Finance. Test from 2021 using the new symbol and you get a stock that appears not to have existed for the first few years. That is not a data error — the rows are under the old name — but a backtest that silently fills the gap with zeros or skips the symbol will mislead you either way.

4. Overfitting

The subtlest one, because it feels like work.

You test a strategy. It is mediocre. You try entering at 9:20 instead of 9:15. Better. You try a 35% stop instead of 30%. Better again. Twenty adjustments later you have something that looks excellent on history and has no reason to work tomorrow.

Nothing dishonest happened. You simply searched a large space of variations and kept the one that best fit the noise in the past.

The standard defence is out-of-sample testing — hold back recent data, tune on the rest, check once. The cheaper and surprisingly effective one is a robustness check: take your tuned strategy and nudge each parameter slightly. Shift entry by ten minutes. Move the stop by five points. Run it on each year separately.

A real edge degrades gracefully. A fitted one falls off a cliff. If your 9:20 entry works and 9:30 is a disaster, you did not find an edge in the time of day — you found a coincidence.

What a trustworthy backtest looks like

  • Every cost applied separately, at real rates, with slippage assumed
  • Corporate actions detected and reported, not absorbed
  • Symbols with unusable data excluded and named, not silently dropped
  • Results shown as a curve over time, not a single number
  • The same strategy tested across several years independently
  • Parameters nudged to see whether the result survives

The last point matters most. The single most useful question about any backtest is not how good is this but how fragile is it.


If you want to see these effects rather than read about them, the backtest tool applies every cost by default and lets you turn them off to see the difference, and the robustness checker runs every entry timing across every year at once.

Sources

Every factual claim above comes from one of these. Check them rather than taking our word for it.