Why Backtests Can Mislead You
This site exists to run backtests, so it is worth being direct about what they cannot do. A backtest is an arithmetic statement about one path history happened to take. Most of the ways people misread them come from treating that as a forecast.
#The window determines the answer
This is the largest problem and the easiest to fall into, because it usually is not deliberate. You pick a round number of years, or the period you have been investing, and the result you get is substantially determined by that choice.
The comparisons on this site are built to show it. QQQ beat the broad market by roughly five percentage points a year over the last decade. The same two strategies over 2000 to 2010 produced an 83 percent drawdown and a decade of almost no return. International lagged the US badly over the last ten years and beat it by nearly eight points a year from 2003 to 2008.
#Survivorship bias in the ticker you chose
Backtesting an individual stock has a problem no calculation can fix: you selected it knowing how it turned out.
The NVDA comparison on this site exists because NVDA won. Nobody builds a page for a semiconductor company that stagnated, and nobody searches for one. The comparison against Intelmakes the point concretely: in 2015 Intel was the larger, safer, more widely recommended of the two, and the plan that chose it ended with less than a thirtieth of the other’s final value.
The honest question is never “what would this stock have returned.” It is “what would the stock I would actually have chosen at the time have returned,” and that set includes all the reasonable-looking choices that went nowhere. Across all individual US stocks over long horizons, the median outcome has historically trailed the index, with aggregate market returns driven by a small minority of extreme winners.
Index funds largely sidestep this, but not entirely. Indexes themselves drop failing constituents and add successful ones, so an index’s long-run record reflects an ongoing selection process rather than a fixed basket held throughout.
#Funds that did not exist yet
You cannot test a fund through a period before it launched, and the launch dates are not randomly distributed. Products tend to be created after a strategy has already performed well, which means their track records disproportionately begin at favorable moments.
- TQQQ launched in February 2010, near the start of a long bull market, and has never operated through a multi-year bear market. See what leveraged ETFs would have done in a real bear market.
- SCHD launched in 2011 and cannot be tested against 2008, which is why the crisis comparison on this site uses VYM instead.
- VXUS launched in 2011, so any international comparison reaching into the 2000s has to substitute EFA.
- SPMO launched in late 2015, so no window including it covers either major crash.
When a backtest silently starts later than you asked, this site flags it as insufficient history and reports the actual inception date rather than quietly producing a shorter run that looks like a complete one.
#What this simulation assumes
Every backtest embeds assumptions. Here are the ones behind the numbers on this site, stated plainly and covered in more detail on How It Works.
- No taxes. Every figure is pre-tax. For dividend-heavy funds in a taxable account over long periods, this is a substantial omission.
- No commissions or spreads. Modern brokerages charge nothing for ETF trades, so this is close to accurate today and less so for older windows.
- Perfect discipline. The simulated investor contributes every single period without fail, including through 2008 and 2020. Real investors frequently stop contributing at exactly the wrong moment.
- Fractional shares.Contributions are invested in full at the day’s closing price. Not every brokerage supports this, though most now do.
- The data is what it is. Prices come from Yahoo Finance, which is generally reliable but occasionally has gaps or errors, particularly for older or less liquid tickers.
#What a backtest is genuinely good for
The point is not that this is all useless. It is that the useful questions are different from the ones people usually ask.
- Understanding mechanisms. Why a contribution plan behaves differently from a lump sum, why volatility drag punishes leveraged funds, why reinvestment compounds. These are structural and transfer across periods.
- Calibrating expectations about risk. Seeing that a broad index fell 55 percent in 2008 and 34 percent in 2020 tells you something real about what holding equities involves.
- Testing a claim you have been told. If you have heard dividend funds are defensive, running them through 2008 is a legitimate test, and it does not support the claim.
- Comparing across many windows. The spread of outcomes across different start dates is far more informative than any single run.
#How to use this tool honestly
Run more than one window. If a conclusion survives 2000 to 2010, 2007 to 2012, and 2016 to today, it is probably about the strategy. If it only appears in the window you happened to pick first, it is about the window.
Look at the drawdown figures before the return figures. A return you would not have held through is not a return you would have received, and the drawdown column is the better predictor of whether you would have stayed invested.
And treat any result that looks spectacular with more suspicion than one that looks ordinary. Spectacular results usually mean the window, the ticker, or both were selected after the fact, which is exactly the condition under which a backtest tells you least.