Skip to content

CFA Level II Exam · Backtesting and Simulation

Backtesting Pitfalls and Biases for CFA Level II

Updated 7 October 2026 · Fact-checked

Backtesting applies a strategy or risk model to past data to see how it would have performed. Pitfalls make the result look better than real life: look-ahead bias, survivorship bias, data snooping, overfitting and ignored costs. To answer a question, find what the test used that was not available or realistic at the time.

Understand Backtesting Pitfalls and Biases

A backtest replays a strategy or model on historical data. You use it to judge whether an idea would have worked and how risky it would have been. The result is only useful if the test mimics what an investor could really have done at each date.

Most pitfalls come from using information or conditions that were not truly available. Look-ahead bias happens when the test uses data that was not public on the decision date. Examples: using year-end financial statements on the day the year ended, or using restated figures instead of the originally reported ones. Survivorship bias happens when the data set holds only firms or funds that still exist today. Failed ones are missing, so average returns look higher and risk looks lower.

Data snooping (data mining) means testing many variables or rules on the same data and reporting the best one. With enough tries, some rule will look significant by chance. Overfitting means the model has too many parameters or is tuned so tightly that it captures noise, not a lasting relationship. It fits the past well and fails on new data.

Other pitfalls are unrealistic transaction cost assumptions (ignoring commissions, bid-ask spreads, market impact, and the fact that large trades in illiquid assets move prices), ignoring taxes and borrowing costs, and using a period that is too short or covers only one market regime. Backtests of risk models such as VaR also depend on the sample window.

The main remedy is out-of-sample testing. The in-sample period is where you build and tune the model. The out-of-sample period is data held back and used once to check it. Related fixes: use point-in-time data, include delisted securities, limit the number of tests and disclose them, require an economic rationale, and apply realistic costs.

Key formulas to remember

Look-ahead bias
Test uses information dated after the decision date
Fix with point-in-time, as-originally-reported data.
Survivorship bias
Sample = only entities that survived to the end date
Overstates average return and understates risk. Fix by including dead and delisted entities.
Data snooping
Many tests on the same data → some pass by chance
Fix with out-of-sample testing, fewer tests, and an economic rationale.
Overfitting
High in-sample fit, weak out-of-sample fit
A large gap between the two signals overfitting.
Net return
Net return = gross return − trading costs − other costs
Costs include commissions, spreads and market impact. Turnover raises their effect.

How to solve Backtesting Pitfalls and Biases questions

Use this method on any vignette that describes how a backtest was built or what it reported.

  1. 1Read the vignette for how the data set, sample period, rules and costs were chosen.
  2. 2Ask at each step: could an investor really have known or done this on that date?
  3. 3If the data used was not yet available (restated, final or future figures), name look-ahead bias.
  4. 4If the sample contains only current members, funds or listed firms, name survivorship bias.
  5. 5If many variables or rules were tried and only the best is reported, name data snooping. If the model is complex and fits the past tightly, name overfitting.
  6. 6Check costs: are spreads, market impact, taxes and liquidity ignored? If so, returns are overstated.
  7. 7Pick the remedy that matches the flaw: point-in-time data, inclusion of failed entities, out-of-sample testing, or realistic costs.
  8. 8Check the direction. Almost all of these biases overstate performance.

Quickest way: Match the clue to the bias

When to use it: When the vignette is short on time and the answer choices are bias names or remedies.

  1. Clue 'data not yet published' or 'restated': look-ahead.
  2. Clue 'current constituents only' or 'funds still operating': survivorship.
  3. Clue 'tested many variations, chose the best': data snooping.
  4. Clue 'many parameters, great fit, poor new data': overfitting.
  5. Clue 'no mention of spreads or impact': unrealistic costs.
  6. Remedy is usually out-of-sample testing or point-in-time, survivor-free data.

Common mistakes in Backtesting Pitfalls and Biases

  • Confusing look-ahead bias with survivorship bias.

    Both make the past look better and both involve data.

    Fix: Look-ahead is about timing of information. Survivorship is about which entities are in the sample.

  • Treating data snooping and overfitting as the same thing.

    Both give strong in-sample results that fade later.

    Fix: Snooping is repeated testing and selecting the best. Overfitting is a model too closely tuned to noise. Match the vignette's wording.

  • Thinking survivorship bias lowers average returns.

    Students focus on 'missing data' and assume less return.

    Fix: Missing entities are the failures, so average returns are overstated and risk is understated.

  • Saying in-sample testing validates a model.

    A good fit feels like proof.

    Fix: Only out-of-sample results show whether the relationship holds on data not used to build the model.

  • Ignoring costs in high-turnover or illiquid strategies.

    Students focus on signal quality, not implementation.

    Fix: When turnover is high or assets are illiquid, expect net returns to fall well below gross returns.

Worked examples

Example 1

A researcher tests a value strategy on global equities from 2005 to 2020. She forms portfolios each 1 January using the prior year's annual earnings, although many firms published them in March. She uses the current index constituent list. Q1: Which bias comes from the earnings timing? Q2: Which bias comes from the constituent list? Q3: How does the second bias affect results?

Show the solution
  1. Q1: Earnings used on 1 January were not public until March. The test uses information not available at the decision date. That is look-ahead bias.
  2. Q2: Using today's constituents drops firms that were delisted or went bankrupt during the period. That is survivorship bias.
  3. Q3: Failed firms are excluded, so average returns are overstated and risk understated.

Answer: Q1: look-ahead bias. Q2: survivorship bias. Q3: it overstates returns and understates risk.

Example 2

An analyst tests 200 technical rules on 10 years of data and reports the best one, with a Sharpe ratio of 1.8. She then runs it on a later 3-year period not used before; the Sharpe ratio is 0.2. Her model assumes zero trading costs, though the rule trades daily in small-cap stocks. Q1: What explains the gap? Q2: What does the later test represent? Q3: What is the effect of the cost assumption?

Show the solution
  1. Q1: Testing 200 rules and keeping the best means some would look strong by chance. That is data snooping, and the fall in performance is consistent with it.
  2. Q2: The 3-year period was not used to select or tune the rule, so it is an out-of-sample test.
  3. Q3: Daily trading in small caps means high turnover, wide spreads and market impact. Assuming zero cost overstates net returns.

Answer: Q1: data snooping (selecting the best of many tests). Q2: out-of-sample test. Q3: returns are overstated, so net performance would be lower than reported.

Exam tips

  • Read the vignette for timing words such as 'restated', 'final' and 'as of year-end'. They often signal look-ahead bias.
  • Name the direction: nearly every bias here overstates return or understates risk.
  • When asked for a remedy, tie it to the flaw: point-in-time data, including delisted entities, out-of-sample testing, or realistic costs.
  • A big gap between in-sample and out-of-sample results points to overfitting or data snooping, not to a bad out-of-sample period.
  • If a choice says costs can be ignored for a backtest, it is almost always wrong.

Backtesting Pitfalls and Biases: frequently asked questions

What is the difference between look-ahead bias and survivorship bias?

Look-ahead bias uses information that was not available at the decision date. Survivorship bias uses a sample that leaves out entities that failed or were delisted. One is about timing and the other is about sample selection.

What is the difference between in-sample and out-of-sample testing?

In-sample data is used to build and tune the model. Out-of-sample data is held back and used to check whether the model works on new data. Out-of-sample results are a better guide to future performance.

How do data snooping and overfitting differ?

Data snooping is running many tests on the same data and reporting the best, so a chance result looks real. Overfitting is building a model so closely matched to past data that it fits noise. Both show strong in-sample and weak out-of-sample results.

How can you avoid backtesting biases?

Use point-in-time data, include delisted and failed entities, and test on out-of-sample data. Limit and disclose the number of tests, require an economic reason for the signal, and include realistic transaction costs.