Skip to content

CFA Level II Exam · Backtesting and Simulation

Limitations of Backtesting and Simulation: Interpreting Results

Updated 7 October 2026 · Fact-checked

Backtesting applies a strategy or risk model to past data; simulation generates many possible future paths from assumed inputs. Both are only as good as their data and assumptions. To interpret results, check sample size, look-ahead and survivorship bias, overfitting, out-of-sample performance, and sensitivity to inputs. Then judge whether the result is robust.

Understand Interpreting Results and Limitations

A backtest asks what would have happened if you had used a rule in the past. A simulation asks what could happen in the future if inputs behave as you assume. Both produce numbers that look precise. The exam tests whether you can see how far those numbers can be trusted.

A backtest uses one historical path. That path may not repeat. Results depend on the period chosen, the data quality and how many rules were tried before one worked. If you test many rules on the same data, some will look good by chance. This is data mining or overfitting. A rule tuned to past noise usually fails on new data.

Common backtest biases include look-ahead bias (using information not available at the time), survivorship bias (using only firms that still exist, which inflates returns), data snooping, and ignoring transaction costs, taxes, liquidity and market impact. A regime change can also make the past a poor guide.

Simulation results depend on the distribution assumed, the correlations and parameters estimated, and the number of trials. Garbage in, garbage out. Parameters estimated from a short or calm period understate tail risk. Assuming normality hides fat tails. Assuming constant correlations hides the fact that correlations rise in a crisis. Simulation also gives no insight into cause: it shows outcomes, not why they happen. More trials reduce sampling error but do not fix a wrong model.

Model risk is the risk of loss from using a wrong, misspecified or misused model. It comes from bad inputs, wrong structure, coding errors and use outside the model's intended purpose. Good practice is to test robustness: use out-of-sample and walk-forward tests, vary parameters (sensitivity analysis), run stress tests and scenario analysis, and compare against other models.

Key formulas to remember

Out-of-sample test
Estimate on in-sample data, then evaluate on data not used in fitting
A large drop in performance out of sample signals overfitting.
Simulation standard error (rule of thumb)
Standard error of the estimated mean ≈ s ÷ √N
s is the sample standard deviation of outcomes, N is the number of trials. To halve the error you need about four times as many trials. This covers sampling error only, not model error.
Expected number of VaR exceptions
Expected exceptions = (1 − confidence level) × number of observations
Compare actual exceptions with this figure when backtesting VaR. Far more exceptions suggest risk is understated.

How to solve Interpreting Results and Limitations questions

Use this sequence for any vignette asking you to judge a backtest or simulation.

  1. 1Identify the method: backtest (historical path) or simulation (assumed inputs).
  2. 2Find the data facts in the vignette: sample period, number of observations, how rules were chosen, whether delisted firms are included.
  3. 3Check for each bias in turn: look-ahead, survivorship, data snooping, overfitting, ignored costs.
  4. 4Compare in-sample and out-of-sample results if both are given. A big gap points to overfitting.
  5. 5For simulation, identify assumed distribution, correlations, parameters and number of trials, and ask which is unrealistic.
  6. 6Decide whether the problem is sampling error (fixed with more trials) or model error (not fixed with more trials).
  7. 7Choose the remedy or conclusion that matches the specific flaw: out-of-sample test, sensitivity analysis, stress test, add costs, or use point-in-time data.

Quickest way: Flaw-to-fix matching

When to use it: When time is short and the question asks for the main limitation or best improvement.

  1. Underline what the vignette says about data and assumptions.
  2. Match the clue: delisted firms missing means survivorship; future data used means look-ahead; many rules tried means data snooping.
  3. Pick the fix tied to that flaw, not a generic answer.
  4. Remember that more simulation trials never cure a wrong model.

Common mistakes in Interpreting Results and Limitations

  • Believing a high backtest return proves the strategy works.

    The number looks precise and the result is in-sample.

    Fix: Ask whether it was tested out of sample and how many rules were tried to find it.

  • Thinking more Monte Carlo trials fix a bad model.

    Confusing sampling error with model error.

    Fix: More trials only reduce random error. Wrong distributions or correlations remain wrong.

  • Mixing up look-ahead and survivorship bias.

    Both inflate results and both involve data problems.

    Fix: Look-ahead uses information not yet known. Survivorship leaves out failed firms.

  • Ignoring transaction costs and liquidity in backtests.

    Candidates focus on statistical issues only.

    Fix: Check whether costs, taxes and market impact were deducted. High-turnover strategies are most affected.

  • Assuming simulation results show likelihood of real events.

    Outputs are presented as probabilities.

    Fix: The probabilities hold only if the assumed inputs are right. Treat them as conditional on assumptions.

Worked examples

Example 1

A quant team tests 200 trading rules on 10 years of data for firms in a current index. The best rule earns 14% a year against 9% for the index. Returns are before costs. Q1: Name two biases. Q2: What test best checks the result?

Show the solution
  1. Firms in the current index are survivors, so survivorship bias is present.
  2. Testing 200 rules on the same data and picking the best is data snooping, which creates overfitting risk.
  3. Costs are ignored, which overstates net returns.
  4. The best check is an out-of-sample test on data not used to choose the rule.

Answer: Q1: Survivorship bias and data snooping (also ignored costs). Q2: An out-of-sample (walk-forward) test.

Example 2

A risk analyst runs a 10,000-trial Monte Carlo simulation of a portfolio using a normal distribution and correlations estimated from a calm two-year period. The 1% VaR is reported as ₹4 crore. Q1: Is the VaR likely understated or overstated? Q2: Would 100,000 trials fix the main problem?

Show the solution
  1. Normality ignores fat tails, so extreme losses are underestimated.
  2. Correlations from a calm period are usually lower than in a crisis, so diversification is overstated.
  3. Both flaws push VaR down, so it is likely understated.
  4. More trials reduce sampling error only. The assumptions stay wrong, so the bias remains.
  5. A better fix is stress tests, a fat-tailed distribution or crisis-period correlations.

Answer: Q1: Likely understated. Q2: No. More trials reduce sampling error but do not correct wrong distribution or correlation assumptions.

Exam tips

  • Read the vignette for clues on how data was chosen. Each clue usually maps to one named bias.
  • When both in-sample and out-of-sample results appear, compare them first.
  • For simulation questions, separate sampling error from model error before choosing an answer.
  • Prefer the answer that fixes the stated flaw, not the one that sounds most thorough.
  • There is no penalty for wrong answers, so always answer every question.

Interpreting Results and Limitations in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Interpreting Results and Limitations: frequently asked questions

What is the main limitation of backtesting?

It uses one historical path, so past results may not repeat. Biases such as look-ahead, survivorship and data snooping can also inflate results. Ignored costs make it worse.

How do you check whether a backtested strategy is robust?

Test it out of sample, run walk-forward tests, vary parameters, and stress it under different conditions. A strategy that holds up across these checks is more credible. Large drops in performance suggest overfitting.

What is model risk in simulation?

It is the risk of loss from a wrong, misspecified or misused model. Causes include bad inputs, wrong distribution assumptions and coding errors. More trials do not remove it.

What is the difference between historical and Monte Carlo simulation limits?

Historical simulation is limited to events in the sample period. Monte Carlo can create new scenarios but depends on assumed distributions and parameters. Both can understate tail risk.