Skip to content

FRM Exam Part I · Machine-Learning Methods

Overfitting, Bias-Variance Tradeoff and Model Validation for FRM Part 1

Updated 11 October 2026 · Fact-checked

Overfitting means a model fits noise in the training data and predicts new data poorly. Underfitting means it is too simple to capture the pattern. The bias-variance tradeoff balances the two. You tune complexity on validation data or with k-fold cross-validation, then judge final performance once on a test set.

Understand Overfitting, Bias-Variance Tradeoff and Model Validation

A machine-learning model is built to predict new data, not to describe the data it was trained on. That one idea drives everything in this topic.

Overfitting happens when a model is too flexible. It learns the noise as well as the signal. Training error is very low, but error on unseen data is high. Underfitting happens when a model is too simple. It misses real structure, so error is high on both training and new data.

Prediction error on new data splits into three parts: bias squared, variance and irreducible noise. Bias is error from wrong or too-simple assumptions. Variance is how much the fitted model changes if you train it on a different sample. Simple models have high bias and low variance. Complex models have low bias and high variance. Adding complexity cuts bias but raises variance. The best model sits where total error on new data is lowest. This is the bias-variance tradeoff.

To find that point you need honest estimates of error on unseen data. So you split the data. The training set fits the parameters. The validation set compares models and tunes hyperparameters (such as the regularization strength or tree depth). The test set is kept aside and used once at the end to estimate real-world performance. If you tune on the test set, it is no longer unseen and your estimate is too optimistic.

When data is scarce, use k-fold cross-validation. Split the data into k equal folds. Train on k − 1 folds, validate on the remaining one, and repeat k times so each fold is the validation set once. Average the k errors. In finance, with time-ordered data, random shuffling can leak future information into training. Use splits that respect time order.

Key formulas to remember

Expected prediction error decomposition
Expected error = Bias² + Variance + Irreducible error
Irreducible error is noise no model can remove. Only bias and variance change with model choice.
Cross-validation error
CV error = (1 ÷ k) × Σ (error on fold i), i = 1 to k
Each fold is the validation set once. This is a simple average when folds are equal in size.
Observations per fold
Fold size = n ÷ k
Each round trains on n × (k − 1) ÷ k observations.
Complexity pattern
Complexity ↑ → Bias ↓, Variance ↑, training error ↓
Validation error is U-shaped: it falls, then rises once overfitting starts.
Data roles
Training = fit; Validation = tune/select; Test = final unbiased estimate
Use the test set once. Using it to tune makes it part of training.

How to solve Overfitting, Bias-Variance Tradeoff and Model Validation questions

Use this routine for any question on fit, bias-variance or validation.

  1. 1Identify what is being asked: diagnose fit, name the data set, or calculate a cross-validation figure.
  2. 2Compare training error with validation (or test) error. Low training and much higher validation error signals overfitting. High error on both signals underfitting.
  3. 3Link the diagnosis to bias and variance. Overfit means low bias, high variance. Underfit means high bias, low variance.
  4. 4Pick the remedy. For overfitting: simplify, regularize, add data, prune or stop early. For underfitting: add features or complexity, or reduce regularization.
  5. 5Check which data set is used for which job: training fits, validation tunes, test evaluates once.
  6. 6For k-fold questions, compute the average of the k fold errors and the training size per round, n × (k − 1) ÷ k.
  7. 7Check for leakage, such as tuning on the test set or shuffling time-series data, and state its effect: optimistic error estimates.

Quickest way: Two-number gap check

When to use it: Use when a question gives training and validation errors and asks for a diagnosis or the best model.

  1. Write training error and validation error for each model.
  2. Pick the model with the lowest validation error, not the lowest training error.
  3. If both errors are high and close, call it underfitting (high bias).
  4. If training error is low and validation error is much higher, call it overfitting (high variance).
  5. For cross-validation, add the fold errors and divide by k.

Common mistakes in Overfitting, Bias-Variance Tradeoff and Model Validation

  • Choosing the model with the lowest training error.

    Training error always falls as complexity rises, so it looks like progress.

    Fix: Select on validation or cross-validated error. Training error says nothing about generalization.

  • Mixing up bias and variance for overfitting.

    Both words sound like 'error', and complex models feel more accurate.

    Fix: Overfitting is high variance and low bias. Underfitting is high bias and low variance.

  • Using the test set to tune hyperparameters.

    It seems efficient to check several settings on the same held-out data.

    Fix: Tune on the validation set or by cross-validation. Touch the test set once for the final estimate.

  • Believing more data always fixes underfitting.

    More data helps variance, so students assume it helps everything.

    Fix: More data mainly reduces overfitting. An underfit model needs more flexibility or better features.

  • Shuffling time-series data randomly before splitting.

    Standard k-fold examples use random folds.

    Fix: Keep chronological order, training on the past and validating on the future, to avoid look-ahead leakage.

  • Miscounting the training size in k-fold.

    Students use n ÷ k, which is the validation fold size.

    Fix: Each round trains on n × (k − 1) ÷ k observations and validates on n ÷ k.

Worked examples

Example 1

A risk team fits three credit-default models. Training and validation errors (misclassification rate) are: Model A 18% and 19%; Model B 6% and 11%; Model C 1% and 17%. Which model should be chosen, and which one is overfit?

Show the solution
  1. Select on validation error: A = 19%, B = 11%, C = 17%. The lowest is Model B.
  2. Check the gap between validation and training error: A = 1 point, B = 5 points, C = 16 points.
  3. Model C has very low training error but much higher validation error. That is overfitting (low bias, high variance).
  4. Model A has high error on both sets, consistent with underfitting (high bias).

Answer: Choose Model B (validation error 11%). Model C is overfit.

Example 2

A dataset has 1,000 observations and you run 5-fold cross-validation. The validation errors (mean squared error) are 4.0, 5.0, 3.5, 4.5 and 6.0. How many observations are used to train in each round, and what is the cross-validation error?

Show the solution
  1. Fold size = 1,000 ÷ 5 = 200 observations.
  2. Training size per round = 1,000 × (5 − 1) ÷ 5 = 800 observations.
  3. Sum of errors = 4.0 + 5.0 + 3.5 + 4.5 + 6.0 = 23.0.
  4. CV error = 23.0 ÷ 5 = 4.6.

Answer: 800 training observations per round; cross-validation MSE = 4.6.

Exam tips

  • Questions usually give a table of training and validation errors. Decide on validation error first, then diagnose.
  • Learn the one-line mapping: overfit = high variance, underfit = high bias. Many MCQs test only this.
  • When an option says 'evaluate on the test set repeatedly to tune', it is almost always the wrong choice.
  • For k-fold, do the arithmetic: fold size n ÷ k, training size n × (k − 1) ÷ k, then average the errors.
  • Watch for time-series wording. The correct answer will preserve time order and avoid look-ahead bias.

Practice questions from Machine-Learning Methods

Overfitting, Bias-Variance Tradeoff and Model Validation in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Overfitting, Bias-Variance Tradeoff and Model Validation: frequently asked questions

What is the difference between overfitting and underfitting?

An overfit model matches the training data too closely, including noise, so it does badly on new data. An underfit model is too simple to capture the pattern, so it does badly on both training and new data. Overfitting is high variance. Underfitting is high bias.

What is the difference between training, validation and test sets?

The training set is used to fit the model parameters. The validation set is used to compare models and tune hyperparameters. The test set is held back and used once at the end to estimate how the final model will perform on unseen data.

How does k-fold cross-validation work?

You split the data into k equal folds. In each of k rounds you train on k − 1 folds and validate on the one left out. You then average the k validation errors. It uses data efficiently when samples are small.

Can a model have both low bias and low variance?

Only up to the limit set by irreducible noise. In practice, reducing one tends to raise the other, so you aim for the complexity that minimizes total error on unseen data.