Skip to content

CFA Level I · CFA Level I Exam · Introduction to Financial Data Science

A model achieves very low error on its training dataset but produces much larger errors on a validation dataset. The model is most likely suffering from:

The model most likely suffers from overfitting. It has fitted noise and idiosyncrasies in the training data, so it performs well in-sample but generalizes poorly to new data. Underfitting would show high error on both training and validation sets, not only on validation.

  1. Aunderfitting, because it is too simple
  2. Boverfitting, because it has learned noise in the training dataCorrect
  3. Cdata leakage, because the validation labels were removed

Explanation

Low training error combined with high out-of-sample error is the classic sign of overfitting, where the model captures noise specific to the training sample. Underfitting would produce high error on both datasets. Removing labels from validation data is not leakage and does not explain the pattern.

Did you get it right without looking?

One question tells you little. A timed set on Introduction to Financial Data Science shows your real accuracy, how long you take and where you lose marks.

More Introduction to Financial Data Science questions