Skip to content

FRM Part II · FRM Exam Part II · Advances in Artificial Intelligence: Implications for Capital Markets Activities

A risk team splits 10 years of monthly data randomly into training (70%) and test (30%) sets to evaluate a machine learning model forecasting volatility. The test results are excellent, yet live performance is poor. Which is the most likely methodological flaw?

The likely flaw is look-ahead leakage from randomly splitting serially correlated time-series data. Neighboring observations appear in both training and test sets, inflating test performance. Chronological or walk-forward validation avoids this and better reflects live conditions.

  1. AThe test set was too large relative to the number of features
  2. BRandom splitting of serially correlated time series leaks information from the future into training, inflating test performanceCorrect
  3. CUsing monthly rather than daily data always causes underfitting
  4. DThe model used regularization, which cannot be validated out of sample

Explanation

With time series, random splits place observations adjacent in time in both sets, so autocorrelation and regime information leak. Test results become optimistic. Proper validation uses chronological or walk-forward splits. Regularization can be validated out of sample.

Did you get it right without looking?

One question tells you little. A timed set on Advances in Artificial Intelligence: Implications for Capital Markets Activities shows your real accuracy, how long you take and where you lose marks.

More Advances in Artificial Intelligence: Implications for Capital Markets Activities questions