Skip to content

IAI Actuarial Core Principles · Risk Modelling and Survival Analysis · Elementary principles of machine learning

An analyst uses k-fold cross-validation with k = 5 on a dataset of 1,000 observations to compare models. Which description of the procedure is correct?

In 5-fold cross-validation the data are divided into five parts. Each part is used once for validation while the model trains on the remaining four, and the five validation errors are averaged to estimate out-of-sample performance.

  1. AThe data are split into 5 parts; each part is used once as the validation set while the model is trained on the other 4 parts, and the 5 errors are averagedCorrect
  2. BThe model is trained on 5 observations and tested on the other 995
  3. CThe model is trained once on 800 observations and tested once on the remaining 200 only
  4. DEach observation is used for validation 5 times and the best result is chosen
  5. The data are split into 5 parts and the model is trained on each part separately and tested on the full data

Explanation

In 5-fold cross-validation, each fold of 200 observations serves as the validation set once, with training on the other 800. The five validation errors are averaged to estimate test error. A single 800/200 split is a hold-out method, not cross-validation.

Did you get it right without looking?

One question tells you little. A timed set on Elementary principles of machine learning shows your real accuracy, how long you take and where you lose marks.

More Elementary principles of machine learning questions