Skip to content

IAI Actuarial Core Principles · Risk Modelling and Survival Analysis · Elementary principles of machine learning

A dataset is split into training, validation and test sets to choose the number of neighbours k in a k-nearest-neighbours model. What is the correct use of the test set?

The test set should be used only once, after k has been chosen on the validation set, to give an unbiased estimate of performance on unseen data. Using it to choose k would leak information and make the estimate optimistically biased.

  1. AChoosing the value of k that minimises test error
  2. BFitting the model parameters alongside the training set
  3. CTuning k and then retraining on the test set
  4. DProviding a final unbiased estimate of performance after k is chosen using the validation setCorrect
  5. Replacing the validation set whenever it is too small

Explanation

Hyperparameters such as k are tuned on the validation set. The test set must remain untouched until the end so that it gives an unbiased estimate of out-of-sample performance. Using it to select k would make the estimate optimistic.

Did you get it right without looking?

One question tells you little. A timed set on Elementary principles of machine learning shows your real accuracy, how long you take and where you lose marks.

More Elementary principles of machine learning questions