Skip to content

CS Professional · Artificial Intelligence, Data Analytics and Cyber Security - Laws and Practice · Data Analytics

A Chennai insurer built a claims-prediction model that scores 97% accuracy on the data used to train it but only 71% on fresh unseen claims. What is the most likely problem, and what lifecycle step exposes it?

The likely problem is overfitting. The model learned the training data too closely and fails to generalise, which is revealed when evaluating it on separate unseen test data during the validation stage. Underfitting would show weak performance on both datasets.

  1. AOverfitting, revealed by evaluation on a separate test datasetCorrect
  2. BUnderfitting, revealed by data collection
  3. CData leakage fixed by deployment monitoring only
  4. DSampling bias, revealed by the visualisation stage

Explanation

High training accuracy with much lower accuracy on unseen data is the classic sign of overfitting: the model memorised training patterns. It is detected during model evaluation or validation using held-out test data. Underfitting would give poor results on both sets.

Did you get it right without looking?

One question tells you little. A timed set on Data Analytics shows your real accuracy, how long you take and where you lose marks.

More Data Analytics questions