IAI Actuarial Core Principles · Risk Modelling and Survival Analysis · Elementary principles of machine learning
A life insurer in Mumbai has data on 20,000 customers (age, premium paid, number of policies, channel) but no outcome variable. An analyst applies k-means to find groups of similar customers. Which statement about this exercise is correct?
It is unsupervised learning. No outcome variable is supplied, and k-means groups customers by similarity of their features, discovering the structure from the data itself. Since the groups are not known labels beforehand and there is no reward signal, it is neither supervised nor reinforcement learning.
- AIt is unsupervised learning because no target variable is provided and the aim is to find structure in the inputsCorrect
- BIt is supervised learning because the groups found act as labels known in advance
- CIt is reinforcement learning because the algorithm improves through repeated rewards
- DIt is supervised regression because the cluster centres are continuous
- It cannot be machine learning because it does not predict anything
Explanation
K-means searches for structure in the inputs without any response variable, so it is unsupervised learning. The clusters are discovered, not given beforehand, so the supervised option is wrong. There is no reward signal, so it is not reinforcement learning.
Did you get it right without looking?
One question tells you little. A timed set on Elementary principles of machine learning shows your real accuracy, how long you take and where you lose marks.
More Elementary principles of machine learning questions
- Which statement about the K-means algorithm is correct?
- When building a predictive model for insurance data, which action best reduces the risk of data leakage?
- A lasso (L1-penalised) regression is fitted to predict lapse rates and the penalty parameter lambda is increased from a small value to a ver…
- A lapse-prediction model is fitted to 1,000 policies. Of the 100 policies that actually lapsed, the model flags 70 as lapses. Of the 900 pol…
- A data scientist at a Mumbai insurer standardises all predictors using the mean and standard deviation of the full dataset, then splits the …
- A regression model is evaluated on a test set of 4 observations with actual values 10, 12, 14, 20 and predicted values 11, 10, 14, 16. What …