Risk Modelling and Survival Analysis · Elementary principles of machine learning
Supervised Learning: Regression and Classification Explained
Updated 11 October 2026 · Fact-checked
Supervised learning fits a model to data where the correct outcome is known, then uses it to predict new cases. Regression predicts a continuous outcome, such as a claim amount. Classification predicts a category, such as claim or no claim. To solve questions, identify the outcome type, pick the method, apply it and check performance.
Understand Supervised Learning: Regression and Classification
In supervised learning you have a set of records. Each record has features (also called predictors or inputs) and a known target (the response or label). The algorithm learns the link between features and target from this training data. You then use the link to predict the target for new records where it is unknown.
The type of target decides the problem type. If the target is a number on a continuous scale, such as a claim size or a premium, it is regression. If the target is a category, such as lapse or no lapse, or low, medium or high risk, it is classification. The same method family can often do both, so always look at the target first.
Four methods matter for the exam. A linear model predicts the target as a weighted sum of features. k-nearest neighbours (k-NN) predicts from the k most similar training records: it takes the majority class for classification, or the average for regression. A decision tree splits the data again and again using simple rules on features, so that each final group (a leaf) is as pure as possible. Naive Bayes uses Bayes' theorem to find the probability of each class given the features, and assumes the features are independent given the class.
Each method has trade-offs. Linear models are simple and easy to explain, but assume a fixed form. k-NN makes no assumption about form, but needs features on comparable scales and is slow on large data. Trees are easy to read and handle mixed data, but can overfit if grown too deep. Naive Bayes is fast and works with little data, but the independence assumption is often false.
In actuarial work, these tools are used for claim severity prediction, fraud flags, lapse prediction and risk grouping. Whatever you build, you judge it on data it has not seen. That links to training, validation and overfitting.
Key rules to remember
- Linear regression model
- y = β0 + β1x1 + β2x2 + ... + βpxp + ε
- Used when the target is continuous. Coefficients are usually fitted by least squares, minimising Σ(yi − ŷi)².
- Euclidean distance (k-NN)
- d(a, b) = √[Σ (aj − bj)²]
- Sum over all features j. Scale features first, or large-valued features dominate.
- k-NN prediction
- Classification: majority class among the k nearest. Regression: ŷ = (1/k) Σ yi over the k nearest
- Small k gives a flexible, noisy fit. Large k gives a smoother fit. Use an odd k for two classes to avoid ties.
- Gini impurity
- G = 1 − Σ pk²
- pk is the proportion of class k in the node. G = 0 means a pure node.
- Entropy
- H = − Σ pk log2(pk)
- Take 0 × log 0 as 0. Information gain is the parent's impurity minus the weighted average impurity of the child nodes.
- Naive Bayes
- P(C | x1,...,xp) ∝ P(C) × Π P(xj | C)
- Assumes features are independent given the class. Compare the product for each class and choose the largest. Divide by their sum to get probabilities.
- Regression tree split criterion
- Choose the split that minimises Σ (yi − ȳleft)² + Σ (yi − ȳright)²
- The leaf prediction is the mean of the target in that leaf.
How to solve Supervised Learning: Regression and Classification questions
Use this order for any supervised learning question.
- 1Read the target variable. Continuous means regression. Categorical means classification.
- 2List the features and note their types (numeric or categorical) and any scale differences.
- 3Name the method asked for, or choose one and justify it from the data and the aim (accuracy or interpretability).
- 4Apply the method with the formula written out: distances for k-NN, impurity for trees, the product of probabilities for naive Bayes.
- 5Write the prediction clearly, with the class or value and the working behind it.
- 6Comment on assumptions and limits, such as independence, scaling, choice of k or tree depth.
- 7Say how you would test it: hold-out or validation data, and a suitable error measure.
Quickest way: Fast route for calculation questions
When to use it: Use this when a small data table is given and you must predict or split in a few minutes.
- Circle the target and decide regression or classification in the first ten seconds.
- For k-NN, compute squared distances only. You do not need the square root to rank neighbours.
- For naive Bayes, compute the unnormalised product for each class. Normalise only if probabilities are asked for.
- For a tree split, compute the weighted impurity of the children and compare it with the others. Lower is better.
- Write the final answer in one line, then add one line on an assumption.
Common mistakes in Supervised Learning: Regression and Classification
Calling a problem regression because it uses the word 'probability'.
Logistic-type models output a number between 0 and 1, so students think the target is continuous.
Fix: Look at the target, not the output. If the target is a category, it is classification, even when the model outputs probabilities.
Using k-NN without scaling the features.
Students go straight to the distance formula.
Fix: Rescale first, for example to a common range or to standard scores. Otherwise a feature such as income in rupees swamps one such as age in years.
Choosing the tree split with the highest impurity.
Students compare the wrong direction or forget to weight the children by size.
Fix: Compute the weighted average impurity of the children and pick the split with the lowest value, which gives the greatest information gain.
Forgetting the prior P(C) in naive Bayes.
Students multiply only the feature likelihoods.
Fix: Always start with P(C). Then multiply by each P(xj | C) and compare across classes.
Judging a model on the training data only.
A deep tree or a k-NN with k = 1 fits the training data perfectly, which looks impressive.
Fix: State that performance must be measured on unseen validation or test data, because a perfect training fit usually signals overfitting.
Claiming naive Bayes needs the features to be truly independent.
The word 'naive' is misread.
Fix: Say it assumes conditional independence given the class. It often still performs reasonably when this does not hold exactly.
Worked examples
Example 1
A k-NN classifier with k = 3 predicts whether a policyholder lapses (L) or stays (S). Scaled training points are (age, premium): A (0.2, 0.3) S; B (0.4, 0.4) L; C (0.5, 0.5) L; D (0.9, 0.8) S; E (0.3, 0.2) S. Classify a new policyholder at (0.4, 0.3) using Euclidean distance.
Show the solution
- Compute squared distances to (0.4, 0.3).
- A: (0.2)² + 0² = 0.04.
- B: 0² + (0.1)² = 0.01.
- C: (0.1)² + (0.2)² = 0.01 + 0.04 = 0.05.
- D: (0.5)² + (0.5)² = 0.50.
- E: (0.1)² + (0.1)² = 0.02.
- Rank: B 0.01, E 0.02, A 0.04, C 0.05, D 0.50.
- The 3 nearest are B (L), E (S) and A (S).
- Majority: S has 2 votes and L has 1.
Answer: The new policyholder is classified as S (stays).
Example 2
A naive Bayes model predicts whether a claim is fraudulent (F) or genuine (G). Prior: P(F) = 0.1, P(G) = 0.9. A claim is reported late and is above ₹1,00,000. P(late | F) = 0.6, P(late | G) = 0.2, P(high | F) = 0.7, P(high | G) = 0.1. Find P(F | late, high) and classify the claim.
Show the solution
- Assume late and high are independent given the class.
- Fraud score: 0.1 × 0.6 × 0.7 = 0.042.
- Genuine score: 0.9 × 0.2 × 0.1 = 0.018.
- Sum = 0.042 + 0.018 = 0.060.
- P(F | late, high) = 0.042 ÷ 0.060 = 0.70.
- Since 0.70 > 0.30, the fraud class has the larger score.
Answer: P(F | late, high) = 0.70, so classify the claim as fraudulent (subject to the independence assumption).
Exam tips
- Start every answer by naming the target and stating regression or classification. This earns easy marks.
- Show the formula before the numbers. Markers give method marks even if arithmetic slips.
- State assumptions explicitly: conditional independence for naive Bayes, scaling for k-NN, stopping rule for trees.
- In discussion questions, compare methods on interpretability, flexibility and overfitting risk, and tie the points to the data in the question.
- Recent Paper A papers open with multiple-choice questions, so be ready for quick definition and distance or probability calculations.
Practice questions from Elementary principles of machine learning
- A model predicting whether a policy lapses is tested on 1,000 policies, of which 100 actually lapse. The model predicts lapse for 80 policie…
- A data scientist fits a very flexible model to a claims dataset. It achieves a very low error on the training data but a much higher error o…
- A analyst applies K-means clustering to policyholder data containing annual premium (in rupees, values in the lakhs) and age (in years, valu…
- In 5-fold cross-validation on 500 observations used to compare two regularised regression models, which statement is correct?
- A model fitted to a training set gives a very low training error but a much higher error on a separate test set. Which is the most likely ex…
Supervised Learning: Regression and Classification in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Supervised Learning: Regression and Classification: frequently asked questions
What is the difference between regression and classification?
Regression predicts a continuous number, such as a claim amount. Classification predicts a category, such as fraud or not fraud. Decide by looking at the target variable, not at the method.
How do you build a decision tree classifier?
Start with all the data in one node. Try each possible split and choose the one that lowers the weighted impurity (Gini or entropy) the most. Repeat in each child node until a stopping rule applies, such as a minimum node size or a maximum depth. Each leaf predicts its majority class.
How does the k-nearest neighbours algorithm work?
For a new record, it finds the k training records closest by a distance measure. For classification it takes the majority class. For regression it averages their target values. Scale the features first, and choose k using validation data.
Why is naive Bayes called naive?
It assumes the features are independent of each other given the class. This is rarely exactly true. It makes the calculation simple, because you just multiply the individual conditional probabilities by the prior.