IAI Actuarial Core Principles · Risk Modelling and Survival Analysis
Elementary Principles of Machine Learning for CS2
Machine learning means building models that learn patterns from data instead of being told fixed rules. For CS2 you must know supervised and unsupervised learning, how to train and validate a model, how overfitting arises, and how to measure performance. Solve questions by naming the task, then the method, then the check.
What this chapter covers
This chapter is the machine learning part of CS2. It introduces the main types of learning, the common methods in each, and the discipline needed to build a model that works on new data. You are not asked to become a programmer. You are asked to understand ideas, explain them clearly, and apply them to short actuarial scenarios.
The chapter has a clear flow. First you learn what machine learning is and how it splits into supervised learning (with a known target) and unsupervised learning (no target). Then you look at regression and classification, and at clustering and dimension reduction. The last two topics cover how you train, validate and judge a model. These last topics apply to every method before them.
It links to the rest of CS2 in two ways. Regression and classification build on the statistical modelling you already know from CS1, such as linear models. The ideas of fitting, parameter estimation and model choice also appear in time series and survival models. In the 2026 syllabus this chapter is a smaller block than stochastic processes or survival models, so it rewards focused effort. Because Paper B is computer-based, expect to see these ideas in practical form as well.
Machine learning carries a smaller syllabus weighting than the larger CS2 blocks, but it is conceptual and easy to learn well, so it is a good place to secure marks. It is also a likely source of multiple-choice questions, where one clear definition earns the full 2 marks. Written questions reward short, precise explanations of overfitting, validation and performance measures. Strong command of this chapter also helps your wider modelling judgement, which the IAI values across Core Principles.
Elementary principles of machine learning: topics in the order to study them
- 1Introduction to Machine Learning and Its TypesYou need the vocabulary and the supervised versus unsupervised split before any method makes sense.
- 2Supervised Learning: Regression and ClassificationIt builds on CS1 regression, so it feels familiar and gives you a base for judging models later.
- 3Unsupervised Learning: Clustering and Dimension ReductionStudy it after supervised learning so you can contrast the two, since no target variable is used.
- 4Model Training, Validation and OverfittingThis explains how to fit any of the methods above without fooling yourself, so it comes once you know the methods.
- 5Model Performance Measures and Practical ConsiderationsIt closes the chapter by showing how to judge models and handle real data issues, tying all earlier topics together.
How to prepare Elementary principles of machine learning
Treat this chapter as concepts first, then application. Short, regular sessions work well if you study alongside work or on your phone.
- Write a one-page map of the chapter: types of learning, example methods under each, and the task each method solves.
- For each method, learn three things: what it does, what data it needs, and one weakness. Say them aloud without notes.
- Link regression and classification to CS1 linear models so you reuse what you know instead of memorising from scratch.
- Learn the training, validation and test split in your own words, then explain overfitting and how to detect and reduce it.
- Practise performance measures on small examples. Build a simple confusion matrix by hand and compute accuracy, and any other measure in your notes, from it.
- Work past-style MCQs and short written answers. State the task, name the method, give the reason, and note an assumption or limitation.
- Repeat the practical ideas in R for Paper B, so you can run a model, split data and read the output.
Common mistakes in Elementary principles of machine learning
Mixing up supervised and unsupervised learning.
Fix: Ask first: is there a known outcome in the data? If yes, it is supervised. If no, it is unsupervised.
Judging a model by its performance on the training data only.
Fix: Always assess on validation or test data and compare with training results to spot overfitting.
Using the test set repeatedly to tune the model.
Fix: Tune on validation data and keep the test set for one final check, so the result stays unbiased.
Relying on accuracy when classes are unbalanced.
Fix: Read the confusion matrix and use measures that reflect the cost of each type of error.
Writing long, vague answers with no definition.
Fix: Start with a one-line definition, then give the reason or example, then one limitation.
Ignoring the practical side for Paper B.
Fix: Practise splitting data, fitting a model and reading the output in R, and be ready to explain the result in words.
Last-day revision: Elementary principles of machine learning
- Supervised learning uses labelled data with a target; unsupervised learning has no target.
- Regression predicts a numerical value; classification predicts a category.
- Clustering groups similar observations; dimension reduction cuts the number of variables while keeping the main information.
- Data is split into training data to fit the model and held-back data to test it on unseen cases.
- A validation set is used to choose between models or tune settings; a test set gives a final unbiased check.
- Overfitting means the model fits noise in the training data and performs poorly on new data.
- Signs of overfitting: very good training results but clearly worse validation results.
- Simpler models, more data and regularisation are common ways to reduce overfitting.
- In classification, a confusion matrix counts correct and incorrect predictions by class.
- Accuracy alone can mislead when one class is rare.
- Always state the assumptions and the purpose of the model before choosing a measure.
- Data quality, such as missing values and scaling, affects every method.
Elementary principles of machine learning practice questions
- A pricing team fits a claim-severity model and finds that its error on the data used for fitting is very small, but its error on a separate …
- A life insurer in Mumbai has data on 20,000 customers (age, premium paid, number of policies, channel) but no outcome variable. An analyst a…
- An actuary fits a flexible model to 200 motor claims and finds it has a very small error on those 200 records but a much larger error on 100…
- A modeller fits polynomial regressions of increasing degree to the same training data. The training error falls steadily with degree, but th…
- An analyst uses k-fold cross-validation with k = 5 on a dataset of 1,000 observations to compare models. Which description of the procedure …
- In ridge regression, a penalty equal to a positive constant λ times the sum of squared coefficients is added to the residual sum of squares.…
- A dataset is split into training, validation and test sets to choose the number of neighbours k in a k-nearest-neighbours model. What is the…
- An insurer has a dataset of past motor policies, each labelled with whether a claim was made, and wants a model to predict claims for new po…
Elementary principles of machine learning in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Elementary principles of machine learning: frequently asked questions
Is machine learning a big part of CS2?
In the 2026 syllabus, machine learning has a weighting of 10%, smaller than stochastic processes or survival models. It is still worth full effort because the content is conceptual and quick to learn. You can see it in both multiple-choice and written questions.
Do I need to code for this chapter?
You need to understand the methods and be able to explain results. Paper B is a computer-based exam, so practise basic R work such as splitting data and fitting a model. Check the current syllabus for the exact expectations.
What is the difference between overfitting and underfitting?
Overfitting means the model follows noise in the training data and does badly on new data. Underfitting means the model is too simple to capture the real pattern, so it does badly on both training and new data.
How should I answer a written question on a machine learning method?
Name the task, such as classification or clustering, then state the method and why it suits the data. Add one assumption or limitation and say how you would check performance on unseen data.