Skip to content

FRM Part I · FRM Exam Part I

Machine-Learning Methods for FRM Part I

Machine-learning methods use data to build models that predict outcomes or find structure, with little hand-coded logic. For FRM Part I, learn the types of learning, overfitting and validation, regularization, classification metrics, trees, clustering, PCA and neural networks. Most questions test concepts and short calculations, so know definitions and formulas precisely.

What this chapter covers

This chapter covers how machines learn patterns from data. You start with the three broad families: supervised learning (labelled data, predicting a target), unsupervised learning (no labels, finding structure) and reinforcement learning (learning by reward). From there the chapter builds a toolkit: how to judge a model, how to prepare data, how to control complexity, and which algorithm suits which problem.

The ideas repeat across the chapter. Every model must balance fit against generalization. That is why overfitting, the bias-variance tradeoff, cross-validation and regularization sit at the centre. Once you understand that tension, ridge, LASSO, pruned trees, ensembles and neural-network tuning all read as variations of one theme.

The chapter connects closely to Quantitative Analysis. Regression, hypothesis testing, and the variance and covariance ideas you already know are the base for regularization, logistic regression and PCA. It also links to Valuation and Risk Models, since credit scoring, fraud detection and market-risk modelling use these methods, and to model risk in Foundations of Risk Management, where poor validation is a real source of loss.

The exam has 100 equally weighted multiple-choice questions in 4 hours, so every topic area counts and a newer chapter like this one can be a cheap source of marks. Questions are usually short: pick the right method, interpret a confusion matrix, or explain what a penalty does to coefficients. If you know the definitions and a handful of formulas, you can answer fast and bank time for longer calculation questions elsewhere. Candidates who skip it as too technical give away marks that need no heavy maths.

Machine-Learning Methods: topics in the order to study them

  1. 1Overview of Machine Learning and Types of LearningIt gives you the vocabulary (supervised, unsupervised, reinforcement, regression versus classification) that every later topic uses.
  2. 2Overfitting, Bias-Variance Tradeoff and Model ValidationThis is the core idea of the chapter. Regularization, trees and neural networks all make sense once you see this tradeoff.
  3. 3Data Preparation and Feature EngineeringScaling, missing values and feature choice affect every model, and you need them before studying penalties and distance-based methods.
  4. 4Regularization: Ridge, LASSO and Elastic NetIt builds on regression and applies the bias-variance idea directly, so it follows validation and data preparation.
  5. 5Logistic Regression and Classification MetricsIt is the first classification model and introduces the confusion matrix, precision, recall and related measures that later topics reuse.
  6. 6Decision Trees, Ensembles and K-Nearest NeighborsThese are flexible classifiers that you can compare against logistic regression, and they show overfitting and variance reduction in practice.
  7. 7Unsupervised Learning: Clustering and PCAIt switches to unlabelled data. PCA uses the variance and covariance ideas you have already revised.
  8. 8Neural Networks and Deep LearningIt comes last because it combines ideas from regression, regularization and validation in the most complex model.

How to prepare Machine-Learning Methods

Aim for clear concepts first, then speed on the small calculations. Most of this chapter can be revised on a phone in short sessions.

  1. Read the GARP Study Guide and Learning Objectives for the current year, since the curriculum is revised annually, and tick off each objective as you cover it.
  2. Study in the order given. After each topic, write a three-line summary: what the method does, when to use it, and its main weakness.
  3. Memorize the formulas you need: accuracy, precision, recall, specificity, F1, the logistic function, and the ridge and LASSO penalty terms. Practise computing them from a small confusion matrix by hand.
  4. Build a comparison table in your notes: for each method list supervised or unsupervised, data needs, interpretability, overfitting risk and a typical risk-management use.
  5. Do practice questions in timed blocks. For every wrong answer, note whether the cause was a concept gap, a formula slip or a misread question.
  6. Two weeks before the exam, rework only your error log and the quick revision list, and link this chapter back to regression and PCA in Quantitative Analysis.

Common mistakes in Machine-Learning Methods

  • Judging a model by its training accuracy

    Fix: Always ask how the model does on data it has not seen. Choose answers based on validation or test performance.

  • Mixing up precision and recall

    Fix: Precision asks: of the cases flagged, how many were right (denominator TP + FP). Recall asks: of the real positives, how many were found (denominator TP + FN).

  • Saying ridge performs variable selection

    Fix: Remember that LASSO's penalty can force coefficients to exactly zero, while ridge only shrinks them toward zero.

  • Forgetting to scale features

    Fix: Link scaling to any method using penalties or distances: ridge, LASSO, KNN, clustering and PCA.

  • Treating accuracy as enough on imbalanced data

    Fix: Check the confusion matrix and use recall, precision or F1 when the positive class is rare.

  • Confusing supervised and unsupervised tasks

    Fix: Ask whether the groups are known in advance with labels. If yes, it is classification; if the groups are discovered from the data, it is clustering.

Last-day revision: Machine-Learning Methods

  • Supervised learning uses labelled data; unsupervised learning finds structure without labels; reinforcement learning learns from rewards.
  • Overfitting means low training error but high error on new data; underfitting means the model is too simple to capture the pattern.
  • Higher model complexity usually lowers bias and raises variance.
  • Use separate training, validation and test data; cross-validation rotates the validation fold.
  • Scale features before using penalized regression, KNN or PCA, because these depend on the size of variables.
  • Ridge shrinks coefficients but does not usually set them to exactly zero; LASSO can set some to exactly zero, so it selects features.
  • Elastic Net mixes the ridge and LASSO penalties.
  • Logistic regression outputs a probability between 0 and 1 using p = 1 ÷ (1 + e^(−z)).
  • Precision = TP ÷ (TP + FP); recall = TP ÷ (TP + FN); accuracy = (TP + TN) ÷ total.
  • Single decision trees overfit easily; ensembles such as bagging and random forests reduce variance by combining many trees.
  • K-means needs the number of clusters set in advance; PCA builds uncorrelated components ordered by variance explained.
  • Neural networks learn through layers, weights and activation functions, and need regularization or early stopping to limit overfitting.

Machine-Learning Methods practice questions

Machine-Learning Methods in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Machine-Learning Methods: frequently asked questions

Do I need to know programming for the machine-learning chapter?

No. The exam is multiple-choice, so you need to understand concepts, interpret results and do short calculations. You do not write code.

How hard are the calculations in this chapter?

They are usually light. Expect things like computing precision and recall from a confusion matrix or evaluating the logistic function. A basic calculator is enough for most of them.

Is this chapter linked to other FRM Part I topics?

Yes. It builds on regression, variance and covariance from Quantitative Analysis, and it connects to credit and market risk modelling and to model risk in Foundations of Risk Management.

Which topics should I prioritize if time is short?

Start with overfitting, the bias-variance tradeoff and validation, then regularization and classification metrics. These ideas feed into almost every other topic in the chapter.