Skip to content

FRM Exam Part I · Machine Learning and Prediction

Machine Learning Basics and Types of Learning for FRM Part 1

Updated 11 October 2026 · Fact-checked

Machine learning lets algorithms find patterns in data and predict, with little hand-written structure. Supervised learning uses labeled outcomes, unsupervised learning finds structure in unlabeled data, and reinforcement learning learns actions by trial and reward. To answer exam questions, identify whether a target variable exists and what the goal is.

Understand Machine Learning Basics and Types of Learning

Machine learning (ML) is a set of methods where a model learns patterns from data instead of following rules written by a person. The focus is on prediction: how well does the model work on data it has not seen?

Traditional econometrics usually starts with a theory and a specified model, such as a linear regression. The focus is on inference: estimating coefficients, testing hypotheses and interpreting effects. ML usually accepts flexible, less interpretable models, many inputs and nonlinear relationships, and judges success by out-of-sample performance. The two overlap. Linear regression is used in both.

There are three main types of learning. In supervised learning, each observation has inputs (features) and a known outcome (the label or target). The model learns the mapping. If the target is a number (a loss amount, a return), it is regression. If the target is a category (default or no default, fraud or not), it is classification. Finance uses include credit scoring, default prediction, fraud detection and forecasting.

In unsupervised learning, there is no target. The model looks for structure in the features alone. Typical tasks are clustering (grouping similar customers, bonds or funds) and dimension reduction such as principal components analysis (PCA), which compresses many correlated variables into a few. Finance uses include grouping borrowers, summarizing yield curve moves and spotting unusual transactions.

In reinforcement learning, an agent takes actions in an environment, receives rewards, and learns a policy that maximizes cumulative reward over time. There is no labeled answer for each step. Finance uses include optimal trade execution, dynamic hedging and portfolio rebalancing. Also note that ML models need careful validation. A flexible model can fit noise, which is why training and test data are kept separate.

Key formulas to remember

Supervised learning
Inputs (features) + known target → learn mapping → predict target
Numeric target = regression. Categorical target = classification.
Unsupervised learning
Inputs only, no target → find structure (clusters, components)
Main tools: clustering and PCA. No right answer is supplied.
Reinforcement learning
Agent → action → environment → reward → update policy
Goal is to maximize cumulative reward, not to match labels.
Data split
Training set (fit) | Validation set (tune) | Test set (final check)
Judge a model on data it did not train on.

How to solve Machine Learning Basics and Types of Learning questions

Use this approach for any question that asks you to classify a problem, a method or a use case.

  1. 1Read what the model is supposed to produce: a prediction, a grouping, or a sequence of decisions.
  2. 2Check whether labeled outcomes exist for the training data.
  3. 3If labels exist, decide whether the target is numeric (regression) or categorical (classification).
  4. 4If no labels exist and the aim is groups or simplification, choose unsupervised learning (clustering or PCA).
  5. 5If the model chooses actions over time and learns from rewards, choose reinforcement learning.
  6. 6If the question contrasts ML with econometrics, decide whether the emphasis is inference and theory or out-of-sample prediction.
  7. 7Eliminate options that misstate a definition, such as saying unsupervised learning needs labels.

Quickest way: Label test

When to use it: Use it when a question asks which type of learning fits a scenario.

  1. Ask: is there a known answer for each training record? Yes means supervised.
  2. Ask: is the goal to find groups or reduce variables with no known answer? Yes means unsupervised.
  3. Ask: does it act, get rewarded and adjust over time? Yes means reinforcement.
  4. Then check the target type to pick regression or classification.

Common mistakes in Machine Learning Basics and Types of Learning

  • Calling clustering a supervised method

    Clusters look like categories, so candidates think they are labels.

    Fix: Clusters are discovered by the algorithm. Classification uses labels that were given in advance.

  • Treating a default prediction model as regression because it outputs a probability

    The output is a number between 0 and 1.

    Fix: Look at the target. Default or not is a category, so it is classification.

  • Saying ML and econometrics are completely different

    Textbooks stress the contrast.

    Fix: They overlap. Linear and logistic regression are used in both. The difference is mainly emphasis: inference versus prediction.

  • Thinking reinforcement learning needs labeled data for every action

    Candidates mix it up with supervised learning.

    Fix: It learns from rewards that may be delayed, not from a correct label at each step.

  • Judging a model by in-sample fit

    A high fit on training data looks like success.

    Fix: Flexible models can memorize noise. Assess performance on held-out data.

Worked examples

Example 1

A bank has 50,000 past loans, each marked as defaulted or repaid, with borrower income, debt and credit history. It builds a model to predict whether new applicants will default. Which type of learning is this, and what kind of task?

Show the solution
  1. Labels exist: every past loan has a known outcome.
  2. So the learning is supervised.
  3. The target (default or repaid) is a category.
  4. So the task is classification, not regression.

Answer: Supervised learning, specifically classification.

Example 2

A risk team has transaction data for 20,000 corporate clients with no default or rating labels. It wants to group clients with similar behavior to review exposures. Which approach fits, and why is regression inappropriate?

Show the solution
  1. No target variable exists, so supervised methods do not apply.
  2. The goal is to find natural groups from the features alone.
  3. This is unsupervised learning, using clustering.
  4. Regression needs a numeric target to predict, and none is available.

Answer: Unsupervised learning (clustering). Regression is inappropriate because there is no target variable.

Exam tips

  • Most questions are scenario matching. Find the label and the goal first.
  • Expect contrasts: inference versus prediction, labeled versus unlabeled, supervised versus reinforcement.
  • Watch for options that wrongly say unsupervised learning uses labeled data.
  • Remember the finance examples: credit scoring and fraud (supervised), clustering and PCA (unsupervised), trade execution and dynamic hedging (reinforcement).

Practice questions from Machine Learning and Prediction

Machine Learning Basics and Types of Learning in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Machine Learning Basics and Types of Learning: frequently asked questions

What is the difference between supervised and unsupervised learning?

Supervised learning trains on data with known outcomes and predicts them for new data. Unsupervised learning has no outcomes and looks for structure such as groups or main components in the features.

How is machine learning different from traditional econometrics?

Econometrics usually starts from theory and a specified model, and focuses on estimating and testing coefficients. ML focuses on predictive accuracy on new data and often uses flexible, harder to interpret models. The two share some tools, such as regression.

What are examples of reinforcement learning in finance?

Common examples are optimal trade execution, dynamic hedging and portfolio rebalancing. In each, an agent takes actions and learns from rewards such as lower costs or better risk-adjusted returns.

Is classification supervised or unsupervised?

Classification is supervised, because it learns from examples with known categories. Clustering is the unsupervised counterpart, where groups are discovered without labels.