Skip to content

FRM Exam Part I · Machine Learning and Prediction

Neural Networks and Deep Learning for FRM Part I

Updated 11 October 2026 · Fact-checked

A neural network is a model built from layers of nodes. Each node computes a weighted sum of its inputs, adds a bias, and passes the result through an activation function. Backpropagation uses the chain rule to find how each weight changes the loss, and gradient descent updates the weights to reduce it.

Understand Neural Networks and Deep Learning

A neural network is a flexible function that maps inputs (features) to an output (a prediction). You can think of it as many small regressions stacked together. Each small unit is called a node or neuron.

The nodes sit in layers. The input layer holds your features, such as leverage, volatility or a credit score. One or more hidden layers transform those features. The output layer gives the prediction: a number for regression, or a probability for classification. A network with many hidden layers is called a deep network, and training it is deep learning.

Inside each node, two things happen. First, the node forms a weighted sum of its inputs plus a bias: z = w1x1 + w2x2 + ... + b. Second, it applies an activation function to z. Without the activation function, stacking layers would still give only a linear model. The activation adds nonlinearity, which lets the network capture interactions and curved relationships that a linear regression misses.

Training means choosing weights and biases that minimise a loss function, such as mean squared error for regression or cross-entropy for classification. Backpropagation is the method that computes the gradient of the loss with respect to every weight. It runs a forward pass to get the prediction and the loss, then a backward pass using the chain rule. Gradient descent then moves each weight a small step against its gradient. The step size is the learning rate.

In finance, networks are used for credit scoring, default prediction, fraud detection, volatility forecasting and pricing complex derivatives. Their strengths are flexibility and predictive power. Their weaknesses are overfitting, large data needs, and poor interpretability (the black-box problem), which matters for model risk and regulators. Controls include data splitting, early stopping, regularisation and dropout.

Key formulas to remember

Node pre-activation
z = Σ (wᵢ × xᵢ) + b
Weighted sum of inputs plus a bias. The activation function is applied to z.
Node output
a = f(z)
f is the activation function. If f is linear, stacked layers collapse to one linear model.
Sigmoid
σ(z) = 1 ÷ (1 + e^(−z))
Output lies between 0 and 1. Used for probabilities, for example default probability. Gradient is small for large positive or negative z.
Sigmoid derivative
σ'(z) = σ(z) × (1 − σ(z))
Maximum value is 0.25 at z = 0. Explains the vanishing gradient problem.
Tanh
tanh(z) = (e^z − e^(−z)) ÷ (e^z + e^(−z))
Output between −1 and 1 and centred on zero.
ReLU
ReLU(z) = max(0, z)
Output is zero for negative z and equal to z otherwise. Cheap to compute and eases vanishing gradients.
Mean squared error loss
L = (1 ÷ n) × Σ (yᵢ − ŷᵢ)²
Common loss for regression. For a single observation with a ½ factor, L = ½(y − ŷ)².
Chain rule for a weight
∂L/∂w = ∂L/∂a × ∂a/∂z × ∂z/∂w
This is backpropagation. For a node, ∂z/∂w equals the input x.
Gradient descent update
w(new) = w(old) − η × ∂L/∂w
η is the learning rate. Too large can overshoot; too small makes training slow.

How to solve Neural Networks and Deep Learning questions

Most exam questions on this topic are conceptual, or ask for a small calculation of one node, one activation or one weight update. Use the same routine each time.

  1. 1Identify what is asked: structure, activation choice, a forward-pass number, a gradient or update, or a risk and use-case judgement.
  2. 2Write down the inputs, weights, bias and activation function given in the question.
  3. 3For a forward pass, compute z = Σ wx + b first, then apply the activation f(z). Keep extra decimals until the end.
  4. 4For a loss, compare the prediction with the target using the stated loss function.
  5. 5For backpropagation, multiply the chain-rule pieces in order: loss gradient, activation derivative, then input. Check each sign.
  6. 6For an update, apply w(new) = w(old) − η × gradient. Subtract the gradient, do not add it.
  7. 7For conceptual choices, match the feature to the task: sigmoid for probabilities, ReLU for hidden layers in deep networks, and regularisation or early stopping for overfitting.
  8. 8Sanity-check the answer: a sigmoid output must be between 0 and 1, and ReLU output cannot be negative.

Quickest way: Match the keyword to the concept

When to use it: Use this for conceptual multiple-choice questions when time is short. Calculation questions still need the full steps.

  1. Nonlinearity or capturing complex patterns points to the activation function.
  2. Output between 0 and 1 or a probability points to sigmoid. Output of exactly zero for negative inputs points to ReLU.
  3. Gradients, chain rule or error flowing backward points to backpropagation.
  4. Step size points to the learning rate. Too high means overshooting and too low means slow training.
  5. Great in-sample fit with poor out-of-sample results points to overfitting. Think validation set, early stopping, dropout or regularisation.
  6. Hard to explain to a regulator points to the black-box and model-risk problem.
  7. For a calculation, do z first, then f(z), then the loss, and only then the gradient.

Common mistakes in Neural Networks and Deep Learning

  • Believing more hidden layers help even if the activation is linear.

    Students assume depth alone creates complexity.

    Fix: Remember that a composition of linear functions is linear. Nonlinear activation functions are what let depth add flexibility.

  • Adding the gradient in the weight update.

    Students forget the aim is to reduce the loss.

    Fix: Use w(new) = w(old) − η × gradient. A positive gradient means the weight should fall.

  • Forgetting the bias in z.

    Attention goes to the weights and inputs only.

    Fix: Always write z = Σ wx + b before computing anything.

  • Treating sigmoid output as a ReLU-style unbounded value, or using ReLU for a probability output.

    Activation shapes are mixed up.

    Fix: Sigmoid is bounded in (0, 1). ReLU is zero or positive with no upper limit. Probabilities need a bounded output.

  • Saying backpropagation is the same as gradient descent.

    The two are taught together.

    Fix: Backpropagation computes the gradients. Gradient descent uses them to update the weights.

  • Judging a network by training error alone.

    Low training loss looks like success.

    Fix: Judge on held-out validation or test data. A large gap between training and test error signals overfitting.

Worked examples

Example 1

A single node has inputs x1 = 2 and x2 = −1, weights w1 = 0.5 and w2 = 0.8, and bias b = 0.3. Find the node output under (a) ReLU and (b) sigmoid. Give the sigmoid to three decimals (e^(−0.5) = 0.60653).

Show the solution
  1. Compute z = w1x1 + w2x2 + b = 0.5 × 2 + 0.8 × (−1) + 0.3.
  2. z = 1.0 − 0.8 + 0.3 = 0.5.
  3. (a) ReLU(0.5) = max(0, 0.5) = 0.5.
  4. (b) Sigmoid = 1 ÷ (1 + e^(−0.5)) = 1 ÷ (1 + 0.60653) = 1 ÷ 1.60653.
  5. 1 ÷ 1.60653 = 0.62246.

Answer: ReLU output is 0.5. Sigmoid output is about 0.622.

Example 2

An output node has pre-activation z = 0 and sigmoid activation, so a = 0.5. The target is y = 1. The loss is L = ½(y − a)². The input to this node from the previous layer is x = 2. Using backpropagation, find ∂L/∂w, then the updated weight if w = 0.4 and the learning rate is η = 0.8.

Show the solution
  1. Loss gradient with respect to the output: ∂L/∂a = −(y − a) = −(1 − 0.5) = −0.5.
  2. Activation derivative: ∂a/∂z = a(1 − a) = 0.5 × 0.5 = 0.25.
  3. Input term: ∂z/∂w = x = 2.
  4. Chain rule: ∂L/∂w = (−0.5) × 0.25 × 2 = −0.25.
  5. Update: w(new) = 0.4 − 0.8 × (−0.25) = 0.4 + 0.2 = 0.6.

Answer: ∂L/∂w = −0.25 and the updated weight is 0.6. The weight rises because the prediction was below the target.

Exam tips

  • Expect conceptual questions on why nonlinear activations matter, what backpropagation does, and the link between overfitting and network size.
  • For calculations, the numbers are usually small. Do the forward pass cleanly and watch the signs in the update.
  • Know the one-line properties of sigmoid, tanh and ReLU: range, typical use, and the vanishing gradient issue for sigmoid.
  • Link neural networks to model risk: interpretability, data requirements and validation on out-of-sample data are favourite angles in finance scenarios.
  • If an option says the network needs no validation because it is flexible, reject it. Flexibility raises overfitting risk.

Practice questions from Machine Learning and Prediction

Neural Networks and Deep Learning in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Neural Networks and Deep Learning: frequently asked questions

How does backpropagation work in simple terms?

The network makes a prediction and measures the error. It then works backward from the output, using the chain rule to find how much each weight contributed to that error. Gradient descent uses those contributions to adjust the weights.

Why do neural networks need activation functions?

Without them, every layer is a linear transformation and the whole network reduces to one linear model. Activation functions add nonlinearity, so the network can learn curved relationships and interactions.

What is the difference between a neural network and deep learning?

A neural network is the model. Deep learning refers to networks with many hidden layers, which can learn more layered representations of the data but need more data and computing power.

How is deep learning used in risk management?

Common uses include credit scoring and default prediction, fraud detection, volatility forecasting and approximating prices of complex derivatives. The main concerns are overfitting, limited interpretability and model risk.