Skip to content

FRM Exam Part I · Machine-Learning Methods

Neural Networks and Deep Learning for FRM Part I

Updated 11 October 2026 · Fact-checked

A neural network is a model built from layers of nodes. Each node takes a weighted sum of its inputs, adds a bias, and passes the result through an activation function. Training uses backpropagation and gradient descent to adjust weights and reduce a loss function. Deep networks have many hidden layers and need controls against overfitting.

Understand Neural Networks and Deep Learning

A neural network is a flexible model that learns a mapping from inputs (features) to outputs (a prediction). Think of it as a chain of small regressions. Each small unit is a neuron or node.

The network has an input layer, one or more hidden layers, and an output layer. A node computes a weighted sum of its inputs plus a bias, then applies an activation function. Without the activation function, stacking layers would just give one big linear model. The activation adds non-linearity, which lets the network capture complex patterns. A deep network is one with many hidden layers.

Training means choosing weights and biases that minimise a loss function, such as mean squared error for a numeric target or cross-entropy for a classification target. The network makes a forward pass to get a prediction and compute the loss. Backpropagation then uses the chain rule to work out how much each weight contributed to the loss. Gradient descent moves each weight a small step against its gradient. The step size is the learning rate. Too large a rate can overshoot. Too small a rate makes training slow.

Deep models have many parameters, so they can overfit: they fit noise in the training data and fail on new data. Common controls are splitting data into training, validation and test sets, early stopping, regularisation (penalties on weights), and dropout (randomly switching off nodes during training). More data also helps.

In finance risk work, networks are used for credit scoring, fraud detection, and forecasting. Their weakness is interpretability. A network is often a black box, so it is hard to explain why it rejected a borrower. This matters for model risk, validation and regulation. Simpler models such as logistic regression are easier to explain, even if they are sometimes less accurate.

Key formulas to remember

Node output
a = f(z), where z = w₁x₁ + w₂x₂ + … + wₙxₙ + b
w are weights, b is the bias, f is the activation function.
Sigmoid activation
σ(z) = 1 ÷ (1 + e^(−z))
Output lies between 0 and 1. Often used for probabilities. Can cause vanishing gradients for large |z|.
ReLU activation
ReLU(z) = max(0, z)
Simple and widely used in hidden layers. Output is zero for negative z.
Tanh activation
tanh(z) = (e^z − e^(−z)) ÷ (e^z + e^(−z))
Output lies between −1 and 1 and is centred on zero.
Mean squared error loss
MSE = (1 ÷ n) × Σ (yᵢ − ŷᵢ)²
Typical loss for numeric targets.
Gradient descent update
w_new = w_old − η × ∂L/∂w
η is the learning rate. L is the loss.
Chain rule in backpropagation
∂L/∂w = ∂L/∂a × ∂a/∂z × ∂z/∂w
Gradients are passed backward from output layer to earlier layers.

How to solve Neural Networks and Deep Learning questions

Use this routine for conceptual and numeric questions on neural networks.

  1. 1Identify what is asked: structure, activation, training step, overfitting, or interpretability.
  2. 2For a numeric node question, compute z = Σ(w × x) + b first.
  3. 3Apply the stated activation function to z. Check the output range of that function.
  4. 4For a training question, name the loss, then use the update w_new = w_old − η × gradient. Watch the minus sign.
  5. 5For a performance gap question, compare training and validation error. Low training error with high validation error signals overfitting.
  6. 6Match the fix to the problem: dropout, early stopping, regularisation, or more data for overfitting.
  7. 7For governance questions, link black-box behaviour to explainability, validation and model risk.
  8. 8Check that your answer is in the right units and range.

Quickest way: Quick exam routine

When to use it: Use when you have about 90 seconds per question and the options look similar.

  1. Scan the stem for keywords: activation, backpropagation, learning rate, overfitting, interpretability.
  2. Numeric: compute z, apply f, stop. Use your calculator's e^x key for sigmoid.
  3. Rule out options that say activation functions are linear, or that backpropagation changes the data.
  4. Training error low and validation error high means overfitting. Both high means underfitting.
  5. Pick the option that matches the standard definition.

Common mistakes in Neural Networks and Deep Learning

  • Forgetting the bias term when computing z.

    Students focus on weights and inputs only.

    Fix: Always write z = Σ(w × x) + b and add b before applying the activation.

  • Saying that deep networks work without non-linear activation functions.

    Students think more layers alone add power.

    Fix: Remember that stacked linear layers collapse into a single linear model. Non-linearity comes from the activation.

  • Confusing backpropagation with gradient descent.

    Both appear in the same training loop.

    Fix: Backpropagation computes the gradients. Gradient descent uses them to update the weights.

  • Adding the gradient instead of subtracting it in the update.

    Students forget that the goal is to reduce the loss.

    Fix: Move against the gradient: w_new = w_old − η × gradient.

  • Treating low training error as proof of a good model.

    Students ignore out-of-sample results.

    Fix: Judge the model on validation or test data. A large gap between training and test error points to overfitting.

  • Assuming the sigmoid output can be negative or above 1.

    Mixing up sigmoid, tanh and ReLU ranges.

    Fix: Sigmoid is between 0 and 1, tanh between −1 and 1, ReLU is 0 or higher.

Worked examples

Example 1

A node has inputs x₁ = 2 and x₂ = −1, weights w₁ = 0.5 and w₂ = 1.5, and bias b = 0.5. The activation is ReLU. What is the node output? What would the output be if the bias were −1.5?

Show the solution
  1. Compute z = (0.5 × 2) + (1.5 × −1) + 0.5.
  2. 0.5 × 2 = 1. 1.5 × −1 = −1.5. So z = 1 − 1.5 + 0.5 = 0.
  3. ReLU(0) = max(0, 0) = 0.
  4. With b = −1.5: z = 1 − 1.5 − 1.5 = −2.
  5. ReLU(−2) = max(0, −2) = 0.

Answer: The output is 0 in the first case (z = 0). It is also 0 in the second case (z = −2), because ReLU sets negative values to zero.

Example 2

A weight is currently w = 0.80. For a given training example, backpropagation gives a gradient ∂L/∂w = 0.50. The learning rate is 0.10. What is the updated weight after one gradient descent step? Then, if a later gradient is −0.30 at the new weight, what is the next weight?

Show the solution
  1. Use w_new = w_old − η × gradient.
  2. First step: 0.80 − (0.10 × 0.50) = 0.80 − 0.05 = 0.75.
  3. Second step: 0.75 − (0.10 × −0.30) = 0.75 + 0.03 = 0.78.

Answer: The weight becomes 0.75 after the first step and 0.78 after the second step.

Exam tips

  • Know the output ranges of sigmoid (0 to 1), tanh (−1 to 1) and ReLU (0 or above). Questions often test these directly.
  • Be able to state in one line what backpropagation does: it computes gradients of the loss with respect to each weight using the chain rule.
  • For overfitting questions, expect answers like dropout, early stopping, regularisation, more data or a simpler network.
  • For risk-management context questions, expect interpretability, model validation and model risk as the core concerns with deep models.
  • Numeric questions are short. Write z, the activation and the answer, and use your calculator for exponentials.

Practice questions from Machine-Learning Methods

Neural Networks and Deep Learning in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Neural Networks and Deep Learning: frequently asked questions

What is an activation function in a neural network?

It is a function applied to a node's weighted sum plus bias. It adds non-linearity so the network can learn complex patterns. Common choices are sigmoid, tanh and ReLU.

How does backpropagation work?

After a forward pass gives a prediction and a loss, backpropagation applies the chain rule layer by layer from the output back to the input. This gives the gradient of the loss for every weight. Gradient descent then uses those gradients to update the weights.

Why are neural networks hard to use in risk management?

They can be hard to interpret, so it is difficult to explain individual decisions to regulators, clients or model validators. They also overfit easily if data is limited. These issues raise model risk.

What is the difference between a neural network and deep learning?

Deep learning refers to neural networks with many hidden layers. A network with a single small hidden layer is a neural network but is not usually called deep.