CFA Level II Exam · Machine Learning
Neural Networks, Deep Learning and Reinforcement Learning
Updated 7 October 2026 · Fact-checked
A neural network is a set of connected nodes in an input layer, hidden layers and an output layer. Each node sums weighted inputs and applies an activation function. Deep learning uses many hidden layers. Reinforcement learning has an agent learn by rewards. To choose an algorithm, match it to the data type and the task.
Understand Neural Networks, Deep Learning and Reinforcement Learning
A neural network (NN) is a supervised learning technique that can model nonlinear relationships between inputs and an output. It has three kinds of layers: an input layer (one node per feature), one or more hidden layers, and an output layer that gives the prediction.
Each hidden node does two jobs. First it forms a weighted sum of its inputs plus a bias (the summation operator). Then it passes that sum through an activation function, which transforms it into the node's output. The activation function is what makes the network nonlinear. Common ones are sigmoid (output between 0 and 1), tanh (between -1 and 1) and ReLU (zero for negative inputs, the input itself for positive ones). The weights are learned by forward propagation (computing output) and backward propagation (adjusting weights to cut the error). The learning rate controls the size of each adjustment.
Deep neural networks (DNNs) have many hidden layers, typically more than two. Extra layers let the network learn complex patterns, such as in image, speech and text tasks. Known variants are the convolutional neural network (suited to images and spatial data) and the recurrent neural network (suited to sequential data such as time series and text). NNs can overfit, need a lot of data, and are hard to interpret, so they are often called black boxes.
Reinforcement learning (RL) has no labelled data. An agent takes actions in an environment, observes the resulting state, and receives a reward. It learns a policy that maximises cumulative reward over time. It balances exploration (trying new actions) with exploitation (using what it already knows). Examples include trading-execution strategies and game-playing. Its weaknesses are that it needs many trials and that results depend on how the environment and reward are specified.
Choosing an algorithm starts with the data and the goal. If the target is labelled, use supervised learning: continuous target means regression; categorical target means classification. If there is no target, use unsupervised learning: reduce dimensions with PCA, or group observations with clustering. Within supervised learning, a small or mid-sized structured dataset with a need for interpretability favours penalized regression or CART. Large datasets with complex nonlinear patterns, images, speech or text favour neural networks and deep learning. Sequential decision-making with rewards favours RL.
Key formulas to remember
- Node calculation
- Node output = f(Σ wᵢxᵢ + b)
- wᵢ are weights, xᵢ inputs, b the bias, f the activation function. The curriculum describes this as a summation operator followed by an activation function.
- Sigmoid activation
- f(x) = 1 ÷ (1 + e^(-x))
- Output lies between 0 and 1. Useful when the output is read as a probability.
- ReLU activation
- f(x) = max(0, x)
- Zero for negative inputs, equal to x for positive inputs. Simple and widely used in hidden layers.
- Layer structure rule
- Input nodes = number of features; hidden layers ≥ 1; DNN = many hidden layers
- The output layer matches the task (one node for a single continuous prediction, for example).
- RL objective
- Maximise expected cumulative (discounted) reward
- The agent learns a policy for choosing actions given the state.
How to solve Neural Networks, Deep Learning and Reinforcement Learning questions
Use this order for any item-set question on neural networks, deep learning, reinforcement learning or algorithm choice.
- 1Read the vignette for the data description: labelled or unlabelled, size, structured or unstructured, continuous or categorical target.
- 2Identify the task: predict a number, classify, find groups, reduce features, or make sequential decisions with rewards.
- 3Map the task to the family: supervised, unsupervised, or reinforcement learning.
- 4Narrow within the family using the clues: nonlinear and large data points to NN or DNN; interpretability needed points to penalized regression or CART; images point to CNN; sequences point to RNN.
- 5For structure questions, count the layers and nodes in the exhibit: input nodes equal features, then hidden layers, then output.
- 6For activation questions, check the output range needed: 0 to 1 suggests sigmoid, any positive value suggests ReLU.
- 7Check for risks the vignette hints at: overfitting, black-box concerns, heavy data needs. Choose the option that fits those clues.
Quickest way: Label, target, then complexity
When to use it: When you have under two minutes for a which-algorithm question.
- Ask: is there a labelled target? No target means unsupervised; reward and actions mean RL.
- If labelled, ask: continuous or categorical?
- Ask: is the data big and unstructured, or is explainability required? Big and unstructured means NN or DNN; explainability means CART or penalized regression.
- Eliminate any option that mismatches the data type, then pick the remaining one.
Common mistakes in Neural Networks, Deep Learning and Reinforcement Learning
Calling reinforcement learning a type of supervised learning because it uses feedback.
A reward looks like a label.
Fix: A reward is a delayed signal from the environment, not a correct answer supplied for each input. RL has no labelled dataset.
Thinking the activation function is optional for a nonlinear model.
Students focus on the weighted sum.
Fix: Without a nonlinear activation, stacked layers collapse to a linear model. The activation provides the nonlinearity.
Saying a deep network is always better.
Deep learning is promoted as advanced.
Fix: Small or structured datasets and interpretability needs favour simpler models. DNNs overfit and need large data.
Mixing up CNN and RNN uses.
Both are deep network variants with similar names.
Fix: CNN for images and spatial patterns; RNN for sequences such as time series and text.
Choosing PCA or clustering for a labelled prediction task.
Students match on keywords like 'groups' or 'many variables'.
Fix: If the vignette gives a labelled target to predict, use supervised learning. PCA and clustering have no target.
Worked examples
Example 1
An analyst at a global asset manager wants to predict next-quarter credit ratings (investment grade or high yield) for 40,000 bond issuers using 25 numeric financial ratios. Past ratings are available as labels. Interpretability is not a requirement, and the relationships appear highly nonlinear. Q1: Which learning family fits? Q2: Which algorithm is most suitable? Q3: How many input nodes would a neural network have?
Show the solution
- Q1: Labelled target exists, so it is supervised learning. The target is categorical, so classification.
- Q2: Large dataset, nonlinear relationships, no interpretability constraint. A neural network (or deep network) suits this better than penalized regression or CART.
- Q3: One input node per feature, so 25 input nodes.
Answer: Q1: Supervised learning (classification). Q2: A neural network. Q3: 25 input nodes.
Example 2
A hedge fund builds a program to execute large orders. The program tries order sizes and timings, observes the price impact after each action, and adjusts to minimise total cost across the day. No labelled dataset exists. Q1: Which approach is this? Q2: What is the role of the cost signal? Q3: What trade-off must the program manage?
Show the solution
- Q1: An agent acts in an environment, gets feedback and learns a policy with no labelled data. This is reinforcement learning.
- Q2: Lower cost acts as the reward (or higher cost as a penalty). The agent seeks to maximise cumulative reward, here minimising total cost.
- Q3: It must balance exploration (trying new sizes and timings) with exploitation (using what has worked best so far).
Answer: Q1: Reinforcement learning. Q2: It is the reward signal the agent maximises cumulatively. Q3: Exploration versus exploitation.
Exam tips
- Every question is tied to the vignette, so underline the data clues first: labelled or not, size, data type, interpretability.
- Memorise one-line use cases: CNN images, RNN sequences, RL sequential decisions with rewards, PCA dimension reduction, clustering grouping.
- Expect conceptual wording, not calculations. Be ready to describe what a hidden node does and why activation matters.
- There is no penalty for wrong answers, so never leave a question blank; eliminate mismatched options and choose.
Neural Networks, Deep Learning and Reinforcement Learning in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Neural Networks, Deep Learning and Reinforcement Learning: frequently asked questions
What is a hidden layer in a neural network?
It is a layer between the input and output layers. Each node there takes a weighted sum of the previous layer's outputs and applies an activation function. Hidden layers let the network learn nonlinear patterns.
What is the difference between a neural network and deep learning?
A neural network has at least one hidden layer. Deep learning refers to networks with many hidden layers. Deep nets can learn more complex patterns but need more data and are harder to interpret.
How does reinforcement learning differ from supervised learning?
Supervised learning trains on labelled examples with known answers. Reinforcement learning has an agent act in an environment and learn from rewards over time. It has no labelled answer for each decision.
How do I choose a machine learning algorithm in the exam?
Start with whether the data has a labelled target. Then check whether the target is continuous or categorical, how large and structured the data is, and whether interpretability matters. Match the algorithm to those clues.