FRM Part I · FRM Exam Part I · Machine Learning and Prediction
A risk analyst builds a feed-forward neural network to predict loan default. The network has an input layer, two hidden layers and an output layer. What is the main role of the nonlinear activation function applied in the hidden layers?
The activation function introduces nonlinearity, letting the network model nonlinear relationships. Without it, multiple layers would collapse into one linear mapping, equivalent to linear regression. It does not ensure zero training error, standardize inputs, or eliminate the need for training data.
- AIt allows the network to model nonlinear relationships between inputs and the outputCorrect
- BIt guarantees that the training error reaches zero
- CIt standardizes the input variables to have zero mean
- DIt removes the need for a training dataset
Explanation
Without nonlinear activation functions, stacked layers collapse into a single linear transformation, so the network would be no more flexible than linear regression. Nonlinear activations let the network capture complex relationships. Activation functions do not guarantee zero training error, do not standardize inputs, and do not remove the need for training data.
Did you get it right without looking?
One question tells you little. A timed set on Machine Learning and Prediction shows your real accuracy, how long you take and where you lose marks.
More Machine Learning and Prediction questions
- An analyst uses ridge regression with one predictor and no intercept, where the predictor has been scaled so that the sum of x squared equal…
- A bank uses the first three principal components of 12 equity factor returns as regressors in a model to predict portfolio losses (principal…
- Compared with a logistic regression using the same inputs, a deep neural network used for a bank's default prediction is most likely to pres…
- A risk analyst fits a single, very deep classification tree to predict loan default. It classifies the training data almost perfectly but pe…
- A risk team runs agglomerative hierarchical clustering on five funds using Euclidean distance. The first merge joins Funds A and B, which ar…
- Compared with a regression using all original predictors, a principal components regression (PCR) that uses the first few components has whi…