FRM Part I · FRM Exam Part I · Machine Learning and Prediction
A neural network with one output neuron uses a squared-error loss L = (y - ŷ)^2 and a linear output ŷ = w·h, where h = 3 is the hidden activation and the current weight is w = 2. The target is y = 10. Using gradient descent with learning rate 0.01, what is the updated weight w?
The updated weight is 2.24. The prediction is 6, the gradient of the squared error is -2×(10-6)×3 = -24, and the update is 2 - 0.01×(-24) = 2.24. Omitting the factor of 2 gives 2.12, and using the wrong sign gives 1.76.
- A2.24Correct
- B2.12
- C1.76
- D2.48
Explanation
ŷ = 2×3 = 6. dL/dw = -2(y - ŷ)h = -2(4)(3) = -24. Update: w = 2 - 0.01(-24) = 2.24. The value 2.12 uses a gradient of -12 (omitting the factor 2). 1.76 has the wrong sign. 2.48 doubles the gradient.
Did you get it right without looking?
One question tells you little. A timed set on Machine Learning and Prediction shows your real accuracy, how long you take and where you lose marks.
More Machine Learning and Prediction questions
- A risk analyst fits a linear model to predict loan losses using 60 correlated explanatory variables and only 120 observations. The analyst w…
- A single neuron receives inputs x1 = 2 and x2 = -1 with weights w1 = 0.5 and w2 = 1.5 and bias b = 0.5. It uses a ReLU activation, f(z) = ma…
- A node in a classification tree holds 100 observations: 50 defaults and 50 non-defaults. A candidate split sends 40 observations left (35 de…
- Which statement correctly distinguishes a random forest from plain bagging of decision trees?
- A neural network hidden node receives inputs x1 = 2 and x2 = -1. The weights are w1 = 0.5 and w2 = 1.5, and the bias is 0.25. The node uses …
- A risk team has 40 yield-curve and macro predictors and is predicting credit spread changes. They run PCA on the full dataset, keep the firs…