FRM Part I · FRM Exam Part I · Machine Learning and Prediction
A single-output neural network uses squared-error loss L = 0.5(y - ŷ)², where ŷ = w·x with no activation and no bias. For one observation, x = 4, y = 10, and the current weight is w = 1. Using gradient descent with learning rate 0.05, what is the updated weight after one step?
The updated weight is 2.2. The prediction is 4, the error is 6, and the gradient is -(6)(4) = -24. Gradient descent gives 1 - 0.05×(-24) = 2.2. Using the wrong gradient sign would give -0.2 instead.
- A1.2
- B2.2Correct
- C4.0
- D-0.2
Explanation
ŷ = 1×4 = 4. The gradient dL/dw = -(y - ŷ)x = -(6)(4) = -24. The update is w_new = 1 - 0.05(-24) = 1 + 1.2 = 2.2. The option 1.2 uses the step size only, forgetting to add it to the old weight. -0.2 comes from the wrong sign of the gradient (1 - 1.2).
Did you get it right without looking?
One question tells you little. A timed set on Machine Learning and Prediction shows your real accuracy, how long you take and where you lose marks.
More Machine Learning and Prediction questions
- A risk analyst groups 500 corporate borrowers into segments using only their financial ratios, with no default labels available. Which descr…
- An analyst at a risk consultancy wants to group 400 corporate borrowers into segments based on leverage, interest coverage and asset volatil…
- A fraud model is evaluated on 10,000 transactions, of which 100 are fraudulent. A naive model that labels every transaction as non-fraud is …
- In a random forest, each tree is trained on a bootstrap sample, leaving roughly one third of observations out of that tree's sample. How are…
- A risk analyst fits a very flexible model to predict loan defaults. The model achieves almost perfect accuracy on the training data but perf…
- A dataset of 1,000 observations is split into training, validation, and test sets in a 60/20/20 ratio. A analyst tunes a hyperparameter by t…