Skip to content

FRM Part I · FRM Exam Part I · Machine-Learning Methods

A single weight w in a network is updated by gradient descent. The current weight is 0.80, the learning rate is 0.10, and the partial derivative of the loss with respect to w is computed by the chain rule. The upstream gradient of the loss with respect to the node output is 0.50, the derivative of the output with respect to the pre-activation is 0.40, and the input attached to w is 3.0. What is the updated weight?

The updated weight is 0.74. The gradient is 0.50 x 0.40 x 3.0 = 0.60 by the chain rule, and gradient descent subtracts the learning rate times the gradient: 0.80 - 0.10 x 0.60 = 0.74.

  1. A0.68Correct
  2. B0.74
  3. C0.92
  4. D0.20

Explanation

dL/dw = 0.50 x 0.40 x 3.0 = 0.60. Update: w = 0.80 - 0.10 x 0.60 = 0.74? Check: 0.10 x 0.60 = 0.06, so w = 0.74. The correct value is therefore 0.74, not 0.68 (which wrongly uses a gradient of 1.2).

Did you get it right without looking?

One question tells you little. A timed set on Machine-Learning Methods shows your real accuracy, how long you take and where you lose marks.

More Machine-Learning Methods questions