FRM Part I · FRM Exam Part I · Machine-Learning Methods
A trading desk builds an algorithm that adjusts its hedge positions daily. It receives a profit-or-loss reward after each action and, by trial and error over many periods, learns a policy that maximizes cumulative reward. Which statement about this method is most accurate?
This is reinforcement learning. The algorithm acts, receives rewards, and learns a policy maximizing cumulative reward by trial and error. Rewards evaluate actions but do not provide labelled correct answers, which distinguishes it from supervised learning.
- AIt is supervised learning because profit and loss serve as labels for each correct hedge
- BIt is unsupervised learning because no explicit dataset of correct hedges is provided
- CIt is reinforcement learning because the agent learns actions from rewards rather than from labelled correct answersCorrect
- DIt is dimensionality reduction because the policy compresses the state variables
Explanation
In reinforcement learning an agent interacts with an environment and learns a policy to maximize cumulative reward. Rewards evaluate actions but do not supply the correct action as a label, so it is not supervised. It also does not aim to find hidden structure, so it is not unsupervised in the usual sense.
Did you get it right without looking?
One question tells you little. A timed set on Machine-Learning Methods shows your real accuracy, how long you take and where you lose marks.
More Machine-Learning Methods questions
- In the logistic regression ln(p/(1-p)) = b0 + b1 x with b1 = 0.693, how does a one-unit increase in x affect the odds of the positive class,…
- Using the confusion matrix of 30 TP, 10 FP, 20 FN and 140 TN from a default classifier at a 0.5 threshold, the bank lowers the cutoff to 0.3…
- A analyst compares penalized regressions on a data set with many correlated predictors, some of which are believed to be irrelevant. She wan…
- A risk team evaluates a fraud classifier where only 1% of transactions are fraudulent. A model that labels every transaction as non-fraud ac…
- A random forest differs from simply bagging many fully grown decision trees because, at each split, a random forest:
- A risk manager lowers the probability threshold at which a logistic model classifies a borrower as a defaulter, from 0.50 to 0.30. Which out…