Skip to content

CFA Level I · CFA Level I Exam · Introduction to Financial Data Science

A data scientist converts a column of text analyst comments into numerical features that a machine learning model can use. Which step in text processing does this most likely represent?

This most likely represents feature extraction through vectorization. Machine learning models need numeric inputs, so tokens from the text are converted into numbers such as counts in a document-term matrix. Cleaning and stemming are preparatory steps that do not themselves create numerical features.

  1. ACleaning
  2. BFeature extraction through vectorizationCorrect
  3. CStemming of tokens to remove stop words

Explanation

Turning text into numerical representations, such as a document-term matrix, is feature extraction or vectorization. Cleaning removes unwanted elements, and stemming reduces words to root forms; neither itself produces the numeric features.

Did you get it right without looking?

One question tells you little. A timed set on Introduction to Financial Data Science shows your real accuracy, how long you take and where you lose marks.

More Introduction to Financial Data Science questions