Skip to content

CFA Level I · CFA Level I Exam · Introduction to Financial Data Science

An analyst preparing unstructured text from analyst reports for natural language processing removes common words such as "the", "is" and "and" before building a document-term matrix. This step is best described as:

The step is stop word removal. It deletes very common words that carry little meaning, such as "the" and "is", to reduce noise and dimensionality before building a document-term matrix. Lemmatization instead reduces words to base forms, and n-grams combine adjacent tokens into phrases.

  1. Alemmatization
  2. Bstop word removalCorrect
  3. Cn-gram creation

Explanation

Removing high-frequency, low-information words such as "the" and "is" is stop word removal. Lemmatization converts words to their base dictionary form, and n-gram creation joins adjacent tokens into multi-word units. Neither of those removes common words.

Did you get it right without looking?

One question tells you little. A timed set on Introduction to Financial Data Science shows your real accuracy, how long you take and where you lose marks.

More Introduction to Financial Data Science questions