Skip to content

CFA Level I · CFA Level I Exam · Introduction to Financial Data Science

An analyst converts a set of earnings call transcripts into a form suitable for modeling. She removes punctuation, converts all letters to lowercase and deletes common words such as "the" and "and". The step of deleting the common words is best described as removal of:

The deletion of common words such as "the" and "and" is removal of stop words. These words add little information for modeling, so they are dropped during text cleansing. N-grams and lemmas refer to token sequences and base word forms, not to this step.

  1. Astop wordsCorrect
  2. Bn-grams
  3. Clemmas

Explanation

Common, low-information words such as "the" and "and" are called stop words, and removing them is a standard text cleansing step. N-grams are sequences of tokens, and lemmatization reduces words to their dictionary base form, which is a different step.

Did you get it right without looking?

One question tells you little. A timed set on Introduction to Financial Data Science shows your real accuracy, how long you take and where you lose marks.

More Introduction to Financial Data Science questions