CFA Level I · CFA Level I Exam · Introduction to Financial Data Science
An analyst converts a set of earnings call transcripts into a form suitable for modeling. She removes punctuation, converts all letters to lowercase and deletes common words such as "the" and "and". The step of deleting the common words is best described as removal of:
The deletion of common words such as "the" and "and" is removal of stop words. These words add little information for modeling, so they are dropped during text cleansing. N-grams and lemmas refer to token sequences and base word forms, not to this step.
- Astop wordsCorrect
- Bn-grams
- Clemmas
Explanation
Common, low-information words such as "the" and "and" are called stop words, and removing them is a standard text cleansing step. N-grams are sequences of tokens, and lemmatization reduces words to their dictionary base form, which is a different step.
Did you get it right without looking?
One question tells you little. A timed set on Introduction to Financial Data Science shows your real accuracy, how long you take and where you lose marks.
More Introduction to Financial Data Science questions
- A dataset contains one feature measured in euros, ranging from 20,000 to 900,000, and another feature measured as a ratio from 0.1 to 2.5. T…
- An analyst collects a dataset containing each firm's credit rating (AAA, AA, A, BBB), where the ratings have a clear ranking but the gaps be…
- A model achieves a very low error on its training data but a much higher error on a validation data set. This result most likely indicates t…
- An analyst wants to group thousands of retail borrowers into segments based on spending behavior, without any predefined labels for the grou…
- A model achieves very low error on its training dataset but produces much larger errors on a validation dataset. The model is most likely su…
- In a machine learning project, the data are divided into a training set, a validation set, and a test set. The test set is most likely used …