CFA Level I · CFA Level I Exam · Introduction to Financial Data Science
An analyst preparing unstructured text from analyst reports for natural language processing removes common words such as "the", "is" and "and" before building a document-term matrix. This step is best described as:
The step is stop word removal. It deletes very common words that carry little meaning, such as "the" and "is", to reduce noise and dimensionality before building a document-term matrix. Lemmatization instead reduces words to base forms, and n-grams combine adjacent tokens into phrases.
- Alemmatization
- Bstop word removalCorrect
- Cn-gram creation
Explanation
Removing high-frequency, low-information words such as "the" and "is" is stop word removal. Lemmatization converts words to their base dictionary form, and n-gram creation joins adjacent tokens into multi-word units. Neither of those removes common words.
Did you get it right without looking?
One question tells you little. A timed set on Introduction to Financial Data Science shows your real accuracy, how long you take and where you lose marks.
More Introduction to Financial Data Science questions
- An analyst trains a model to predict whether a borrower will default, using historical loans that are each labeled "default" or "no default.…
- An analyst converts a set of earnings call transcripts into a form suitable for modeling. She removes punctuation, converts all letters to l…
- A data scientist wants to reduce the words "earned", "earning" and "earnings" to a common base form, using a rule-based procedure that strip…
- A fund uses a model that classifies firms as likely to default or not, using labeled historical data on past defaults and firm characteristi…
- An analyst receives a dataset in which each observation is a customer's credit rating recorded as AAA, AA, A, BBB, and so on. The analyst ne…
- An analyst trains a model on labeled data in which each borrower record includes a known outcome of default or no default. The model is then…