CFA Level I · CFA Level I Exam · Introduction to Financial Data Science
An analyst prepares text data from company filings for a model and removes common words such as 'the', 'and', and 'of' before counting word frequencies. This step is best described as:
This step is stop-word removal. It deletes very common words that add little informational value so that the word counts reflect meaningful terms. Stemming and lemmatization instead reduce words to a base or dictionary form without removing the words.
- Astemming
- Blemmatization
- Cstop-word removalCorrect
Explanation
Removing high-frequency words that carry little meaning is stop-word removal. Stemming trims words to a base form by cutting endings, and lemmatization converts words to their dictionary form. Neither deletes common function words.
Did you get it right without looking?
One question tells you little. A timed set on Introduction to Financial Data Science shows your real accuracy, how long you take and where you lose marks.
More Introduction to Financial Data Science questions
- In a machine learning project, the data are divided into a training set, a validation set, and a test set. The test set is most likely used …
- An analyst computes term frequency-inverse document frequency (TF-IDF) for the word "covenant" in a loan filing. The word appears 6 times in…
- A fund uses a model that classifies firms as likely to default or not, using labeled historical data on past defaults and firm characteristi…
- An analyst's dataset of customer incomes contains several extreme values far above the rest, and the analyst wants to reduce the influence o…
- An analyst applies a clustering algorithm to the return histories of 500 stocks to find groups that behave similarly, without predefining an…
- An analyst builds a bag-of-words document term matrix from 2,000 corporate disclosures to predict whether a firm will cut its dividend. Comp…