CFA Level I · CFA Level I Exam · Introduction to Financial Data Science
A data scientist converts a column of text analyst comments into numerical features that a machine learning model can use. Which step in text processing does this most likely represent?
This most likely represents feature extraction through vectorization. Machine learning models need numeric inputs, so tokens from the text are converted into numbers such as counts in a document-term matrix. Cleaning and stemming are preparatory steps that do not themselves create numerical features.
- ACleaning
- BFeature extraction through vectorizationCorrect
- CStemming of tokens to remove stop words
Explanation
Turning text into numerical representations, such as a document-term matrix, is feature extraction or vectorization. Cleaning removes unwanted elements, and stemming reduces words to root forms; neither itself produces the numeric features.
Did you get it right without looking?
One question tells you little. A timed set on Introduction to Financial Data Science shows your real accuracy, how long you take and where you lose marks.
More Introduction to Financial Data Science questions
- An analyst applies a clustering algorithm to the return histories of 500 stocks to find groups that behave similarly, without predefining an…
- An analyst builds a bag-of-words document term matrix from 2,000 corporate disclosures to predict whether a firm will cut its dividend. Comp…
- A team builds a model to forecast bond returns using 200 features and only 150 observations. To reduce overfitting, the team adds a penalty …
- An analyst prepares text data from company filings for a model and removes common words such as 'the', 'and', and 'of' before counting word …
- To assess how well a model will generalize, an analyst splits the data into a training set, a validation set, and a test set. The test set i…
- A researcher wants to show the distribution, median, quartiles and outliers of monthly returns for several funds side by side. The visualiza…