CFA Level I · CFA Level I Exam · Introduction to Financial Data Science
An analyst computes term frequency-inverse document frequency (TF-IDF) for the word "covenant" in a loan filing. The word appears 6 times in a filing of 200 tokens, and it appears in 10 of 1,000 filings in the corpus. Using IDF = ln(total documents / documents containing the term), the TF-IDF score is closest to:
The TF-IDF score is about 0.14. Term frequency is 6/200 = 0.03, and inverse document frequency is ln(1,000/10) = 4.605. Multiplying gives 0.138. The value 0.03 ignores the IDF component, which rewards rare words.
- A0.03
- B0.14Correct
- C0.28
Explanation
TF = 6/200 = 0.03. IDF = ln(1,000/10) = ln(100) = 4.605. TF-IDF = 0.03 × 4.605 = 0.138, about 0.14. The 0.03 option uses TF alone, and 0.28 doubles the result.
Did you get it right without looking?
One question tells you little. A timed set on Introduction to Financial Data Science shows your real accuracy, how long you take and where you lose marks.
More Introduction to Financial Data Science questions
- A feature has values ranging from 10 to 50, with a minimum of 10 and a maximum of 50. An analyst applies min-max normalization. The normaliz…
- An analyst builds a term frequency-inverse document frequency (TF-IDF) measure for a corpus of 1,000 earnings call transcripts. The word "re…
- Which feature of big data is most likely described by an investment firm collecting real-time social media posts, satellite images and trans…
- A data scientist finds that 4% of the entries in the 'annual income' column of a client dataset are blank. The blanks appear unrelated to an…
- A data scientist has a dataset of daily closing prices for 500 stocks over 10 years, with each stock observed on the same trading days. This…
- A data scientist builds a model that fits the training data almost perfectly but produces poor predictions on new data. The problem is most …