Skip to content

CFA Level II Exam · Big Data Projects

Text Exploration and Feature Selection for CFA Level II

Updated 7 October 2026 · Fact-checked

Text exploration summarises cleaned text with word counts, word clouds and term frequency. Feature selection keeps only useful tokens, using frequency measures (TF, DF, TF-IDF), chi-square or mutual information. Feature engineering adds n-grams, named entity recognition and parts-of-speech tags. You solve questions by matching the method to the task described in the vignette.

Understand Text Exploration and Feature Selection

Text is unstructured. Before a model can use it, you clean it into tokens and build a document term matrix (DTM). In a DTM, each row is a document and each column is a token. Each cell holds a count or a weight. This is the structured form a machine learning model needs.

Text exploration comes after cleansing and wrangling. You look at the data to understand it. Common tools are word clouds, where bigger words appear more often, and simple counts. Exploration also guides which tokens to keep.

Feature selection removes tokens that add noise. Fewer features means a simpler model, faster training and less overfitting. It also removes very common and very rare tokens. Frequency measures include term frequency (TF), the count of a token in a document or corpus, and document frequency (DF), the share of documents containing the token. TF-IDF multiplies TF by inverse document frequency, so a token that is frequent in one document but rare across the corpus gets a high weight.

Two statistical methods also appear. Chi-square tests whether a token is independent of a class. A high chi-square value means the token's occurrence depends on the class, so it is useful. Mutual information (MI) measures how much a token tells you about the class. MI is 0 when the token's distribution is the same in all classes. It is close to 1 when the token appears in only one class, so it is a strong class indicator.

Feature engineering creates new features from existing tokens. N-grams are sequences of n words kept together, such as the bigram "interest rate". They keep word order and context that single words lose. Named entity recognition (NER) tags tokens as entities such as organisation, person, money or date, using context. Parts-of-speech (POS) tagging labels words as noun, verb and so on. Numbers can be replaced with a token such as a number tag.

Key formulas to remember

Term frequency (TF)
TF = count of token in a text ÷ total tokens in that text
Frequency measure. Can also be a raw count depending on the question's definition. Use the vignette's definition.
Document frequency (DF)
DF = number of documents containing the token ÷ total documents
Very high DF tokens (stop-word-like) and very low DF tokens add little value.
Inverse document frequency (IDF)
IDF = ln(total documents ÷ documents containing the token)
Rare tokens get a higher IDF. Some texts add 1 to the denominator. Follow what the exhibit gives.
TF-IDF
TF-IDF = TF × IDF
High when a token is frequent in one document but rare across the corpus.
N-gram
n = 1 unigram, n = 2 bigram, n = 3 trigram
A sequence of n adjacent tokens treated as one feature.
Mutual information reading
MI = 0 → token independent of class; MI near 1 → strong class indicator
Higher MI means the token is more informative about the class.
Chi-square reading
Higher χ² → token and class more dependent → keep the feature
Low χ² suggests the token is independent of the class and can be dropped.

How to solve Text Exploration and Feature Selection questions

Use this order for any item-set question on text exploration or feature selection. Read what the vignette asks the model to do before you pick a method.

  1. 1Identify the stage: exploration, feature selection or feature engineering.
  2. 2Find the goal in the vignette, such as classifying sentiment, reducing features, or capturing word order or entities.
  3. 3If it is selection, find the data: counts, document totals, chi-square or MI values in the exhibit.
  4. 4For TF-IDF, compute TF, then IDF from the document counts, then multiply. Use the formula the exhibit gives.
  5. 5For chi-square or MI, rank tokens by the statistic and keep the highest. Drop tokens with low values.
  6. 6If the goal is context or phrases, choose n-grams. If it is to identify companies, people or amounts, choose NER.
  7. 7Check the conclusion: very frequent and very rare tokens are weak features, and fewer features reduce overfitting.
  8. 8Match your answer to the option wording and confirm direction (high versus low).

Quickest way: Match the tool to the problem

When to use it: Use when the question asks which technique fits a situation and no calculation is needed.

  1. Too many tokens or overfitting: feature selection.
  2. Which tokens relate to the class: chi-square or mutual information.
  3. Token important to one document but rare overall: TF-IDF.
  4. Phrases or word order matter: n-grams.
  5. Need companies, people, dates or money: NER.
  6. Need word type such as noun or verb: POS tagging.
  7. Just viewing common words: word cloud.

Common mistakes in Text Exploration and Feature Selection

  • Treating a high mutual information value as meaning the token is unimportant.

    Students confuse MI with a p-value, where low means significant.

    Fix: For MI, higher means more informative. MI of 0 means the token is independent of the class.

  • Keeping tokens with low chi-square values.

    Students forget that chi-square tests dependence between token and class.

    Fix: Keep high chi-square tokens. Low values mean the token's occurrence does not depend on class.

  • Thinking the most frequent tokens are the best features.

    Frequency feels like importance.

    Fix: Very frequent tokens behave like noise and very rare tokens are not generalisable. Filter both ends using DF.

  • Using TF alone when the question asks about uniqueness across documents.

    Students forget the IDF part.

    Fix: TF-IDF rewards tokens frequent in one document and rare across the corpus.

  • Mixing up n-grams and NER.

    Both create features from text.

    Fix: N-grams join adjacent words into phrases. NER labels entities such as organisation or money using context.

  • Applying feature selection before cleansing.

    Students skip workflow order.

    Fix: Cleanse and wrangle first, then explore, then select and engineer features.

Worked examples

Example 1

An analyst has a corpus of 200 earnings call transcripts. The token "litigation" appears in 8 transcripts. The token "revenue" appears in 160. In one transcript of 500 tokens, "litigation" appears 10 times and "revenue" appears 10 times. Use IDF = ln(total documents ÷ documents containing the token). Q1: Which token has the higher TF-IDF in that transcript? Q2: Compute the TF-IDF of "litigation" to two decimals.

Show the solution
  1. TF for each token = 10 ÷ 500 = 0.02. They are equal.
  2. IDF for litigation = ln(200 ÷ 8) = ln(25) = 3.2189.
  3. IDF for revenue = ln(200 ÷ 160) = ln(1.25) = 0.2231.
  4. TF-IDF litigation = 0.02 × 3.2189 = 0.0644.
  5. TF-IDF revenue = 0.02 × 0.2231 = 0.0045.
  6. So litigation is higher because it is rare across the corpus.

Answer: Q1: "litigation" has the higher TF-IDF. Q2: 0.02 × 3.2189 ≈ 0.06.

Example 2

A model classifies news articles as positive or negative for a company. The analyst has 6,000 tokens and wants to cut them. Token A has mutual information of 0.00 with the class. Token B has 0.41. Token C has chi-square value of 1.2, and Token D has chi-square value of 18.6. Q1: Which of A or B should be kept? Q2: Which of C or D is the stronger feature? Q3: The analyst wants the model to recognise the phrase "not profitable" as a unit. What should be used?

Show the solution
  1. Q1: MI of 0 means token A is independent of class, so it carries no class information. Keep B.
  2. Q2: A higher chi-square means stronger dependence between token and class. D is stronger than C.
  3. Q3: Keeping adjacent words together as one feature is an n-gram, here a bigram. This keeps the negation attached to the word.

Answer: Q1: Keep token B. Q2: Token D. Q3: Use bigrams (n-grams).

Exam tips

  • Read the vignette for the goal first. The method follows from the goal.
  • If an exhibit gives IDF or TF-IDF formulas, use exactly that version, including any added constants.
  • Remember direction: high chi-square and high MI are good, MI of 0 means no information.
  • Expect conceptual wording such as which feature engineering step captures an entity. NER is for entities, n-grams for phrases.
  • Check workflow order: cleansing, wrangling, exploration, then selection and engineering.

Text Exploration and Feature Selection: frequently asked questions

What is the difference between feature selection and feature engineering?

Feature selection chooses which existing tokens to keep, to reduce noise and overfitting. Feature engineering creates new features, such as n-grams or entity tags, from the text. Both prepare data for the model.

How do chi-square and mutual information differ for text features?

Both measure how related a token is to a class. Chi-square tests dependence, so higher means more dependence. Mutual information measures information shared with the class, from 0 (none) towards 1 (strong indicator).

Why is TF-IDF better than raw term frequency?

Raw frequency favours common words that appear everywhere. TF-IDF multiplies by inverse document frequency, so tokens that are distinctive to certain documents get higher weight.

When should I use n-grams?

Use n-grams when word order or phrases carry meaning, such as "interest rate" or "not profitable". They keep context that single words lose, but they also increase the number of features.