Skip to content

CFA Level I Exam · Introduction to Financial Data Science

Text Analytics, NLP and Data Visualization for CFA Level 1

Updated 7 October 2026 · Fact-checked

Text analytics turns unstructured text into structured data a model can use. You clean the text, split it into tokens, remove noise, reduce words to stems or lemmas, then build a bag-of-words or similar matrix. NLP uses this for sentiment and classification. Visualizations such as word clouds and charts help you explore the data.

Understand Text Analytics, NLP and Data Visualization

Most financial information is unstructured: earnings call transcripts, news, filings, social media posts. Models need structured data, meaning rows and columns of numbers. Text analytics is the set of steps that converts text into that form. Natural language processing (NLP) is the use of computers to interpret human language, for tasks such as sentiment scoring, topic classification and summarizing.

The workflow has two stages. First comes text preprocessing (cleansing): remove HTML tags, punctuation, numbers and extra white space. Then comes text wrangling (preparation): tokenize the text, normalize it, and convert it into a document-term structure.

A token is a piece of text, usually a word. Tokenization splits text into tokens. Normalization reduces tokens to a consistent form: lowercasing, removing stop words (very common words such as 'the' or 'is' that carry little meaning), and stemming or lemmatization. Stemming chops word endings by rule, so 'earning' and 'earnings' may both become 'earn'. The result may not be a real word. Lemmatization uses a dictionary and word context to return the real base form, the lemma, so 'was' becomes 'be'. Lemmatization is more accurate but more computationally costly.

The cleaned tokens form a bag-of-words (BOW): a collection of words that ignores order and grammar. It is arranged as a document term matrix (DTM), with one row per document, one column per token, and cell values showing counts or frequencies. An n-gram is a sequence of n tokens, such as a bigram 'interest rate'. N-grams keep some word order that BOW loses. Features can then be selected by removing very rare and very frequent terms, and by using frequency measures (such as document frequency and TF-IDF), chi-square tests and mutual information to keep tokens that help classify.

Sentiment analysis assigns a tone (positive, negative, neutral) to text, using word lists or trained models. Text is also used in supervised learning, such as predicting a stock reaction to a filing.

Data visualization helps you explore data before modelling. A word cloud shows tokens sized by frequency. Histograms and box plots show distribution and outliers. Scatter plots show relationships between two variables. Line charts show trends over time. Bar charts compare categories. Heat maps show values or correlations by colour. Choose the chart to fit the data type and the message.

Key formulas to remember

Text processing order
Raw text → cleansing → tokenization → normalization (lowercase, stop words, stem/lemma) → BOW / DTM → feature selection → model
Know the sequence and what each step does.
Stemming vs lemmatization
Stemming = rule-based suffix removal (may not be a real word); Lemmatization = dictionary-based base form (real word)
Lemmatization is more accurate and more costly.
Document term matrix
Rows = documents; Columns = tokens; Cells = count or frequency
BOW ignores word order. N-grams retain some order.
Term frequency
TF = count of token in a document ÷ total tokens in that document
Shows how often a token appears relative to document length.
Chart choice
Distribution → histogram / box plot; Relationship → scatter plot; Trend → line chart; Categories → bar chart; Text frequency → word cloud
Match chart to the question being asked.

How to solve Text Analytics, NLP and Data Visualization questions

Use this method for any question on text analytics, NLP or charts.

  1. 1Identify what is being asked: a preprocessing step, a definition, a technique choice or a chart choice.
  2. 2Check whether the data is unstructured text or already structured numbers.
  3. 3If it is a preprocessing question, place the step in the workflow: cleansing, tokenization, normalization or matrix building.
  4. 4For stemming versus lemmatization, ask: is the output a real dictionary word? If yes, lemmatization; if it is a crude chopped form, stemming.
  5. 5For BOW questions, ask whether word order matters. BOW ignores it; n-grams keep some.
  6. 6For chart questions, name the data type and the message (distribution, relationship, trend, frequency), then pick the matching chart.
  7. 7Eliminate the two options that describe a different step or a different chart, then confirm the remaining one.

Quickest way: Keyword matching for text and chart questions

When to use it: Use when you have about 90 seconds and the question is definitional.

  1. Spot the key word in the stem: 'stop words', 'stem', 'lemma', 'token', 'order', 'frequency', 'trend'.
  2. Link it to its one-line meaning: stop words are common low-information words; stemming is crude chopping; lemma is a real base word.
  3. Link the chart cue: frequency of words means word cloud; trend means line chart; outliers mean box plot.
  4. Remove options that reverse the idea, such as BOW preserving word order.
  5. Pick the remaining option. There is no penalty for a wrong answer, so never leave it blank.

Common mistakes in Text Analytics, NLP and Data Visualization

  • Saying stemming always returns a valid word.

    Students mix it up with lemmatization.

    Fix: Remember that stemming just trims endings by rule, so outputs like 'analyz' can occur. Lemmatization returns real words.

  • Thinking bag-of-words keeps word order.

    The word 'bag' is not taken literally.

    Fix: A bag has no order. If order or phrases matter, use n-grams.

  • Putting tokenization before cleansing or confusing the order of steps.

    Students memorize terms but not the sequence.

    Fix: Learn the flow: cleanse, tokenize, normalize, build the matrix, select features.

  • Treating stop words as the most important words because they are the most frequent.

    Frequency is assumed to mean importance.

    Fix: Stop words are frequent but carry little meaning, so they are usually removed.

  • Choosing a word cloud to show a time trend or a relationship between two variables.

    Word clouds look attractive and familiar.

    Fix: A word cloud shows only word frequency. Use a line chart for trends and a scatter plot for relationships.

  • Assuming unstructured text can go directly into a model.

    Students overlook the need for numeric representation.

    Fix: Text must be converted to a structured form such as a document term matrix first.

Worked examples

Example 1

An analyst processes the sentence 'The earnings were rising strongly.' She lowercases it, removes punctuation, tokenizes it, and drops the words 'the' and 'were'. How many tokens remain, and what is the dropped-word category? A. 3 tokens, stop words. B. 4 tokens, stop words. C. 5 tokens, stemmed words.

Show the solution
  1. After cleansing and lowercasing, the text is: the earnings were rising strongly.
  2. Tokenization gives 5 tokens: the, earnings, were, rising, strongly.
  3. Removing 'the' and 'were' leaves 3 tokens: earnings, rising, strongly.
  4. The removed words are common, low-information words, which are stop words.
  5. Option A matches. Option B is wrong because 4 tokens would mean only one word was dropped, but two were. Option C is wrong because 5 tokens is the count before removal, and dropping words is not stemming.

Answer: A. 3 tokens remain, and the dropped words are stop words.

Example 2

An analyst wants to show which terms appear most often across 500 central bank speeches, and also to show how the count of the word 'inflation' changed year by year. Which pair of charts fits best? A. Scatter plot and word cloud. B. Word cloud and line chart. C. Histogram and heat map.

Show the solution
  1. Overall term frequency across many documents is shown by a word cloud, which sizes words by frequency.
  2. A change in one count over successive years is a trend over time, which a line chart shows.
  3. A scatter plot shows the relationship between two numeric variables, so it does not fit the first need.
  4. A histogram shows a distribution of one numeric variable and a heat map shows values by colour in a grid, and neither is the standard choice for the two tasks here.
  5. Option B fits both tasks.

Answer: B. Word cloud for term frequency and line chart for the yearly trend.

Exam tips

  • Questions are definitional. Learn the one-line difference between stemming, lemmatization, tokenization and stop words.
  • Memorize the preprocessing order, because items often ask which step comes first or next.
  • For chart items, name the data type and the message first, then choose the chart.
  • Watch for traps that reverse a property, such as BOW preserving order or lemmatization giving non-words.
  • Each question has three options, so eliminate the two that describe the wrong step and choose the remaining one.

Practice questions from Introduction to Financial Data Science

Text Analytics, NLP and Data Visualization: frequently asked questions

How is text data converted to structured data in CFA Level I?

You cleanse the text, tokenize it, normalize the tokens, and then build a bag-of-words arranged as a document term matrix. Each row is a document and each column is a token. The cells hold counts or frequencies, which a model can use.

What is the difference between stemming and lemmatization?

Stemming removes word endings by simple rules and may produce non-words. Lemmatization uses a dictionary and context to return the real base word. Lemmatization is more accurate but takes more computing effort.

What does a word cloud show?

A word cloud displays tokens with size reflecting how often they occur in the text. It is useful for a quick view of the most common terms. It does not show trends, relationships or word order.

Why remove stop words?

Stop words such as 'the' and 'is' appear very often but carry little information about meaning. Removing them reduces noise and the size of the document term matrix. Analysts may keep them in special cases where they matter.