Skip to content

CFA Level II Exam · Big Data Projects

Text Wrangling: Tokenization, Stemming and Lemmatization

Updated 7 October 2026 · Fact-checked

Text wrangling turns raw text into clean, structured input for a model. You split text into tokens, lowercase it, remove stop words, then reduce words to a base form by stemming or lemmatization. The tokens are then counted into a bag of words and arranged as a document term matrix.

Understand Text Wrangling: Tokenization, Stemming and Lemmatization

Raw text is unstructured. A model cannot read sentences, so you must convert them into numbers. Text wrangling is the set of steps that does this, and it is part of the preprocessing stage of a big data project.

Tokenization splits text into pieces called tokens. A token is usually a word, but it can be a character or a sentence. Each token becomes a possible feature.

Normalization reduces the number of distinct tokens so that similar words count as one. The common steps are:

  • Lowercasing: so that "Profit" and "profit" are the same token.
  • Stop word removal: dropping very common words such as "the", "is" and "and" that add little meaning.
  • Stemming: a rule-based process that chops suffixes to reach a base form called the stem. "Analyzed" and "analyzing" may both become "analyz". The stem may not be a real word. A word such as "analysis" may not reduce to the same stem. Stemming is fast but crude.
  • Lemmatization: reduces a word to its lemma, the dictionary form, using vocabulary and word structure. "Was" becomes "be" and "better" can become "good". It is more accurate but more computationally expensive.

After normalization, you build a bag of words (BOW). This is a collection of the distinct tokens in the cleaned text, with no regard to word order. A document term matrix (DTM) puts documents in rows and tokens in columns. Each cell holds a count of how often the token appears in that document. The DTM is the structured data a model uses.

Because BOW ignores order, it loses context. Phrases such as "not good" can be split into "not" and "good". An n-gram keeps adjacent words together, for example a bigram such as "not_good". Using n-grams keeps some word order and meaning.

Key formulas to remember

Tokenization
Text → list of tokens (words, characters or sentences)
First step. Each distinct token is a candidate feature.
Normalization steps
Lowercase, remove stop words, stem or lemmatize
Goal is fewer distinct tokens carrying the same meaning. The curriculum does not prescribe a strict sequence for these cleansing steps.
Stemming vs lemmatization
Stemming = rule-based suffix chopping; Lemmatization = dictionary-based reduction to lemma
Stemming is faster and cruder and the stem may not be a word. Lemmatization is more accurate and costlier.
Bag of words
BOW = set of distinct tokens from the cleaned text, order ignored
Built after cleaning and normalization.
Document term matrix
Rows = documents; columns = tokens; cell = count of token in document
Structured numeric output that feeds the model.
N-grams
n-gram = sequence of n adjacent tokens
Bigrams and trigrams keep some word order that plain BOW loses.

How to solve Text Wrangling: Tokenization, Stemming and Lemmatization questions

Most questions give a short text or a list of terms and ask you to identify the step, the output, or the better technique. Work through the pipeline in order.

  1. 1Identify what the vignette shows: raw text, tokens, normalized tokens, a BOW or a DTM.
  2. 2Name the step that has just been applied, such as tokenization, lowercasing, stop word removal, stemming or lemmatization.
  3. 3If a word was reduced, check whether the result is a real dictionary word. A non-word stem points to stemming. A proper dictionary form points to lemmatization.
  4. 4Compare the options on accuracy versus cost: stemming is simpler and faster, lemmatization is more accurate.
  5. 5For a DTM question, read rows as documents and columns as tokens, then read the cell as a count.
  6. 6Check whether word order matters to the stated goal. If it does, BOW alone is weak and n-grams are the fix.
  7. 7Choose the answer that matches the vignette data, not a general definition.

Quickest way: Real word or not: the one-glance test

When to use it: Use when a question asks whether stemming or lemmatization was applied, or which is better for a stated need.

  1. Look at the reduced word. Not a real word, such as "analyz": stemming.
  2. Real dictionary form, such as "be" from "was": lemmatization.
  3. If the need is speed on large text, lean to stemming. If the need is accuracy, lean to lemmatization.
  4. For a DTM, count rows as documents and columns as distinct tokens, then read the cell.

Common mistakes in Text Wrangling: Tokenization, Stemming and Lemmatization

  • Saying stemming always returns a real word.

    Students assume any base form is a proper word.

    Fix: Stemming only chops suffixes by rules, so stems like "analyz" can be non-words. Only lemmatization uses a dictionary.

  • Calling lemmatization the faster method.

    It sounds more refined, so students link it with efficiency.

    Fix: Lemmatization is more accurate but more computationally expensive. Stemming is faster and cruder.

  • Thinking a bag of words keeps word order.

    The word "bag" is misread as a sentence store.

    Fix: BOW ignores order and only records which tokens appear. Use n-grams to retain some order.

  • Mixing up rows and columns in a document term matrix.

    Students picture words as rows, as in a word list.

    Fix: Documents are rows and tokens are columns. Each cell is a count.

  • Treating stop word removal as a way to fix word meaning, or expecting it to merge related words.

    Students lump all cleansing and normalization steps together as doing the same job.

    Fix: Stop word removal only drops very common, low-information words. Merging related forms such as "analyzed" and "analyzing" is the job of stemming or lemmatization.

Worked examples

Example 1

An analyst processes earnings call transcripts. After tokenizing and lowercasing, she applies a procedure that turns "analyzed" and "analyzing" into "analyz". (1) Which technique was used? (2) Is it more or less costly than the alternative? (3) What is the drawback?

Show the solution
  1. Check the output. "analyz" is not a dictionary word.
  2. A non-word stem produced by chopping suffixes by rule indicates stemming, not lemmatization.
  3. Stemming is simple and fast, so it is less costly than lemmatization, which needs a dictionary and word structure.
  4. The drawback is accuracy: different words can be collapsed wrongly or left unmerged ("analysis" may not reduce to the same stem), and the stem may not be a real word.

Answer: (1) Stemming. (2) Less computationally expensive than lemmatization. (3) Lower accuracy and stems that may not be real words.

Example 2

Two documents remain after cleaning. Document 1: "profit rise profit". Document 2: "loss rise". Build the document term matrix using tokens in the order profit, rise, loss. (1) What are the rows of Document 1 and Document 2? (2) How many distinct tokens form the bag of words? (3) What word order information is lost?

Show the solution
  1. List distinct tokens across both documents: profit, rise, loss. That gives 3.
  2. Document 1: profit appears 2 times, rise 1, loss 0, giving the row 2, 1, 0.
  3. Document 2: profit 0, rise 1, loss 1, giving the row 0, 1, 1.
  4. The matrix has 2 rows for documents and 3 columns for tokens.
  5. Counts record only frequency, so the order of words within each document is lost.

Answer: (1) Document 1 = (2, 1, 0); Document 2 = (0, 1, 1). (2) 3 distinct tokens. (3) All word order, since BOW records only counts. N-grams could retain some of it.

Exam tips

  • Expect the vignette to give example words. Decide stem versus lemma by whether the output is a real word.
  • Know the pipeline broadly: tokenize first, then cleanse and normalize (lowercase, remove stop words, stem or lemmatize), then build BOW and DTM. Do not rely on a rigid order among the cleansing steps.
  • Remember the trade-off: stemming is faster and cruder, lemmatization is more accurate and costlier.
  • For DTM questions, read the cells carefully. Rows are documents, columns are tokens, and cells are counts.
  • If a question stresses context or negation, think n-grams rather than plain BOW.

Text Wrangling: Tokenization, Stemming and Lemmatization: frequently asked questions

What is the difference between stemming and lemmatization?

Stemming removes suffixes using rules and can give a non-word stem. Lemmatization uses vocabulary and word structure to return the dictionary form. Lemmatization is more accurate and more costly.

What is tokenization in a CFA Level II big data project?

Tokenization splits text into units called tokens, usually words. Each token can become a feature once the text is cleaned and normalized.

Why remove stop words?

Stop words such as "the" and "is" appear very often and carry little meaning. Removing them reduces noise and the number of features in the document term matrix.

What is a document term matrix?

It is a table where each row is a document and each column is a token. Each cell shows how many times that token appears in that document. It turns text into structured numeric data for a model.