Skip to content

CFA Level II Exam · Big Data Projects

Text Data Preparation and Cleansing for CFA Level II

Updated 7 October 2026 · Fact-checked

Text cleansing is the first step in preparing unstructured text for analysis. You remove noise such as HTML tags, punctuation, numbers and extra whitespace, using regular expressions or similar tools. The output is clean text ready for tokenization and normalization. In an exam, match each noise type to its removal step.

Understand Text Data Preparation and Cleansing

Raw text is unstructured. It comes from filings, news, earnings call transcripts, social media and web pages. A model cannot learn from it directly, because it holds noise that carries no meaning for the task. Text preparation turns this raw text into a clean, consistent form.

In the CFA big data workflow, the text steps are: text problem formulation, text curation, text preparation (cleansing), text wrangling (preprocessing), text exploration, and model training. Cleansing and wrangling are the two stages that turn raw text into model-ready form. Cleansing removes unwanted elements from raw text. Wrangling transforms the cleaned text into a usable structure. Wrangling typically involves tokenization and then normalization. Normalization includes lowercasing, removing stop words, stemming and lemmatization. Do not mix cleansing and wrangling.

Typical noise includes HTML tags from web-scraped pages, punctuation, numbers, and extra white spaces such as double spaces, tabs and line breaks. Cleansing is usually done with regular expressions (regex), which are patterns that find and remove or replace these items.

There is a judgment call. Not all noise is useless. Punctuation such as a percent sign or a number such as a year may matter if the task needs it. When numbers matter, you may replace them with a token like "number" instead of deleting them. Some punctuation, like a question mark, may carry signal for certain tasks. The right choice depends on the purpose of the model.

Case differences are not handled in cleansing. Lowercasing is generally classed as a normalization step in wrangling, so that "Revenue" and "revenue" count as the same word. Poor cleansing leaves noise that adds useless features and can hurt model performance. Over-cleansing can remove real information. Always tie each step to what the task needs.

Key formulas to remember

Text preparation sequence
Text problem formulation → Text curation → Text preparation (cleansing) → Text wrangling (preprocessing) → Text exploration → Model training
Cleansing and wrangling are the two stages that prepare the text. Cleansing comes after curation and before wrangling. Wrangling works on the cleaned text.
Noise and removal
HTML tags, punctuation, numbers, extra white spaces → removed or replaced, typically with regular expressions
Numbers may be replaced with a token such as "number" if their presence matters.
Cleansing vs wrangling
Cleansing = remove unnecessary elements; Wrangling = tokenization and normalization (lowercasing, stop word removal, stemming, lemmatization)
Tokenization and all normalization steps, including lowercasing and stop word removal, belong to wrangling. Normalization is typically done after tokenization.

How to solve Text Data Preparation and Cleansing questions

Use this method for any item-set question about preparing or cleaning text data.

  1. 1Read the vignette and identify the source of the text (web pages, filings, transcripts, social media) and the goal of the model.
  2. 2List the noise visible in the sample text: tags, punctuation, numbers, extra spaces, case differences, errors.
  3. 3Decide which stage each fix belongs to. Removing noise is cleansing. Splitting into tokens, lowercasing or reducing words to roots is wrangling.
  4. 4Match each noise type to its removal tool, usually a regular expression.
  5. 5Check whether the task needs the element. If numbers or symbols carry meaning, replace them with a token instead of deleting.
  6. 6Choose the answer that fits both the noise present and the model goal, and watch for options that mix up stages.

Quickest way: Noise-to-fix matching

When to use it: Use when the question lists a sample of messy text and asks which step cleans it.

  1. Underline each noisy element in the sample.
  2. Name it: tag, punctuation, number or whitespace.
  3. Pick the option that removes exactly those items and nothing the task needs.
  4. Reject any option that describes tokenization, stemming or stop word removal as cleansing.

Common mistakes in Text Data Preparation and Cleansing

  • Treating cleansing and wrangling as the same step.

    Both prepare text, and the names sound alike.

    Fix: Remember: cleansing removes noise from raw text; wrangling transforms cleaned text into tokens and normalized form.

  • Always deleting numbers and punctuation.

    Students memorize the list of removals as a fixed rule.

    Fix: Check the model goal. If numbers matter, replace them with a token. Removal is the default, not a law.

  • Placing stop word removal or lowercasing inside cleansing.

    They also reduce noise in the text.

    Fix: Treat normalization and stop word removal as wrangling steps.

  • Forgetting HTML tags in web-scraped data.

    Candidates focus on punctuation and numbers.

    Fix: Whenever the source is a web page, expect tags and remove them first.

  • Assuming cleaning improves results no matter how much is removed.

    More cleaning feels safer.

    Fix: Over-cleansing can strip useful signal. Choose the steps the task needs.

Worked examples

Example 1

An analyst scrapes earnings call summaries from web pages to build a sentiment model. A sample reads: "<p>Revenue grew 12% in 2023!</p>". Question 1: Which noise elements must cleansing deal with? Question 2: Which tool is typically used? Question 3: Is converting the text to lowercase a cleansing step?

Show the solution
  1. Q1: The sample has HTML tags (<p> and </p>), punctuation (% and !), a number (12 and 2023) and a double space between "grew" and "12".
  2. Q2: These patterns are found and removed or replaced with regular expressions.
  3. Q3: Lowercasing is normalization, which belongs to wrangling, not cleansing.

Answer: Q1: HTML tags, punctuation, numbers and extra whitespace. Q2: Regular expressions. Q3: No, it is a wrangling (normalization) step.

Example 2

A fund builds a model to flag filings that mention large monetary amounts. The team cleans filing text by deleting all numbers. Question 1: Is this appropriate? Question 2: What would be a better treatment? Question 3: Which stage is this decision part of?

Show the solution
  1. Q1: The model goal depends on amounts being present. Deleting every number removes the signal the model needs, so it is not appropriate.
  2. Q2: Replace numbers with a token such as "number" so the model can see that an amount appears.
  3. Q3: Removing or replacing numbers is a noise-handling decision, so it is part of cleansing.

Answer: Q1: Not appropriate. Q2: Replace numbers with a placeholder token. Q3: Cleansing.

Exam tips

  • Know the order: cleansing first, then wrangling. Questions often test which stage a step belongs to.
  • Read the model goal in the vignette before choosing what to remove. The best answer depends on it.
  • Expect web-scraped text to need HTML tag removal, and regular expressions as the tool.
  • Eliminate options that call tokenization, stemming or lemmatization cleansing.

Text Data Preparation and Cleansing: frequently asked questions

What are the text cleansing steps in CFA Level II?

You remove noise from raw text: HTML tags, punctuation, numbers and extra white spaces. Regular expressions are the usual tool. Numbers may be replaced with a token if their presence matters.

What is the difference between text cleansing and text wrangling?

Cleansing is the text preparation step. It removes unwanted elements from raw data. Wrangling (preprocessing) then transforms the cleaned text into a usable form, such as tokens after normalization. Cleansing comes first.

Should I always remove numbers and punctuation?

No. Removal is the common default, but if the task depends on them, keep them or replace them with a token. Decide from the model goal.

What are regular expressions used for in text cleansing?

They are patterns that locate specific text such as tags, symbols or digits. You use them to delete or replace those items across a large body of text.