Skip to content

CFA Level II · CFA Level II Exam

Big Data Projects for CFA Level II: Chapter Guide

Big Data Projects covers how analysts turn large, varied data into model-based predictions. You follow a workflow: define the problem, collect data, prepare and wrangle text, explore and select features, train the model, then evaluate and tune it. In the exam, you read the vignette and judge which step or metric fits.

What this chapter covers

This chapter is about the process of using big data and machine learning in investment work. Big data is characterized by volume, velocity and variety. Veracity, meaning the reliability of the data, is a further consideration. The chapter then walks through a project in order: conceptualize the task, collect data, prepare and wrangle it, explore it, and train and evaluate a model. Much of the focus is on unstructured text, such as news, filings and transcripts.

The topics build on one another. Text must be cleaned before it is tokenized. Tokens become a bag of words and a document term matrix. Features are then selected and engineered. Only after that do you train a model and measure how well it performs. Fit and tuning close the loop by asking why a model fails on new data and how to fix it.

This chapter links to the machine learning and quantitative methods material, where supervised and unsupervised models are described. It also connects to equities and portfolio work, where text-derived sentiment can feed forecasts. Because every Level II question sits in a vignette, you will usually be handed a short project description and asked to spot the stage, the flaw, or the right metric.

The chapter's ideas are testable in a mechanical way. A vignette describes a text project, and the questions ask you to name the step, pick the cleaning or wrangling action, read a confusion matrix, or diagnose overfitting. Quantitative Methods is 5-10% of the exam as a whole, and Big Data Projects is only one chapter within it. So do not expect a large number of questions on this chapter alone, and do not assume any marks are guaranteed. The content is learnable in a short time, and knowing the workflow and the formulas helps you win points that do not depend on long calculations. There is no penalty for wrong answers, so you should always answer.

Big Data Projects: topics in the order to study them

  1. 1Big Data Characteristics and Project WorkflowIt gives you the map of stages and the three Vs (volume, velocity, variety), with veracity as a further consideration, so every later topic has a place to sit.
  2. 2Text Data Preparation and CleansingThis is the first hands-on stage, and it explains why raw text must be cleaned before any analysis.
  3. 3Text Wrangling: Tokenization, Stemming and LemmatizationIt turns cleaned text into tokens and a structured form, so you need cleansing first.
  4. 4Text Exploration and Feature SelectionIt uses the wrangled tokens to find useful features, such as term frequency and chosen words.
  5. 5Model Training and Performance EvaluationWith features ready, you learn how to train a model and judge it with precision, recall, accuracy, F1 and related measures.
  6. 6Model Fit and TuningIt comes last because it uses the evaluation tools to diagnose overfitting and underfitting and to adjust the model.

How to prepare Big Data Projects

Plan for a short, focused study block. The chapter rewards clear definitions and a few formulas, then practice on vignettes.

  1. Write the project workflow on one page from memory: conceptualization, data collection, preparation and wrangling, exploration, model training. Note what happens at each stage.
  2. Learn the three Vs (volume, velocity, variety), note veracity as a further consideration, and learn the difference between structured and unstructured data, so you can classify any data source in a vignette.
  3. For text, list the cleansing actions (removing HTML, punctuation, numbers, white space) and the wrangling actions (tokenization, lowercasing, stop word removal, stemming, lemmatization). Practise telling them apart from short examples.
  4. Study feature selection and feature engineering for text: term frequency, document frequency and chi-square or mutual information, plus n-grams and named entity recognition. Link each to why it helps a model.
  5. Memorize the confusion matrix and the formulas: precision = TP ÷ (TP + FP), recall = TP ÷ (TP + FN), accuracy = (TP + TN) ÷ (TP + FP + TN + FN), F1 = 2 × precision × recall ÷ (precision + recall). Do several calculations by hand.
  6. Learn how to read fit problems: low training error with high test error means overfitting, and high error on both means underfitting. Match each to its remedy and to the ideas of cross-validation and regularization.
  7. Finish with timed item sets. Underline the stage described in the vignette first, then read the question.

Common mistakes in Big Data Projects

  • Mixing up stemming and lemmatization.

    Fix: Remember that stemming is a rough rule-based cut that can leave non-words, while lemmatization uses vocabulary and gives a real base word.

  • Confusing cleansing with wrangling.

    Fix: Treat cleansing as removing unwanted material and wrangling as transforming the cleaned text into usable tokens and structure.

  • Swapping the denominators of precision and recall.

    Fix: Precision asks how many predicted positives were right, so it uses FP. Recall asks how many actual positives were found, so it uses FN.

  • Trusting accuracy on unbalanced data.

    Fix: Check class balance in the vignette and prefer F1, precision or recall when one class is rare.

  • Misreading overfitting versus underfitting.

    Fix: Always compare the two. A big gap points to overfitting; poor results on both point to underfitting.

Last-day revision: Big Data Projects

  • Big data is characterized by volume, velocity and variety; veracity (data reliability) is a further consideration.
  • Workflow order: conceptualize, collect data, prepare and wrangle, explore, train the model.
  • Cleansing removes noise such as HTML tags, punctuation and extra white space.
  • Tokenization splits text into tokens, usually words.
  • Stemming cuts words to a base form by rule and may produce non-words; lemmatization returns the dictionary form and is more costly.
  • Stop words are very common words removed to cut noise.
  • A bag of words ignores order and feeds a document term matrix.
  • Precision = TP ÷ (TP + FP); recall = TP ÷ (TP + FN).
  • Accuracy = (TP + TN) ÷ total; F1 = 2 × precision × recall ÷ (precision + recall).
  • Use F1 when classes are unbalanced, since accuracy can mislead.
  • Overfitting: strong on training data, weak on new data; underfitting: weak on both.
  • Cross-validation and regularization help control overfitting.

Big Data Projects in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Big Data Projects: frequently asked questions

How are Big Data Projects questions asked in the CFA Level II exam?

They appear inside item sets. A vignette describes a data project and the questions ask you to identify a stage, choose a preparation step, or compute a metric from a confusion matrix. You answer from the vignette, not from stand-alone recall.

Do I need to know how to code for this chapter?

No. You need to understand the concepts and steps, not write code. Focus on what each technique does and when you would use it.

Which formulas matter most in Big Data Projects?

The confusion matrix measures matter most: precision, recall, accuracy and F1. Practise computing them quickly from a table of true and false positives and negatives.

How long should I spend on this chapter?

It is a small chapter, so a few focused sessions can be enough for first study. Spend the extra time on practice item sets and on repeating the formulas until they are automatic.