Skip to content

CFA Level II · CFA Level II Exam

Big Data Projects: formula sheet

Full chapter guide

Key formulas

Four Vs of big data
Volume (quantity) | Velocity (speed of arrival) | Variety (formats and sources) | Veracity (reliability)
Match each exam description to one V. Credibility, noise and bias point to veracity.
Structured-data workflow
Conceptualization → Data collection → Data preparation and wrangling → Data exploration → Model training
Order matters. Exploration comes before training, and training includes evaluation and tuning.
Text-data workflow
Text problem formulation → Data curation → Text preparation and wrangling → Text exploration → Model training
This is the unstructured text counterpart of the structured workflow.
Data types
Structured (fixed schema) | Semi-structured (tags or fields, no strict table) | Unstructured (no organization)
Text, images and video are unstructured. Convert them to structured form before modeling.
Text preparation sequence
Text problem formulation → Text curation → Text preparation (cleansing) → Text wrangling (preprocessing) → Text exploration → Model training
Cleansing and wrangling are the two stages that prepare the text. Cleansing comes after curation and before wrangling. Wrangling works on the cleaned text.
Noise and removal
HTML tags, punctuation, numbers, extra white spaces → removed or replaced, typically with regular expressions
Numbers may be replaced with a token such as "number" if their presence matters.
Cleansing vs wrangling
Cleansing = remove unnecessary elements; Wrangling = tokenization and normalization (lowercasing, stop word removal, stemming, lemmatization)
Tokenization and all normalization steps, including lowercasing and stop word removal, belong to wrangling. Normalization is typically done after tokenization.
Tokenization
Text → list of tokens (words, characters or sentences)
First step. Each distinct token is a candidate feature.
Normalization steps
Lowercase, remove stop words, stem or lemmatize
Goal is fewer distinct tokens carrying the same meaning. The curriculum does not prescribe a strict sequence for these cleansing steps.
Stemming vs lemmatization
Stemming = rule-based suffix chopping; Lemmatization = dictionary-based reduction to lemma
Stemming is faster and cruder and the stem may not be a word. Lemmatization is more accurate and costlier.
Bag of words
BOW = set of distinct tokens from the cleaned text, order ignored
Built after cleaning and normalization.
Document term matrix
Rows = documents; columns = tokens; cell = count of token in document
Structured numeric output that feeds the model.
N-grams
n-gram = sequence of n adjacent tokens
Bigrams and trigrams keep some word order that plain BOW loses.
Term frequency (TF)
TF = count of token in a text ÷ total tokens in that text
Frequency measure. Can also be a raw count depending on the question's definition. Use the vignette's definition.
Document frequency (DF)
DF = number of documents containing the token ÷ total documents
Very high DF tokens (stop-word-like) and very low DF tokens add little value.
Inverse document frequency (IDF)
IDF = ln(total documents ÷ documents containing the token)
Rare tokens get a higher IDF. Some texts add 1 to the denominator. Follow what the exhibit gives.
TF-IDF
TF-IDF = TF × IDF
High when a token is frequent in one document but rare across the corpus.
N-gram
n = 1 unigram, n = 2 bigram, n = 3 trigram
A sequence of n adjacent tokens treated as one feature.
Mutual information reading
MI = 0 → token independent of class; MI near 1 → strong class indicator
Higher MI means the token is more informative about the class.
Chi-square reading
Higher χ² → token and class more dependent → keep the feature
Low χ² suggests the token is independent of the class and can be dropped.
Precision (P)
P = TP ÷ (TP + FP)
Share of predicted positives that are truly positive. Penalises false positives (Type I errors).
Recall (R)
R = TP ÷ (TP + FN)
Also the true positive rate. Penalises false negatives (Type II errors).
Accuracy
Accuracy = (TP + TN) ÷ (TP + FP + TN + FN)
Share of all predictions that are correct. Misleading when classes are imbalanced.
F1 score
F1 = (2 × P × R) ÷ (P + R)
Harmonic mean of precision and recall. Ranges from 0 to 1.
False positive rate (FPR)
FPR = FP ÷ (FP + TN)
Horizontal axis of the ROC curve.
RMSE
RMSE = √[ Σ(predicted − actual)² ÷ n ]
For numeric predictions. Lower is better. Same units as the target variable.
AUC interpretation
AUC = 0.5 means random; AUC = 1 means perfect
Higher AUC means better separation of classes; the curve is plotted of TPR against FPR.
Total error decomposition
Total error = Bias error + Variance error + Base error
Base error is random noise and cannot be reduced. Tuning trades bias against variance.
Underfitting pattern
High training error and high validation error (high bias)
Fix by adding complexity or better features.
Overfitting pattern
Low training error and high validation error (high variance)
Fix with regularization, cross-validation, more data or fewer features.
Good fit pattern
Low training error and low validation error, small gap
Out-of-sample performance is the test of a model.
Grid search
Best hyperparameters = combination with the best validation score across the grid
Number of runs = product of the number of values tried for each hyperparameter.
Ceiling analysis
Gain of a step = pipeline accuracy with that step made perfect − current accuracy
Fix the step with the largest gain first.

Quick revision

  • Big data is characterized by volume, velocity and variety; veracity (data reliability) is a further consideration.
  • Workflow order: conceptualize, collect data, prepare and wrangle, explore, train the model.
  • Cleansing removes noise such as HTML tags, punctuation and extra white space.
  • Tokenization splits text into tokens, usually words.
  • Stemming cuts words to a base form by rule and may produce non-words; lemmatization returns the dictionary form and is more costly.
  • Stop words are very common words removed to cut noise.
  • A bag of words ignores order and feeds a document term matrix.
  • Precision = TP ÷ (TP + FP); recall = TP ÷ (TP + FN).
  • Accuracy = (TP + TN) ÷ total; F1 = 2 × precision × recall ÷ (precision + recall).
  • Use F1 when classes are unbalanced, since accuracy can mislead.
  • Overfitting: strong on training data, weak on new data; underfitting: weak on both.
  • Cross-validation and regularization help control overfitting.

Common mistakes

  • Confusing variety with veracity. Fix: Variety is about different formats and sources. Veracity is about whether the data can be trusted.
  • Treating velocity as the same as volume. Fix: Volume is how much data exists. Velocity is how quickly it arrives and must be processed.
  • Treating cleansing and wrangling as the same step. Fix: Remember: cleansing removes noise from raw text; wrangling transforms cleaned text into tokens and normalized form.
  • Always deleting numbers and punctuation. Fix: Check the model goal. If numbers matter, replace them with a token. Removal is the default, not a law.
  • Saying stemming always returns a real word. Fix: Stemming only chops suffixes by rules, so stems like "analyz" can be non-words. Only lemmatization uses a dictionary.
  • Calling lemmatization the faster method. Fix: Lemmatization is more accurate but more computationally expensive. Stemming is faster and cruder.
  • Treating a high mutual information value as meaning the token is unimportant. Fix: For MI, higher means more informative. MI of 0 means the token is independent of the class.
  • Keeping tokens with low chi-square values. Fix: Keep high chi-square tokens. Low values mean the token's occurrence does not depend on class.
  • Mixing up precision and recall. Fix: Precision divides by predicted positives (TP + FP). Recall divides by actual positives (TP + FN). Ask what the denominator counts.
  • Recommending accuracy for an imbalanced data set. Fix: When one class is rare, use F1, recall, precision or AUC. State that a trivial model can have high accuracy.

Exam tips

  • Read the question stem before the vignette so you know whether to look for a V, a data type or a workflow step.
  • Reduce each V to one word in your head: size, speed, formats, trust.
  • Memorize the workflow order for both structured and text data. Many questions are just ordering checks.
  • When two workflow options seem right, pick the one whose activity is cleaning (preparation) versus inspecting (exploration).
  • Do not leave any question blank. There is no penalty for wrong answers.
  • Know the order: cleansing first, then wrangling. Questions often test which stage a step belongs to.
  • Read the model goal in the vignette before choosing what to remove. The best answer depends on it.
  • Expect web-scraped text to need HTML tag removal, and regular expressions as the tool.