CFA Level II · CFA Level II Exam
Big Data Projects: formula sheet
Key formulas
- Four Vs of big data
- Volume (quantity) | Velocity (speed of arrival) | Variety (formats and sources) | Veracity (reliability)
- Match each exam description to one V. Credibility, noise and bias point to veracity.
- Structured-data workflow
- Conceptualization → Data collection → Data preparation and wrangling → Data exploration → Model training
- Order matters. Exploration comes before training, and training includes evaluation and tuning.
- Text-data workflow
- Text problem formulation → Data curation → Text preparation and wrangling → Text exploration → Model training
- This is the unstructured text counterpart of the structured workflow.
- Data types
- Structured (fixed schema) | Semi-structured (tags or fields, no strict table) | Unstructured (no organization)
- Text, images and video are unstructured. Convert them to structured form before modeling.
- Text preparation sequence
- Text problem formulation → Text curation → Text preparation (cleansing) → Text wrangling (preprocessing) → Text exploration → Model training
- Cleansing and wrangling are the two stages that prepare the text. Cleansing comes after curation and before wrangling. Wrangling works on the cleaned text.
- Noise and removal
- HTML tags, punctuation, numbers, extra white spaces → removed or replaced, typically with regular expressions
- Numbers may be replaced with a token such as "number" if their presence matters.
- Cleansing vs wrangling
- Cleansing = remove unnecessary elements; Wrangling = tokenization and normalization (lowercasing, stop word removal, stemming, lemmatization)
- Tokenization and all normalization steps, including lowercasing and stop word removal, belong to wrangling. Normalization is typically done after tokenization.
- Tokenization
- Text → list of tokens (words, characters or sentences)
- First step. Each distinct token is a candidate feature.
- Normalization steps
- Lowercase, remove stop words, stem or lemmatize
- Goal is fewer distinct tokens carrying the same meaning. The curriculum does not prescribe a strict sequence for these cleansing steps.
- Stemming vs lemmatization
- Stemming = rule-based suffix chopping; Lemmatization = dictionary-based reduction to lemma
- Stemming is faster and cruder and the stem may not be a word. Lemmatization is more accurate and costlier.
- Bag of words
- BOW = set of distinct tokens from the cleaned text, order ignored
- Built after cleaning and normalization.
- Document term matrix
- Rows = documents; columns = tokens; cell = count of token in document
- Structured numeric output that feeds the model.
- N-grams
- n-gram = sequence of n adjacent tokens
- Bigrams and trigrams keep some word order that plain BOW loses.
- Term frequency (TF)
- TF = count of token in a text ÷ total tokens in that text
- Frequency measure. Can also be a raw count depending on the question's definition. Use the vignette's definition.
- Document frequency (DF)
- DF = number of documents containing the token ÷ total documents
- Very high DF tokens (stop-word-like) and very low DF tokens add little value.
- Inverse document frequency (IDF)
- IDF = ln(total documents ÷ documents containing the token)
- Rare tokens get a higher IDF. Some texts add 1 to the denominator. Follow what the exhibit gives.
- TF-IDF
- TF-IDF = TF × IDF
- High when a token is frequent in one document but rare across the corpus.
- N-gram
- n = 1 unigram, n = 2 bigram, n = 3 trigram
- A sequence of n adjacent tokens treated as one feature.
- Mutual information reading
- MI = 0 → token independent of class; MI near 1 → strong class indicator
- Higher MI means the token is more informative about the class.
- Chi-square reading
- Higher χ² → token and class more dependent → keep the feature
- Low χ² suggests the token is independent of the class and can be dropped.
- Precision (P)
- P = TP ÷ (TP + FP)
- Share of predicted positives that are truly positive. Penalises false positives (Type I errors).
- Recall (R)
- R = TP ÷ (TP + FN)
- Also the true positive rate. Penalises false negatives (Type II errors).
- Accuracy
- Accuracy = (TP + TN) ÷ (TP + FP + TN + FN)
- Share of all predictions that are correct. Misleading when classes are imbalanced.
- F1 score
- F1 = (2 × P × R) ÷ (P + R)
- Harmonic mean of precision and recall. Ranges from 0 to 1.
- False positive rate (FPR)
- FPR = FP ÷ (FP + TN)
- Horizontal axis of the ROC curve.
- RMSE
- RMSE = √[ Σ(predicted − actual)² ÷ n ]
- For numeric predictions. Lower is better. Same units as the target variable.
- AUC interpretation
- AUC = 0.5 means random; AUC = 1 means perfect
- Higher AUC means better separation of classes; the curve is plotted of TPR against FPR.
- Total error decomposition
- Total error = Bias error + Variance error + Base error
- Base error is random noise and cannot be reduced. Tuning trades bias against variance.
- Underfitting pattern
- High training error and high validation error (high bias)
- Fix by adding complexity or better features.
- Overfitting pattern
- Low training error and high validation error (high variance)
- Fix with regularization, cross-validation, more data or fewer features.
- Good fit pattern
- Low training error and low validation error, small gap
- Out-of-sample performance is the test of a model.
- Grid search
- Best hyperparameters = combination with the best validation score across the grid
- Number of runs = product of the number of values tried for each hyperparameter.
- Ceiling analysis
- Gain of a step = pipeline accuracy with that step made perfect − current accuracy
- Fix the step with the largest gain first.
Quick revision
- Big data is characterized by volume, velocity and variety; veracity (data reliability) is a further consideration.
- Workflow order: conceptualize, collect data, prepare and wrangle, explore, train the model.
- Cleansing removes noise such as HTML tags, punctuation and extra white space.
- Tokenization splits text into tokens, usually words.
- Stemming cuts words to a base form by rule and may produce non-words; lemmatization returns the dictionary form and is more costly.
- Stop words are very common words removed to cut noise.
- A bag of words ignores order and feeds a document term matrix.
- Precision = TP ÷ (TP + FP); recall = TP ÷ (TP + FN).
- Accuracy = (TP + TN) ÷ total; F1 = 2 × precision × recall ÷ (precision + recall).
- Use F1 when classes are unbalanced, since accuracy can mislead.
- Overfitting: strong on training data, weak on new data; underfitting: weak on both.
- Cross-validation and regularization help control overfitting.
Common mistakes
- Confusing variety with veracity. Fix: Variety is about different formats and sources. Veracity is about whether the data can be trusted.
- Treating velocity as the same as volume. Fix: Volume is how much data exists. Velocity is how quickly it arrives and must be processed.
- Treating cleansing and wrangling as the same step. Fix: Remember: cleansing removes noise from raw text; wrangling transforms cleaned text into tokens and normalized form.
- Always deleting numbers and punctuation. Fix: Check the model goal. If numbers matter, replace them with a token. Removal is the default, not a law.
- Saying stemming always returns a real word. Fix: Stemming only chops suffixes by rules, so stems like "analyz" can be non-words. Only lemmatization uses a dictionary.
- Calling lemmatization the faster method. Fix: Lemmatization is more accurate but more computationally expensive. Stemming is faster and cruder.
- Treating a high mutual information value as meaning the token is unimportant. Fix: For MI, higher means more informative. MI of 0 means the token is independent of the class.
- Keeping tokens with low chi-square values. Fix: Keep high chi-square tokens. Low values mean the token's occurrence does not depend on class.
- Mixing up precision and recall. Fix: Precision divides by predicted positives (TP + FP). Recall divides by actual positives (TP + FN). Ask what the denominator counts.
- Recommending accuracy for an imbalanced data set. Fix: When one class is rare, use F1, recall, precision or AUC. State that a trivial model can have high accuracy.
Exam tips
- Read the question stem before the vignette so you know whether to look for a V, a data type or a workflow step.
- Reduce each V to one word in your head: size, speed, formats, trust.
- Memorize the workflow order for both structured and text data. Many questions are just ordering checks.
- When two workflow options seem right, pick the one whose activity is cleaning (preparation) versus inspecting (exploration).
- Do not leave any question blank. There is no penalty for wrong answers.
- Know the order: cleansing first, then wrangling. Questions often test which stage a step belongs to.
- Read the model goal in the vignette before choosing what to remove. The best answer depends on it.
- Expect web-scraped text to need HTML tag removal, and regular expressions as the tool.