Skip to content

CFA Level I · CFA Level I Exam

Introduction to Financial Data Science: formula sheet

Full chapter guide

Key formulas

Big data characteristics (the Vs)
Volume = how much; Velocity = how fast; Variety = what formats; Veracity = how reliable
Learn each V with a one-line test. Volume is size, velocity is speed of generation and processing, variety is mix of structured and unstructured forms, veracity is quality and trustworthiness.
Structured vs unstructured data
Structured = fixed fields (tables, prices); Unstructured = no fixed format (text, images, audio)
Alternative data is mostly unstructured and needs processing before analysis.
Use case matching
Text/news/filings → NLP; pattern finding and prediction → machine learning; fast rule-based order execution → algorithmic trading
Most application questions reduce to picking the right tool for the described task.
Normalization (min-max scaling)
X_norm = (X − X_min) ÷ (X_max − X_min)
Result lies between 0 and 1. Sensitive to outliers because min and max are used.
Standardization (z-score scaling)
X_std = (X − μ) ÷ σ
Result has mean 0 and standard deviation 1. It does not bound values to a fixed range.
IQR outlier rule
Outlier if X < Q1 − 1.5 × IQR or X > Q3 + 1.5 × IQR, where IQR = Q3 − Q1
Common rule of thumb; a larger multiplier such as 3 flags only extreme outliers.
Cleansing problems checklist
Missing, invalid, inaccurate, non-uniform, duplicate
Know the example for each type and its usual fix.
Bias-variance idea
Total error = bias error + variance error + base (irreducible) error
High bias error means underfitting. High variance error means overfitting. Complexity trades one against the other.
K-fold cross-validation
Each fold size ≈ N ÷ k; k training rounds; each fold used once for validation
Training uses k − 1 folds each round. Average the k validation results.
Fit pattern: overfit
Low training error, high out-of-sample error
The model memorized noise and does not generalize.
Fit pattern: underfit
High training error, high out-of-sample error
The model is too simple. Add features or complexity.
Fit pattern: good fit
Low training error, similar low out-of-sample error
The model generalizes well.
Classification metrics
Precision = TP ÷ (TP + FP); Recall = TP ÷ (TP + FN); Accuracy = (TP + TN) ÷ (TP + FP + TN + FN); F1 = 2 × P × R ÷ (P + R)
Use precision when false positives are costly and recall when false negatives are costly. F1 balances both.
Text processing order
Raw text → cleansing → tokenization → normalization (lowercase, stop words, stem/lemma) → BOW / DTM → feature selection → model
Know the sequence and what each step does.
Stemming vs lemmatization
Stemming = rule-based suffix removal (may not be a real word); Lemmatization = dictionary-based base form (real word)
Lemmatization is more accurate and more costly.
Document term matrix
Rows = documents; Columns = tokens; Cells = count or frequency
BOW ignores word order. N-grams retain some order.
Term frequency
TF = count of token in a document ÷ total tokens in that document
Shows how often a token appears relative to document length.
Chart choice
Distribution → histogram / box plot; Relationship → scatter plot; Trend → line chart; Categories → bar chart; Text frequency → word cloud
Match chart to the question being asked.

Quick revision

  • Structured data fits fixed fields and tables; unstructured data such as text, images and audio does not.
  • Supervised learning uses labelled data to predict a target; unsupervised learning finds patterns without labels.
  • Regression predicts a continuous target; classification predicts a category.
  • Clustering and dimension reduction are typical unsupervised tasks.
  • Overfitting means the model fits training data too closely and performs poorly on new data.
  • Underfitting means the model is too simple to capture the pattern, even in training data.
  • Split data into training, validation and test sets; the test set checks performance on unseen data.
  • Data cleaning deals with missing values, outliers, duplicates and inconsistent formats.
  • Tokenization splits text into units; stop-word removal and stemming or lemmatization reduce noise.
  • NLP turns text into structured form so it can be analysed, for example for sentiment.
  • Pick the chart for the message: trends over time, comparisons, distributions or relationships.
  • Always apply Ethics thinking: data privacy, bias and understanding a model before relying on it.

Common mistakes

  • Confusing velocity with volume Fix: Volume is how much data exists. Velocity is how quickly it is created and must be processed. Look for words like real-time or streaming.
  • Thinking big data means only structured data Fix: Variety includes unstructured sources such as text, images and audio. Alternative data is mostly unstructured.
  • Calling JSON or XML data unstructured. Fix: If it has tags or key-value pairs that give organization, it is semi-structured.
  • Mixing up normalization and standardization. Fix: Normalization uses min and max and gives 0 to 1. Standardization uses mean and standard deviation and gives mean 0 and standard deviation 1.
  • Calling clustering a supervised method because it produces groups. Fix: Check for labels in advance. Clustering discovers groups without labels; classification predicts known labels.
  • Saying a model with the lowest training error is the best model. Fix: Judge on out-of-sample error. A very low training error with high test error signals overfitting.
  • Saying stemming always returns a valid word. Fix: Remember that stemming just trims endings by rule, so outputs like 'analyz' can occur. Lemmatization returns real words.
  • Thinking bag-of-words keeps word order. Fix: A bag has no order. If order or phrases matter, use n-grams.

Exam tips

  • Expect short scenario questions: the stem gives a clue, and you name the V or the tool. Practise the clue-to-V conversion until it is automatic.
  • Watch for options that overstate technology, such as 'eliminates bias' or 'guarantees accuracy'. These are usually wrong.
  • Know the difference between structured and unstructured data and give a quick example of each.
  • With no penalty for wrong answers, never leave a question blank. Eliminate one option you can rule out and choose between the other two.
  • Read for limitations in the stem. A hint of poor data quality, overfitting or opaque models usually shapes the right answer.
  • Expect definition and classification questions. Memorize one example for each data type and each cleansing problem.
  • Know the contrast: normalization is bounded 0 to 1 and outlier-sensitive; standardization is not bounded and handles outliers better.
  • In a calculation, do the subtraction first and write it down. Mistakes usually come from wrong inputs, not from division.