CFA Level I · CFA Level I Exam
Introduction to Financial Data Science: formula sheet
Key formulas
- Big data characteristics (the Vs)
- Volume = how much; Velocity = how fast; Variety = what formats; Veracity = how reliable
- Learn each V with a one-line test. Volume is size, velocity is speed of generation and processing, variety is mix of structured and unstructured forms, veracity is quality and trustworthiness.
- Structured vs unstructured data
- Structured = fixed fields (tables, prices); Unstructured = no fixed format (text, images, audio)
- Alternative data is mostly unstructured and needs processing before analysis.
- Use case matching
- Text/news/filings → NLP; pattern finding and prediction → machine learning; fast rule-based order execution → algorithmic trading
- Most application questions reduce to picking the right tool for the described task.
- Normalization (min-max scaling)
- X_norm = (X − X_min) ÷ (X_max − X_min)
- Result lies between 0 and 1. Sensitive to outliers because min and max are used.
- Standardization (z-score scaling)
- X_std = (X − μ) ÷ σ
- Result has mean 0 and standard deviation 1. It does not bound values to a fixed range.
- IQR outlier rule
- Outlier if X < Q1 − 1.5 × IQR or X > Q3 + 1.5 × IQR, where IQR = Q3 − Q1
- Common rule of thumb; a larger multiplier such as 3 flags only extreme outliers.
- Cleansing problems checklist
- Missing, invalid, inaccurate, non-uniform, duplicate
- Know the example for each type and its usual fix.
- Bias-variance idea
- Total error = bias error + variance error + base (irreducible) error
- High bias error means underfitting. High variance error means overfitting. Complexity trades one against the other.
- K-fold cross-validation
- Each fold size ≈ N ÷ k; k training rounds; each fold used once for validation
- Training uses k − 1 folds each round. Average the k validation results.
- Fit pattern: overfit
- Low training error, high out-of-sample error
- The model memorized noise and does not generalize.
- Fit pattern: underfit
- High training error, high out-of-sample error
- The model is too simple. Add features or complexity.
- Fit pattern: good fit
- Low training error, similar low out-of-sample error
- The model generalizes well.
- Classification metrics
- Precision = TP ÷ (TP + FP); Recall = TP ÷ (TP + FN); Accuracy = (TP + TN) ÷ (TP + FP + TN + FN); F1 = 2 × P × R ÷ (P + R)
- Use precision when false positives are costly and recall when false negatives are costly. F1 balances both.
- Text processing order
- Raw text → cleansing → tokenization → normalization (lowercase, stop words, stem/lemma) → BOW / DTM → feature selection → model
- Know the sequence and what each step does.
- Stemming vs lemmatization
- Stemming = rule-based suffix removal (may not be a real word); Lemmatization = dictionary-based base form (real word)
- Lemmatization is more accurate and more costly.
- Document term matrix
- Rows = documents; Columns = tokens; Cells = count or frequency
- BOW ignores word order. N-grams retain some order.
- Term frequency
- TF = count of token in a document ÷ total tokens in that document
- Shows how often a token appears relative to document length.
- Chart choice
- Distribution → histogram / box plot; Relationship → scatter plot; Trend → line chart; Categories → bar chart; Text frequency → word cloud
- Match chart to the question being asked.
Quick revision
- Structured data fits fixed fields and tables; unstructured data such as text, images and audio does not.
- Supervised learning uses labelled data to predict a target; unsupervised learning finds patterns without labels.
- Regression predicts a continuous target; classification predicts a category.
- Clustering and dimension reduction are typical unsupervised tasks.
- Overfitting means the model fits training data too closely and performs poorly on new data.
- Underfitting means the model is too simple to capture the pattern, even in training data.
- Split data into training, validation and test sets; the test set checks performance on unseen data.
- Data cleaning deals with missing values, outliers, duplicates and inconsistent formats.
- Tokenization splits text into units; stop-word removal and stemming or lemmatization reduce noise.
- NLP turns text into structured form so it can be analysed, for example for sentiment.
- Pick the chart for the message: trends over time, comparisons, distributions or relationships.
- Always apply Ethics thinking: data privacy, bias and understanding a model before relying on it.
Common mistakes
- Confusing velocity with volume Fix: Volume is how much data exists. Velocity is how quickly it is created and must be processed. Look for words like real-time or streaming.
- Thinking big data means only structured data Fix: Variety includes unstructured sources such as text, images and audio. Alternative data is mostly unstructured.
- Calling JSON or XML data unstructured. Fix: If it has tags or key-value pairs that give organization, it is semi-structured.
- Mixing up normalization and standardization. Fix: Normalization uses min and max and gives 0 to 1. Standardization uses mean and standard deviation and gives mean 0 and standard deviation 1.
- Calling clustering a supervised method because it produces groups. Fix: Check for labels in advance. Clustering discovers groups without labels; classification predicts known labels.
- Saying a model with the lowest training error is the best model. Fix: Judge on out-of-sample error. A very low training error with high test error signals overfitting.
- Saying stemming always returns a valid word. Fix: Remember that stemming just trims endings by rule, so outputs like 'analyz' can occur. Lemmatization returns real words.
- Thinking bag-of-words keeps word order. Fix: A bag has no order. If order or phrases matter, use n-grams.
Exam tips
- Expect short scenario questions: the stem gives a clue, and you name the V or the tool. Practise the clue-to-V conversion until it is automatic.
- Watch for options that overstate technology, such as 'eliminates bias' or 'guarantees accuracy'. These are usually wrong.
- Know the difference between structured and unstructured data and give a quick example of each.
- With no penalty for wrong answers, never leave a question blank. Eliminate one option you can rule out and choose between the other two.
- Read for limitations in the stem. A hint of poor data quality, overfitting or opaque models usually shapes the right answer.
- Expect definition and classification questions. Memorize one example for each data type and each cleansing problem.
- Know the contrast: normalization is bounded 0 to 1 and outlier-sensitive; standardization is not bounded and handles outliers better.
- In a calculation, do the subtraction first and write it down. Mistakes usually come from wrong inputs, not from division.