CFA Level I Exam · Introduction to Financial Data Science
Machine Learning Approaches and Model Fit for CFA Level 1
Updated 7 October 2026 · Fact-checked
Machine learning finds patterns in data without being explicitly programmed. Supervised learning uses labeled data to predict a target. Unsupervised learning finds structure in unlabeled data. Deep learning uses multi-layer neural networks. Overfitting means a model fits training noise and fails on new data. Underfitting means it is too simple. Validation data and cross-validation check fit.
Understand Machine Learning Approaches and Model Fit
Machine learning (ML) lets a computer learn patterns from data instead of following fixed rules written by a person. In finance, you might use it to predict defaults, group similar stocks, or read text in earnings calls.
The first split is by whether the data has answers. In supervised learning, each observation has inputs (features) and a known output (the target or label). The model learns the link between them. If the target is a number, such as next month's return, it is regression. If the target is a category, such as default or no default, it is classification. Examples include penalized regression, support vector machines, k-nearest neighbor, classification and regression trees (CART), and random forests.
In unsupervised learning, there is no target. The model looks for structure on its own. Two main tasks: dimension reduction (for example principal components analysis, PCA, which compresses many correlated features into a few) and clustering (for example k-means and hierarchical clustering, which group similar observations). Deep learning uses neural networks with many hidden layers. It can be used for supervised or unsupervised tasks, and it suits complex data such as images and text. Reinforcement learning is a third type, where an agent learns by trial and reward.
Now model fit. A good model must work on new data, not just the data it learned from. Overfitting happens when the model is too complex and learns noise in the training set. It shows low training error but high error on new data. Underfitting happens when the model is too simple to capture the real pattern. It shows high error on both.
To check fit, split the data into a training set (to fit the model), a validation set (to tune it) and a test set (to judge final performance). K-fold cross-validation repeats this idea: split the data into k equal parts, train on k − 1 parts, validate on the remaining one, and rotate so every part is used once for validation. Then average the results. Overfitting is also reduced by penalizing complexity (regularization), using simpler models, and gathering more data.
Key formulas to remember
- Bias-variance idea
- Total error = bias error + variance error + base (irreducible) error
- High bias error means underfitting. High variance error means overfitting. Complexity trades one against the other.
- K-fold cross-validation
- Each fold size ≈ N ÷ k; k training rounds; each fold used once for validation
- Training uses k − 1 folds each round. Average the k validation results.
- Fit pattern: overfit
- Low training error, high out-of-sample error
- The model memorized noise and does not generalize.
- Fit pattern: underfit
- High training error, high out-of-sample error
- The model is too simple. Add features or complexity.
- Fit pattern: good fit
- Low training error, similar low out-of-sample error
- The model generalizes well.
- Classification metrics
- Precision = TP ÷ (TP + FP); Recall = TP ÷ (TP + FN); Accuracy = (TP + TN) ÷ (TP + FP + TN + FN); F1 = 2 × P × R ÷ (P + R)
- Use precision when false positives are costly and recall when false negatives are costly. F1 balances both.
How to solve Machine Learning Approaches and Model Fit questions
Use this order for any question on ML approaches or model fit.
- 1Ask: does the data have a labeled target? Yes means supervised. No means unsupervised.
- 2If supervised, decide the target type: a continuous number means regression, a category means classification.
- 3If unsupervised, decide the goal: grouping observations means clustering; reducing the number of features means dimension reduction.
- 4If the data is images, text or very complex patterns, think deep learning or neural networks.
- 5For fit questions, compare training error with out-of-sample error. Low then high means overfit. High on both means underfit.
- 6Match the fix to the problem: overfitting needs simpler models, regularization, more data or cross-validation; underfitting needs more complexity or features.
- 7Check the data split: training fits, validation tunes, test gives the final unbiased estimate.
- 8Eliminate the two options that contradict the earlier steps and choose the remaining one.
Quickest way: Label test and error-gap test
When to use it: Use for most three-option conceptual questions where you have about 90 seconds.
- Find the word 'labeled', 'target' or 'known outcome'. Present means supervised; absent means unsupervised.
- Look for the error gap. Big gap between training and test performance means overfitting.
- If both errors are high, choose underfitting.
- If the question asks for the cure, pick the option that reduces complexity for overfitting or adds it for underfitting.
Common mistakes in Machine Learning Approaches and Model Fit
Calling clustering a supervised method because it produces groups.
Groups look like classes, which classification also produces.
Fix: Check for labels in advance. Clustering discovers groups without labels; classification predicts known labels.
Saying a model with the lowest training error is the best model.
Low error feels like good performance.
Fix: Judge on out-of-sample error. A very low training error with high test error signals overfitting.
Treating overfitting and underfitting as the same problem.
Both give poor predictions.
Fix: Overfit means complex model, large train-test gap. Underfit means simple model, high error everywhere.
Using the test set repeatedly to tune the model.
It seems a convenient source of feedback.
Fix: Tune on validation data or with cross-validation. Keep the test set untouched for the final check.
Confusing precision and recall.
Both involve true positives and the names sound alike.
Fix: Precision divides by all predicted positives (TP + FP). Recall divides by all actual positives (TP + FN).
Thinking deep learning is a separate category from supervised and unsupervised learning.
It is often listed alongside them.
Fix: Deep learning is a type of model (multi-layer neural networks) that can be trained in supervised or unsupervised settings.
Worked examples
Example 1
An analyst trains a model on 10 years of data. It has a 1% error on the training set but a 22% error on new data. Which description is most accurate? A. Underfitting, because the model is too simple. B. Overfitting, because the model learned noise in the training data. C. A good fit, because the training error is low.
Show the solution
- Compare errors: training 1%, new data 22%. The gap is large.
- A large gap with very low training error means the model memorized the training set.
- Underfitting would show high error on both sets, so A is out.
- A good fit needs low error on new data too, so C is out.
Answer: B
Example 2
A portfolio manager has 5,000 stocks described by 60 correlated financial ratios and no predefined labels. She wants to group similar stocks. Which approach fits best? A. Supervised classification. B. Unsupervised clustering. C. Supervised regression.
Show the solution
- Check labels: there are none, so the problem is unsupervised.
- The goal is grouping similar observations, which is clustering.
- Classification and regression both require a labeled target, so A and C are eliminated.
Answer: B
Exam tips
- The first question to ask is always whether labels exist. It settles most supervised versus unsupervised items.
- Items focus on concepts and matching: which method suits which task, and which fit problem matches which error pattern. Few require calculation.
- Remember that no option will say 'all of the above', so find the single best match and eliminate the two clear mismatches.
- For k-fold cross-validation, remember each data point is used for validation exactly once and for training k − 1 times.
- With no penalty for wrong answers, always answer. If unsure, eliminate the option that confuses training error with out-of-sample error.
Practice questions from Introduction to Financial Data Science
- An analyst prepares text data from company filings for a model and removes common words such as 'the', 'and', and 'of' before counting word …
- A data scientist converts a column of text analyst comments into numerical features that a machine learning model can use. Which step in tex…
- An analyst wants to show how the frequency of individual words varies across a large set of central bank statements, so readers can quickly …
- A researcher wants to show the distribution, median, quartiles and outliers of monthly returns for several funds side by side. The visualiza…
- A model predicting loan defaults achieves very low error on the training data but performs poorly on new data. This outcome is best describe…
Machine Learning Approaches and Model Fit in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Machine Learning Approaches and Model Fit: frequently asked questions
What is the difference between supervised and unsupervised learning?
Supervised learning trains on data with known outputs and predicts a target, either a number or a category. Unsupervised learning has no target and finds structure such as clusters or reduced dimensions. The presence of labels is the test.
How do I tell overfitting from underfitting?
Compare training and out-of-sample error. Overfitting shows low training error but much higher error on new data. Underfitting shows high error on both.
How does k-fold cross-validation help avoid overfitting?
It splits the data into k parts and rotates which part is held out for validation. Every observation is used for both training and validation across rounds. The averaged result gives a more reliable estimate of out-of-sample performance than a single split.
Is deep learning supervised or unsupervised?
It can be either. Deep learning means neural networks with many hidden layers, and these can be trained with labeled data or without it.