Risk Modelling and Survival Analysis · Elementary principles of machine learning
Model Training, Validation and Overfitting in Machine Learning
Updated 11 October 2026 · Fact-checked
You fit a model on a training set, tune it on a validation set and judge it once on a test set. Overfitting means the model fits noise, so training error is low but new-data error is high. You limit it with cross-validation, simpler models and regularisation such as LASSO or ridge.
Understand Model Training, Validation and Overfitting
A machine learning model is built to predict new data, not to describe the data you already have. So you cannot judge it by how well it fits the data used to build it. A very flexible model can match every training point, noise included, and then predict badly. This is overfitting.
To measure real performance, you split the data. The training set is used to estimate the parameters. The validation set is used to compare models or choose tuning settings (such as the number of parameters or the penalty size). The test set is kept aside and used once at the end to give an honest estimate of error on unseen data. If you use the test set to make choices, it is no longer unbiased.
When data is scarce, k-fold cross-validation makes better use of it. You split the data into k parts (folds). You train on k − 1 folds and test on the remaining one. You repeat this k times so each fold is the validation fold once. You then average the k error values. A special case is leave-one-out, where k equals the number of observations.
The bias-variance trade-off explains why flexibility cuts both ways. Bias is the error from a model that is too simple to capture the true pattern (underfitting). Variance is how much the fitted model changes if you train on a different sample (overfitting). Making a model more flexible lowers bias but raises variance. Expected test error is roughly bias squared plus variance plus irreducible noise. You want the flexibility that minimises test error, not training error.
Regularisation controls flexibility by penalising large coefficients. Ridge adds a penalty on the sum of squared coefficients. It shrinks coefficients towards zero but does not normally set them exactly to zero. LASSO adds a penalty on the sum of absolute values of coefficients. It can set some coefficients exactly to zero, so it also selects variables. A larger penalty parameter λ means more shrinkage: higher bias, lower variance. You usually choose λ by cross-validation.
Key rules to remember
- Expected test error decomposition
- Expected squared prediction error = Bias² + Variance + Irreducible error
- For squared-error loss. Irreducible error is the noise variance and cannot be removed by any model.
- Ridge regression
- Minimise Σ (yᵢ − ŷᵢ)² + λ Σ βⱼ²
- The sum is over the slope coefficients. The intercept is normally not penalised. λ ≥ 0. When λ = 0 you get ordinary least squares.
- LASSO regression
- Minimise Σ (yᵢ − ŷᵢ)² + λ Σ |βⱼ|
- Can set coefficients exactly to zero, so it performs variable selection.
- k-fold cross-validation error
- CV = (1 ÷ k) Σ Eᵢ, for i = 1 to k
- Eᵢ is the error measured on fold i when the model is trained on the other k − 1 folds.
- Effect of λ
- λ ↑ ⇒ bias ↑, variance ↓
- Choose λ that minimises cross-validated error.
How to solve Model Training, Validation and Overfitting questions
Use this approach for any question on training, validation and overfitting.
- 1Identify what the question asks: describe a concept, diagnose a problem from given errors, or compute a cross-validation or penalised value.
- 2Name the data sets in use. State what each is for: training to fit, validation to tune or choose, test for final unbiased assessment.
- 3Compare training and validation errors. Low training error with much higher validation error signals overfitting. High error on both signals underfitting (high bias).
- 4Link the diagnosis to bias and variance. Overfitting means low bias, high variance. Underfitting means high bias, low variance.
- 5Choose the remedy and justify it: more data, a simpler model, fewer variables, cross-validation, or regularisation (ridge or LASSO).
- 6For calculations, write the formula first, substitute the numbers, and average over folds or evaluate the penalised objective carefully.
- 7State your conclusion in context, including any assumption such as the data being independent and representative.
Quickest way: Diagnose by comparing errors
When to use it: Use for multiple-choice questions and short written parts where training and validation errors, or model flexibility, are given.
- Training error much lower than validation error: overfitting, so reduce flexibility or add a penalty.
- Both errors high and close: underfitting, so add flexibility or variables.
- Need variable selection or sparse model: LASSO. Many correlated predictors to be shrunk: ridge.
- Larger λ means simpler model, more bias, less variance.
- Never use the test set to tune.
Common mistakes in Model Training, Validation and Overfitting
Choosing the model with the lowest training error.
Training error always falls as flexibility rises, so it looks like improvement.
Fix: Compare models on validation or cross-validated error, never on training error alone.
Using the test set repeatedly to tune the model.
Students treat validation and test sets as the same thing.
Fix: Tune on validation data or by cross-validation. Use the test set once at the end.
Saying ridge sets coefficients to zero.
Both methods shrink coefficients, so the difference is blurred.
Fix: Ridge shrinks towards zero but usually not exactly to zero. LASSO can give exact zeros because of its absolute-value penalty.
Mixing up bias and variance when λ increases.
Students remember that a penalty 'reduces overfitting' but not the direction of each effect.
Fix: Larger λ gives a simpler model: bias goes up and variance goes down.
Dividing by the wrong number in k-fold cross-validation.
Confusion between k folds and the size of the training set.
Fix: Compute the error for each of the k folds, then average over k. Check that each observation is validated exactly once.
Penalising the intercept or ignoring variable scaling.
The penalty is applied mechanically to every parameter.
Fix: State that predictors are usually standardised before penalising and that the intercept is normally left unpenalised.
Worked examples
Example 1
A model is fitted with polynomial degrees 1, 2, 3 and 4. Mean squared errors are: degree 1 training 9.0, validation 9.4; degree 2 training 5.1, validation 5.6; degree 3 training 2.0, validation 7.8; degree 4 training 0.8, validation 12.5. Which degree would you choose, and what do the results show?
Show the solution
- Compare models on validation error, not training error.
- Validation errors are 9.4, 5.6, 7.8 and 12.5. The lowest is 5.6, at degree 2.
- Degree 1 has high error on both sets, so it underfits (high bias).
- Degrees 3 and 4 have much lower training error than validation error, and validation error rises, so they overfit (high variance).
- Degree 2 balances bias and variance.
Answer: Choose degree 2 (validation MSE 5.6). Degree 1 underfits. Degrees 3 and 4 overfit. A final test set should then be used once to estimate performance.
Example 2
A model is assessed by 4-fold cross-validation. The mean squared errors on the four validation folds are 12, 15, 9 and 16. (a) Calculate the cross-validation error. (b) Explain why the test set should not be used to choose between models.
Show the solution
- (a) The formula is CV = (1 ÷ k) Σ Eᵢ with k = 4.
- Sum the errors: 12 + 15 + 9 + 16 = 52.
- Divide by 4: 52 ÷ 4 = 13.
- (b) If the test set guides model choice, the chosen model is partly tailored to it.
- Its test error then understates the true error on genuinely new data, so the estimate is optimistic.
Answer: (a) CV error = 13. (b) Using the test set for selection makes it part of the fitting process, so it no longer gives an unbiased estimate of performance on unseen data. Use validation data or cross-validation for selection.
Exam tips
- Always state which data set a quoted error comes from. Examiners award marks for separating training, validation and test errors.
- In written answers, name both bias and variance and say which way each moves as flexibility or λ changes.
- For LASSO versus ridge, give the penalty form and the exact-zero property. These are the standard comparison marks.
- Show the k-fold working: each fold used once for validation, then the average. Keep it short and exact.
- In the computer-based paper, say how you chose λ or the number of folds and why, not just the output.
Practice questions from Elementary principles of machine learning
- Which description best distinguishes a classification problem from a regression problem in supervised learning?
- When a validation set is used to choose between several candidate models and the chosen model is then assessed, why is a separate final test…
- A life insurer in Mumbai has data on 20,000 customers (age, premium paid, number of policies, channel) but no outcome variable. An analyst a…
- An actuary fits a flexible model to 200 motor claims and finds it has a very small error on those 200 records but a much larger error on 100…
- An analyst uses k-fold cross-validation with k = 5 on a dataset of 1,000 observations to compare models. Which description of the procedure …
Model Training, Validation and Overfitting in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Model Training, Validation and Overfitting: frequently asked questions
How do I avoid overfitting in machine learning?
Use a simpler model, fewer variables or more data. Validate with a separate set or cross-validation, and apply regularisation such as ridge or LASSO. Choose tuning settings by validation error, not training error.
What is the bias-variance trade-off in simple terms?
Simple models miss the pattern and have high bias but low variance. Flexible models follow the sample too closely and have low bias but high variance. The best model balances the two to minimise error on new data.
How does k-fold cross-validation work?
You split the data into k folds. You train on k − 1 folds, test on the remaining fold, and repeat so each fold is tested once. You then average the k errors.
What is the difference between LASSO and ridge regression?
Ridge penalises the sum of squared coefficients and shrinks them without usually reaching zero. LASSO penalises the sum of absolute coefficients and can set some exactly to zero, so it selects variables as well as shrinking them.