CFA Level II Exam · Evaluating Regression Model Fit and Interpreting Model Results
Influence Analysis: Cook's Distance, Leverage and Outliers
Updated 7 October 2026 · Fact-checked
Influence analysis checks whether a few observations distort a regression. Leverage (h) flags unusual X values. Studentized residuals flag unusual Y values. Cook's distance combines both and measures how much the fitted model changes if one observation is removed. Large values mark influential points that need investigation, not automatic deletion.
Understand Influence Analysis and Outliers
An outlier is an observation that does not fit the pattern of the rest of the data. But not every outlier changes your regression much. The key question is influence: if you delete this point, do the coefficients move materially?
There are two ways a point can be unusual. A high-leverage point has an extreme value on an independent variable (X). It sits far from the other X values, so it has the potential to pull the fitted line toward itself. An outlier in Y has a dependent value far from what the model predicts, so its residual is large.
Leverage is measured by h (the hat value, from the diagonal of the hat matrix). h lies between 0 and 1. The average h across all observations is (k + 1) ÷ n, where k is the number of independent variables and n the number of observations. A common rule of thumb flags h above 3 × (k + 1) ÷ n.
To spot Y outliers, use the studentized residual. You delete one observation, refit the model, and compare the actual Y for that observation with the prediction from the refit model, scaled by its standard error. A common rule of thumb flags an absolute value above 3. You can also compare it with the critical t-value with n − k − 2 degrees of freedom.
Cook's distance (D) puts both pieces together. It measures the combined effect of an observation's leverage and its residual on the fitted values. A point with high leverage but a small residual, or a big residual but low leverage, may have modest influence. A point with both is the dangerous one. The curriculum's common flag is D above √(k ÷ n). Other conventions flag D above 0.5 (worth a look) or D above 1 (likely influential). Once you find an influential point, check for data errors first. Then ask whether it is a genuine part of the population. Do not delete it just because it hurts your fit.
Key formulas to remember
- Average leverage
- average h = (k + 1) ÷ n
- k = number of independent variables, n = number of observations. The hat values sum to k + 1.
- Leverage rule of thumb
- Potentially influential if h > 3 × (k + 1) ÷ n
- Flags unusual X values. It is a screening rule, not proof of influence.
- Studentized residual
- t_i* = e_i* ÷ s_e*, where e_i* = Y_i − Ŷ_i(i) is the residual of observation i computed from the model fitted without it, and s_e* is its standard error
- Flags Y outliers. Rule of thumb: |t*| > 3, or compare with the critical t-value with n − k − 2 degrees of freedom.
- Cook's distance
- D_i = Σ(Ŷ_j − Ŷ_j(i))² ÷ ((k + 1) × MSE), summed over all observations j, where Ŷ_j(i) is the fitted value for j from the model refit without observation i
- Combines residual size and leverage. Larger D means greater influence on the fitted values. An equivalent form is D_i = (e_i² ÷ ((k + 1) × MSE)) × (h_i ÷ (1 − h_i)²), where e_i is the ordinary residual from the full model.
- Cook's distance rules of thumb
- D > √(k ÷ n) is the curriculum's common flag; other conventions: D > 0.5 worth a look, D > 1 likely influential
- These are conventions. The exam will tell you which threshold to use or give a clear gap.
How to solve Influence Analysis and Outliers questions
Use this order for any influence-analysis item. It tells you which statistic answers which question.
- 1Read the question stem. Decide whether it asks about unusual X values (leverage), unusual Y values (studentized residual) or overall influence (Cook's distance).
- 2Find k and n in the vignette. Count independent variables only for k.
- 3For leverage, compute the threshold 3 × (k + 1) ÷ n and compare each observation's h with it.
- 4For studentized residuals, compare the absolute value with 3 or with the critical t-value the vignette gives.
- 5For Cook's distance, compare D with the stated threshold, such as 0.5, 1 or √(k ÷ n).
- 6Combine the results. High h plus large residual plus large D means an influential observation. High h alone only means potential influence.
- 7Choose the action. Verify the data for errors, then consider whether the point is a valid observation. Do not simply delete it.
Quickest way: Three-flag screen
When to use it: Use when an exhibit lists h, studentized residual and Cook's D for several observations and you must pick the influential one.
- Compute the leverage cutoff once: 3 × (k + 1) ÷ n.
- Scan the Cook's D column first. If one value clearly exceeds the threshold, it is usually your answer.
- Check that observation's h and residual to explain why it is influential.
- Eliminate options that call an observation influential on leverage alone, or that recommend automatic deletion.
Common mistakes in Influence Analysis and Outliers
Treating every outlier as influential.
An outlier looks dramatic on a scatter plot, so it seems to matter.
Fix: Influence needs the point to actually move the fitted model. A Y outlier at a central X value has low leverage and often small influence. Check Cook's D.
Treating high leverage as proof of a problem.
Students read h above the cutoff as a bad point.
Fix: High leverage is potential influence. If the point lies close to the fitted line, its residual is small and it may not distort the model.
Using n instead of k + 1 in the leverage cutoff, or counting the intercept in k.
The formula uses k + 1 and k in different places.
Fix: k is the number of slope variables. The leverage average and cutoff use k + 1. Write k and n down before computing.
Deleting influential observations automatically.
Removing the point improves R-squared, so it seems like the fix.
Fix: First check for data errors. If the point is valid, it carries real information. Deleting it to improve fit is poor practice.
Mixing up studentized residual with the ordinary residual.
Both are residuals, and the names sound alike.
Fix: The studentized residual is scaled, and it uses a model fitted without that observation. That is why it compares to t-values and a rule of 3.
Worked examples
Example 1
An analyst regresses monthly fund returns on 3 independent variables using 60 observations. Observation 17 has h = 0.31. Observation 42 has h = 0.12. (1) What is the leverage cutoff? (2) Which observations exceed it?
Show the solution
- k = 3 and n = 60, so k + 1 = 4.
- Cutoff = 3 × 4 ÷ 60 = 12 ÷ 60 = 0.20.
- Observation 17: 0.31 > 0.20, so it is a high-leverage point.
- Observation 42: 0.12 < 0.20, so it is not.
Answer: The cutoff is 0.20. Only observation 17 is a high-leverage point, which means it has potential to be influential, not that it is.
Example 2
Using the same regression (k = 3, n = 60), observation 17 has a studentized residual of 0.8 and Cook's D of 0.15. Observation 5 has h = 0.08, a studentized residual of 4.1 and Cook's D of 1.4. (1) Is observation 17 influential using D > 1? (2) What does observation 5 show? (3) What is the best next step?
Show the solution
- Observation 17: D = 0.15 is below 0.5 and below 1. Its studentized residual 0.8 is below 3. Despite high leverage it fits the model well, so it is not influential.
- Observation 5: |t*| = 4.1 > 3 marks a Y outlier. D = 1.4 > 1 marks high influence. Its h = 0.08 is below the 0.20 cutoff, so its influence comes from the very large residual.
- Next step: check observation 5 for a data error, such as a mistyped return. If it is valid, assess it before deciding how to treat it. Do not delete it automatically.
Answer: Observation 17 is not influential. Observation 5 is influential because of a large residual with low leverage. Investigate the data before deciding on any change.
Exam tips
- Know which statistic maps to which question: leverage for X, studentized residual for Y, Cook's D for overall influence.
- Compute k + 1 carefully. Most numerical errors come from using k alone in the leverage cutoff.
- When an option says to delete the observation automatically, it is almost always wrong. The answer is usually to investigate first.
- Expect an exhibit with several observations. Scan the Cook's D column first, then check h and the residual to justify.
- If the vignette gives a threshold, use it. Do not substitute your own rule of thumb.
Influence Analysis and Outliers in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Influence Analysis and Outliers: frequently asked questions
What is Cook's distance in CFA Level II?
Cook's distance measures how much the fitted regression values change if one observation is removed. It combines the observation's residual and its leverage. Larger values mean the point is more influential.
What is a high-leverage point in regression?
It is an observation with an extreme value on one or more independent variables. It has the potential to pull the regression line toward it. It is flagged when h exceeds 3 × (k + 1) ÷ n.
How is a studentized residual different from a normal residual?
A studentized residual divides the residual by a standard error computed from a model fitted without that observation. This makes it comparable with t-values. A common flag is an absolute value above 3.
Should I remove influential observations from the data?
Not automatically. First check whether the data point is an error. If it is valid, it is real information, and removing it only to improve fit is poor practice.