CFA Level II Exam · Extensions of Multiple Regression
Influential Observations in Regression: Outliers and Leverage
Updated 7 October 2026 · Fact-checked
An influential observation is a data point that changes the regression results when removed. A high-leverage point has an unusual X value. An outlier has an unusual Y given X. You check leverage, studentized residuals and Cook's D against cutoffs, then investigate any flagged points.
Understand Influence Analysis: Outliers and Leverage
A regression line is fitted to all observations together. Most points pull on the line a little. A few points can pull on it a lot. Influence analysis finds those points and asks how much they change your estimates.
Two different things can make a point unusual. A high-leverage point has an extreme value on one or more independent variables (X). It sits far from the other X values, so it has the potential to pull the fitted line toward itself. An outlier has an extreme value of the dependent variable (Y) relative to what the model predicts. It sits far from the fitted line, so it has a large residual.
A point can be one, both or neither. A high-leverage point that lies close to the line the other points suggest may do no harm. An outlier with ordinary X values usually has limited effect on the slope. The danger is a point that is both high leverage and an outlier. That point is likely to be influential, meaning that deleting it would materially change the estimated coefficients.
Three tools measure this. Leverage (h) is a number between 0 and 1 for each observation that measures how far its X values are from the average X values. Studentized residuals rescale residuals so you can judge how large they are. Cook's distance (Cook's D) combines both ideas into one measure of how much the fitted values change if the observation is dropped.
Flagging a point is not a reason to delete it. You should find out why it is unusual. It may be a data entry error, a genuine rare event, or a sign that the model is misspecified. Delete only when you have a sound reason, such as a clear error.
Key formulas to remember
- Leverage rule of thumb
- Potentially high leverage if hᵢᵢ > 3 × (k + 1) ÷ n
- k is the number of independent variables and n is the number of observations. (k + 1) ÷ n is the average leverage. Leverage values lie between 0 and 1.
- Studentized residual
- tᵢ* = eᵢ ÷ (standard error of the residual estimated with observation i deleted)
- Compare |tᵢ*| with the critical t-value from a t-distribution with n − k − 2 degrees of freedom. A value beyond the critical value flags an outlier. A cutoff of |tᵢ*| > 3 is also used as a common rule of thumb.
- Cook's D rules of thumb
- Dᵢ > 0.5: possibly influential. Dᵢ > 1: very likely influential. Another rule: Dᵢ > √(k ÷ n)
- Values above 0.5 (possibly influential) and above 1 (very likely influential) are the main guides. The curriculum also gives √(k ÷ n) as another cutoff, where k is the number of independent variables and n is the number of observations. It is not the standard for every sample size. Cook's D measures the overall change in fitted values when observation i is dropped. Use the cutoff the question supplies.
- Classification rule
- High leverage = unusual X; outlier = unusual Y given X; influential = changes the results when removed
- The cutoffs are guides, not hard laws. Always say 'flag for investigation'.
How to solve Influence Analysis: Outliers and Leverage questions
Use this order for any influence-analysis question in a vignette.
- 1Find n and k in the vignette. k is the number of independent variables, not counting the intercept.
- 2Read the statistic given for each observation: leverage, studentized residual or Cook's D.
- 3Match each statistic to its cutoff. Compute 3 × (k + 1) ÷ n for leverage. Use the critical t-value or the cutoff given for studentized residuals. Use √(k ÷ n) for Cook's D unless the question directs you to 0.5 or 1.
- 4Classify each observation: high leverage (leverage above cutoff), outlier (studentized residual beyond cutoff), influential (Cook's D above cutoff).
- 5Check for the combination. An observation that fails more than one test is the strongest candidate to be influential.
- 6State the action: investigate the data point, correct errors, and compare the model with and without it. Do not recommend automatic deletion.
Quickest way: Three numbers, three cutoffs
When to use it: Use this when an exhibit lists leverage, studentized residual and Cook's D for several observations and you have about two minutes.
- Compute the leverage cutoff once: 3 × (k + 1) ÷ n. Compute the Cook's D cutoff once: √(k ÷ n), unless the question gives one.
- Scan each column for values above its cutoff and note which observation numbers fail.
- Leverage fail alone means high leverage. Residual fail alone means outlier. Cook's D above the cutoff flags the point as likely influential; investigate it and compare results with and without it.
- Pick the answer that matches the label for the flagged observation. Beware options that call a high-leverage point an outlier.
Common mistakes in Influence Analysis: Outliers and Leverage
Treating high leverage and outlier as the same thing.
Both describe an unusual point, so the words blur together.
Fix: Leverage is about X values. An outlier is about Y given X, shown by a large residual.
Using k instead of k + 1 in the leverage cutoff.
Students forget the intercept counts as a parameter.
Fix: Use 3 × (k + 1) ÷ n. For 2 independent variables, use 3 × 3 ÷ n. Note that the Cook's D cutoff √(k ÷ n) uses k.
Deleting every flagged observation.
A flag feels like proof of a problem.
Fix: A flag means investigate. Remove a point only if it is an error or does not belong to the population studied.
Assuming a high-leverage point is always influential.
Leverage measures potential to pull, not actual pull.
Fix: Check Cook's D or the change in coefficients. A high-leverage point close to the fitted line may have little effect.
Judging an outlier by the ordinary residual size alone.
Raw residuals depend on scale and on the point's leverage.
Fix: Use the studentized residual and compare it with the critical t-value.
Worked examples
Example 1
An analyst regresses a stock's return on 3 independent variables using 40 observations. Observation 12 has leverage 0.31, a studentized residual of 1.1 and Cook's D of 0.20. Use √(k ÷ n) as the Cook's D cutoff. (1) What is the leverage cutoff? (2) How should observation 12 be classified?
Show the solution
- k = 3 and n = 40, so the cutoff is 3 × (3 + 1) ÷ 40 = 12 ÷ 40 = 0.30.
- Leverage 0.31 is above 0.30, so observation 12 is a high-leverage point.
- The studentized residual of 1.1 is small, so it is not an outlier.
- The question specifies the Cook's D cutoff √(k ÷ n) = √(3 ÷ 40) = √0.075 ≈ 0.274. Cook's D of 0.20 is below it, so the point is not likely influential. The 0.5 guide gives the same conclusion, since 0.20 is below 0.5.
Answer: The leverage cutoff is 0.30. Observation 12 is a high-leverage point but not an outlier and not likely influential. Investigate it but do not delete it automatically.
Example 2
A regression has 2 independent variables and 60 observations. Observation 7 has leverage 0.04, a studentized residual of 3.4 and Cook's D of 0.15. Observation 31 has leverage 0.20, a studentized residual of 3.2 and Cook's D of 0.90. Use a leverage cutoff of 3 × (k + 1) ÷ n, a studentized residual cutoff of 3 and a Cook's D cutoff of √(k ÷ n). (1) Which observation is a high-leverage point? (2) Which is the stronger candidate for being influential?
Show the solution
- Leverage cutoff = 3 × (2 + 1) ÷ 60 = 9 ÷ 60 = 0.15.
- Cook's D cutoff = √(2 ÷ 60) = √0.0333 ≈ 0.183.
- Observation 7: leverage 0.04 is below 0.15, so not high leverage. Residual 3.4 is above 3, so it is an outlier. Cook's D 0.15 is below 0.183, so not likely influential.
- Observation 31: leverage 0.20 is above 0.15, so high leverage. Residual 3.2 is above 3, so an outlier. Cook's D 0.90 is above 0.183, so likely influential.
- Observation 31 fails all three tests, so it is the stronger candidate.
Answer: Observation 31 is the high-leverage point. It is also an outlier with Cook's D above the cutoff, so it is the stronger candidate for being influential. Observation 7 is only an outlier.
Exam tips
- Write the leverage cutoff first. It is the one number you must calculate yourself.
- Read option wording closely. Examiners swap 'outlier' and 'high leverage' to build wrong answers.
- If the vignette gives its own cutoff, use it instead of a rule of thumb.
- When asked what the analyst should do, choose investigate and compare results with and without the point, not delete.
- Remember k excludes the intercept but the leverage cutoff uses k + 1.
Influence Analysis: Outliers and Leverage: frequently asked questions
What is leverage in multiple regression?
Leverage measures how far an observation's independent variable values are from their averages. It ranges from 0 to 1. A high value means the point has more potential to pull the fitted line.
What is the difference between an outlier and a high-leverage point?
An outlier has an unusual dependent variable value given its X values, shown by a large residual. A high-leverage point has unusual X values. A point can be either, both or neither.
What is Cook's distance used for?
Cook's D measures how much the fitted regression values change when one observation is removed. It combines leverage and residual size. A value above √(k ÷ n) is taken as a sign the point is likely influential. Values above 0.5 or 1 are other common guides.
Should I delete influential observations?
Not automatically. First check for data errors and ask whether the point belongs to the population you study. Delete only with a sound reason, and compare results with and without it.