FRM Exam Part I · Regression Diagnostics
Outliers, Residual Plots and Regression Diagnostics for FRM
Updated 11 October 2026 · Fact-checked
Regression diagnostics check whether a fitted model is trustworthy. Plot residuals against fitted values to spot patterns, compute leverage to find extreme x-values, and use Cook's distance to measure how much one point moves the fit. An outlier has a large residual; an influential point changes the estimates.
Understand Outliers, Residual Plots and Model Diagnostics
A regression line is only a summary. Before you trust the coefficients, you check what the residuals say. The residual is the actual y minus the fitted y. If the model is well specified, residuals look like random noise with no pattern and roughly constant spread.
A residual plot graphs residuals against fitted values (or against an explanatory variable, or against time). A random horizontal band around zero is good. A curve or U-shape suggests a missing nonlinear term or an omitted variable. A funnel shape, with spread growing or shrinking, suggests heteroskedasticity. Runs of positive then negative residuals in time order suggest serial correlation.
Three ideas are easy to confuse. An outlier is a point with an unusually large residual: its y is far from the line. Leverage measures how extreme a point's x-value is relative to the other x-values. A high-leverage point has the potential to pull the line. An influential point actually changes the coefficients when you remove it. A point can be an outlier with low leverage (little influence), high leverage with a small residual (it sits on the line), or both (highly influential).
Cook's distance combines residual size and leverage into one measure of influence: how much the fitted values change when observation i is dropped. A common rule of thumb flags values above 1, or above 4 ÷ n for a stricter screen. These are guides, not laws. Another approach is to compare the coefficients with and without the point.
When you find an influential point, do not just delete it. Check for data errors first. If the point is genuine, such as a crisis-period return, removing it may hide real risk. Report results with and without it, consider robust methods, or revisit the model specification.
Key formulas to remember
- Residual
- eᵢ = Yᵢ − Ŷᵢ
- Actual minus fitted. OLS residuals sum to zero when the model has an intercept.
- Standardized residual (approximate)
- eᵢ ÷ s, where s = √(SSR ÷ (n − k − 1))
- s is the standard error of the regression. Absolute values above about 2 to 3 are often flagged as outliers. Exact versions also adjust for leverage.
- Leverage average
- Average hᵢᵢ = (k + 1) ÷ n
- k is the number of explanatory variables. Leverages sum to k + 1. Values well above the average, such as more than twice it, are often flagged.
- Leverage in simple regression
- hᵢᵢ = 1/n + (Xᵢ − X̄)² ÷ Σ(Xⱼ − X̄)²
- Grows as Xᵢ moves away from the mean of X.
- Cook's distance
- Dᵢ = (eᵢ² ÷ ((k + 1) × s²)) × (hᵢᵢ ÷ (1 − hᵢᵢ)²)
- Large when both the residual and the leverage are large. Common flags: Dᵢ > 1, or Dᵢ > 4 ÷ n.
- Residual-plot reading rule
- Pattern in mean → misspecification; pattern in spread → heteroskedasticity; pattern in time order → serial correlation
- Use this to link a picture to the problem.
How to solve Outliers, Residual Plots and Model Diagnostics questions
Use this method for any question on outliers, leverage, influence or residual plots.
- 1Identify what the question gives you: a residual plot description, residual values, leverage values, Cook's distance values, or coefficients with and without a point.
- 2Classify the pattern or point. Curve in residuals: nonlinearity or omitted variable. Funnel: heteroskedasticity. Runs over time: serial correlation. Large residual: outlier. Extreme x: high leverage.
- 3If numbers are given, compute what is needed: standardized residual as e ÷ s, average leverage as (k + 1) ÷ n, or Cook's distance from its formula.
- 4Compare with the rule-of-thumb threshold, such as Cook's D above 1 or 4 ÷ n, and treat the threshold as a flag, not proof.
- 5Decide whether the point is influential by asking whether removing it would change the coefficients materially. Both a large residual and high leverage are needed for strong influence.
- 6State the consequence: biased or unstable coefficients, invalid standard errors, or unreliable inference.
- 7Choose the remedy: check the data, add a missing variable or transformation, use robust standard errors or robust regression, or report results with and without the point. Avoid deleting data without justification.
Quickest way: Two-question screen for outliers and influence
When to use it: Use when an MCQ asks which point or plot indicates a problem and you have under two minutes.
- Ask: is the pattern in the residual mean, spread, or order? Mean means misspecification, spread means heteroskedasticity, order means serial correlation.
- For a single point, ask: is it far in y (outlier), far in x (leverage), or both (influential)?
- Eliminate options that call a high-leverage point automatically influential, or an outlier automatically influential.
- If Cook's D is quoted, compare it with 1 or 4 ÷ n and pick the larger D as most influential.
- Prefer answers that investigate before deleting.
Common mistakes in Outliers, Residual Plots and Model Diagnostics
Treating every outlier as influential.
A large residual looks dramatic on a plot.
Fix: An outlier near the mean of x has low leverage and may barely move the line. Influence needs a large residual combined with leverage.
Assuming high leverage means a bad point.
Leverage sounds like a defect.
Fix: A high-leverage point that lies close to the fitted line has a small residual and low Cook's distance. It is only potential influence.
Reading a funnel-shaped residual plot as nonlinearity.
Both are visible patterns.
Fix: A curve in the average residual signals nonlinearity or an omitted variable. A changing spread signals heteroskedasticity.
Deleting any point flagged by Cook's distance.
Students want a clean fit.
Fix: Check for data errors first. Real extreme events carry risk information. Report results with and without the point or use robust methods.
Treating thresholds like 1 or 4 ÷ n as exact tests.
Rules of thumb are memorized as rules.
Fix: They are screening guides. Look at the actual change in coefficients and the context.
Using the wrong average leverage by forgetting the intercept.
Students use k ÷ n.
Fix: The average leverage is (k + 1) ÷ n, because the intercept counts as a parameter.
Worked examples
Example 1
A regression with one explanatory variable (k = 1) uses n = 20 observations. Observation 7 has leverage 0.40. Is its leverage high compared with the average, using twice the average as a flag?
Show the solution
- Average leverage = (k + 1) ÷ n = 2 ÷ 20 = 0.10.
- Flag level = 2 × 0.10 = 0.20.
- Observation 7 has 0.40, which is above 0.20 (four times the average).
Answer: Yes. Its leverage of 0.40 is four times the average of 0.10 and exceeds the 0.20 flag, so it is a high-leverage point.
Example 2
In a regression with k = 1, n = 25, standard error of regression s = 2.0. Observation 5 has residual e = 6.0 and leverage h = 0.20. Compute Cook's distance and say whether it exceeds the 4 ÷ n and 1 thresholds.
Show the solution
- Formula: D = (e² ÷ ((k + 1) × s²)) × (h ÷ (1 − h)²).
- e² = 36. (k + 1) × s² = 2 × 4 = 8. First term = 36 ÷ 8 = 4.5.
- h ÷ (1 − h)² = 0.20 ÷ (0.80)² = 0.20 ÷ 0.64 = 0.3125.
- D = 4.5 × 0.3125 = 1.40625.
- 4 ÷ n = 4 ÷ 25 = 0.16. D = 1.41 is above 0.16 and above 1.
Answer: Cook's distance is about 1.41. It exceeds both 0.16 and 1, so observation 5 is influential and needs checking.
Exam tips
- Match each picture to a problem: curve to misspecification, funnel to heteroskedasticity, runs over time to serial correlation.
- Remember the logic: outlier is about y, leverage is about x, influence needs the effect on the fit.
- Know the average leverage (k + 1) ÷ n and the common Cook's D flags, and treat them as guides.
- Expect conceptual answers that say investigate before removing, and report with and without the point.
- If Cook's D is computed, do the arithmetic in steps: residual term first, then the leverage term, then multiply.
Practice questions from Regression Diagnostics
- In a cross-sectional regression of household food spending on household income, the dispersion of the residuals grows steadily as income ris…
- A risk manager estimates a time-series regression by OLS and finds the residuals are positively autocorrelated, while the regressors are not…
- After fitting a multiple regression, an analyst finds that a plot of residuals against a regressor shows a systematic U-shaped pattern. Whic…
- In a multiple regression of a fund's excess return on several factor returns, an analyst finds that the R-squared is high and the F-test is …
- A risk analyst is concerned about severe multicollinearity between two regressors, X1 and X2, in a model used to forecast portfolio losses. …
Outliers, Residual Plots and Model Diagnostics in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Outliers, Residual Plots and Model Diagnostics: frequently asked questions
What is the difference between an outlier and an influential observation?
An outlier has an unusually large residual. An influential observation materially changes the fitted coefficients when removed. An outlier with low leverage is often not influential.
What does Cook's distance measure?
It measures how much the fitted values change when one observation is dropped, combining its residual size and its leverage. Larger values mean more influence. Common screens are above 1 or above 4 ÷ n.
How do I read a residual plot?
Look for randomness around zero with constant spread. A curve means a missing nonlinear term or variable, a funnel means heteroskedasticity, and runs of same-sign residuals over time mean serial correlation.
Should I delete an influential outlier?
Not automatically. Check for data errors first. If the point is genuine, removal can hide real risk, so report results with and without it or use robust methods.