Skip to content

CFA Level II · CFA Level II Exam

Evaluating Regression Model Fit and Interpreting Model Results

This chapter teaches you to judge whether a multiple regression is useful and reliable. You check fit with R-squared and adjusted R-squared, test coefficients with t-tests, test the whole model with the F-test, then diagnose problems: heteroskedasticity, serial correlation, multicollinearity, outliers and misspecification. In the exam you read the output, pick the right test, and state the consequence.

What this chapter covers

This chapter sits in Quantitative Methods and builds on multiple regression. You already know how to read slope coefficients. Here you ask a harder question: can you trust the model? You learn to measure fit, test whether coefficients and groups of coefficients differ from zero, and spot the common ways a regression breaks.

The chapter has two halves. The first half is about numbers you read from a table: R-squared, adjusted R-squared, standard errors, t-statistics, the F-statistic and the ANOVA table. The second half is diagnosis. Each violation has a cause, a detection method, a consequence for the standard errors or coefficients, and a fix. Dummy variables and misspecification close the chapter by showing how model design itself can go wrong.

The ideas connect widely. Regression appears in Equities (factor models, valuation), Fixed Income (term structure and spread models), Economics (forecasting), Portfolio Construction (factor exposures) and Alternative Investments. At Level II every question sits in an item set, so a vignette may show regression output and ask you to interpret it, then use it elsewhere in the same set. You must find the right figure in the exhibit and apply the rule, not recall a definition.

Quantitative Methods is a smaller topic area by weight (5-10%), but regression output shows up as an exhibit in other topic areas too, so the skills carry beyond this chapter. The questions are also learnable. They follow a small set of patterns: compute a test statistic, compare it with a critical value, name a violation from a clue, and state its effect on standard errors and tests. If you practise these patterns, you can collect points quickly and save time for heavier item sets.

Evaluating Regression Model Fit and Interpreting Model Results: topics in the order to study them

  1. 1Goodness of Fit: R-squared and Adjusted R-squaredStart with how fit is measured, since it introduces the sums of squares and the ANOVA table used by every later topic.
  2. 2Hypothesis Testing of Regression CoefficientsNext, learn the t-test on a single coefficient. It uses the standard errors that the violations later distort.
  3. 3F-test and Joint Hypothesis TestsMove from one coefficient to several at once. It reuses the ANOVA table and the idea of a test statistic against a critical value.
  4. 4Violations: Heteroskedasticity and Serial CorrelationNow you know the tests, so you can see how errors that break assumptions make them unreliable, and how to detect and correct this.
  5. 5MulticollinearityThis is a third assumption problem, but with a different symptom: high R-squared and F with weak individual t-statistics.
  6. 6Influence Analysis and OutliersOnce you can diagnose assumption problems, learn to find single observations that drive the results.
  7. 7Dummy Variables and MisspecificationFinish with model design: how to code categories and which specification errors cause biased or inconsistent estimates.

How to prepare Evaluating Regression Model Fit and Interpreting Model Results

Treat this chapter as a toolkit. For each tool you need three things: when to use it, how to compute it, and what the result means. Work with real output tables, since the exam gives you exhibits.

  1. Read the ANOVA table until you can fill any missing cell: SST = RSS + SSE, R² = RSS ÷ SST, and the mean squares divide by their degrees of freedom.
  2. Memorise the core formulas and the degrees of freedom: t = (b̂ − hypothesised value) ÷ standard error, with n − k − 1 degrees; F = MSR ÷ MSE, with k and n − k − 1 degrees.
  3. Build a one-page table of the violations: cause, detection test, effect on coefficients and standard errors, and correction. Add multicollinearity and misspecification to it.
  4. Practise reading a vignette exhibit. Underline the coefficient, standard error, sample size and number of variables before you calculate anything.
  5. Do item-set practice with four questions per vignette. After each set, note whether you missed the data, the method or the interpretation.
  6. Redo missed questions after two days. Focus on the direction of effects, such as whether standard errors are too small or too large and whether Type I errors rise.
  7. In the last week, do timed sets mixing this chapter with others so you recognise regression clues quickly.

Common mistakes in Evaluating Regression Model Fit and Interpreting Model Results

  • Using a high R² as proof that the model is good.

    Fix: Compare adjusted R², check coefficient significance, and test for violations before you trust the fit.

  • Using the wrong degrees of freedom in the t-test or F-test.

    Fix: Write n, k and the degrees of freedom first, directly from the exhibit, before you look up a critical value.

  • Concluding from the F-test that every variable is significant.

    Fix: Remember it says at least one slope is non-zero. Use the individual t-tests to look at each coefficient.

  • Mixing up the effects of heteroskedasticity, serial correlation and multicollinearity.

    Fix: Learn each as cause, detection, effect and fix in one table. Remember that multicollinearity is about the independent variables, not the errors.

  • Getting the direction of the bias in standard errors wrong.

    Fix: Remember that positive serial correlation typically understates standard errors, which inflates t-statistics and raises Type I error, and that multicollinearity inflates them.

  • Misreading dummy variable coefficients.

    Fix: Use n − 1 dummies, and read each coefficient as the difference from the omitted category, holding other variables constant.

Last-day revision: Evaluating Regression Model Fit and Interpreting Model Results

  • R² = explained variation ÷ total variation; it never falls when you add a variable.
  • Adjusted R² penalises extra variables and can fall; it is always at most R².
  • t-statistic = (estimated coefficient − hypothesised value) ÷ standard error, with n − k − 1 degrees of freedom.
  • F = MSR ÷ MSE tests whether all slope coefficients are jointly zero; it is a one-tailed test with k and n − k − 1 degrees of freedom.
  • In a multiple regression, a significant F-test only tells you at least one slope is non-zero, not which one.
  • Heteroskedasticity: error variance changes with the independent variables; coefficients stay consistent, but standard errors are unreliable. Use the Breusch-Pagan test and robust (White) standard errors.
  • Serial correlation: errors are correlated across time; positive serial correlation usually makes standard errors too small and t-statistics too large. Use the Durbin-Watson test and Newey-West standard errors.
  • Multicollinearity: highly correlated independent variables inflate standard errors; the signal is a high R² and significant F with insignificant t-statistics. Check variance inflation factors.
  • Outliers and high-leverage points can drive results; leverage and Cook's distance help you find influential observations.
  • With a dummy variable for n categories, use n − 1 dummies; the omitted category is the benchmark.
  • Misspecification, such as an omitted variable or wrong functional form, can make coefficients biased and inconsistent.

Evaluating Regression Model Fit and Interpreting Model Results in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Evaluating Regression Model Fit and Interpreting Model Results: frequently asked questions

What is the difference between R-squared and adjusted R-squared?

R-squared is the share of variation in the dependent variable that the model explains, and it never falls when you add a variable. Adjusted R-squared adjusts for the number of variables, so it can fall if a new variable adds little. Use it to compare models with different numbers of independent variables.

When do I use the F-test instead of the t-test?

Use the t-test for a single coefficient. Use the F-test to test whether all slopes are jointly zero, or whether a group of coefficients is jointly zero. The joint test matters because individual t-tests can mislead when variables are related.

How do I tell heteroskedasticity from serial correlation in a vignette?

Heteroskedasticity means the error variance changes with the level of the independent variables, often in cross-sectional data. Serial correlation means errors are related over time, often in time-series data. The vignette usually names the test: Breusch-Pagan points to heteroskedasticity, and Durbin-Watson to serial correlation.

What is the sign of multicollinearity in regression output?

The classic sign is a high R² and a significant F-statistic together with insignificant individual t-statistics. Variance inflation factors confirm it. Coefficient estimates remain unbiased, but their standard errors are inflated.

Do I need to memorise critical values for this chapter?

Usually the vignette provides the critical values or the table. Your task is to compute the test statistic correctly, use the right degrees of freedom and compare them properly. Practise reading the exhibit rather than memorising tables.