Skip to content

CFA Level II Exam · Evaluating Regression Model Fit and Interpreting Model Results

R-squared and Adjusted R-squared: Goodness of Fit in Regression

Updated 7 October 2026 · Fact-checked

Goodness of fit measures how much of the variation in the dependent variable a regression explains. R-squared = explained variation ÷ total variation, but it never falls when you add variables. Adjusted R-squared penalizes extra variables. AIC and BIC compare models; lower values are better, and BIC penalizes more heavily.

Understand Goodness of Fit: R-squared and Adjusted R-squared

A regression splits the total variation in Y into two parts. Regression sum of squares (RSS) is the variation the model explains. Sum of squared errors (SSE) is the variation left over. Their total is the total sum of squares (SST): SST = RSS + SSE.

R-squared is the share of total variation that the model explains: R² = RSS ÷ SST. It lies between 0 and 1. An R² of 0.60 means the independent variables explain 60% of the variation in Y. It says nothing about whether the coefficients are significant or the model is correctly specified.

The problem with R² is that it never decreases when you add an independent variable, even a useless one. So it rewards overfitting. Adjusted R² fixes this by adjusting for degrees of freedom. It rises only if the new variable improves the fit by more than chance would. It can be lower than R², and it can even be negative. When k ≥ 1, adjusted R² is always at or below R².

The standard error of estimate (SEE) is the standard deviation of the residuals: √(SSE ÷ (n − k − 1)). Lower is better, and it is in the units of Y.

AIC and BIC are used to choose between models. Both start from the SSE and add a penalty for the number of parameters. Lower is better for both. AIC is preferred when the aim is prediction. BIC is preferred when the aim is a good description of the data, because its penalty is larger and it favors smaller models. These criteria only rank models fitted to the same dependent variable and data. They have no meaning on their own.

Key formulas to remember

Sum of squares identity
SST = RSS + SSE
RSS is explained variation; SSE is unexplained variation.
R-squared
R² = RSS ÷ SST = 1 − SSE ÷ SST
Never falls when a variable is added.
Adjusted R-squared
Adjusted R² = 1 − [(n − 1) ÷ (n − k − 1)] × (1 − R²)
n = observations, k = number of independent variables.
Standard error of estimate
SEE = √[SSE ÷ (n − k − 1)]
Also √MSE. Lower means a tighter fit.
Mean square error
MSE = SSE ÷ (n − k − 1)
Degrees of freedom for SSE are n − k − 1.
AIC
AIC = n × ln(SSE ÷ n) + 2(k + 1)
Lower is better. Preferred for prediction.
BIC
BIC = n × ln(SSE ÷ n) + ln(n) × (k + 1)
Lower is better. Penalty is heavier when n ≥ 8, so it favors simpler models.

How to solve Goodness of Fit: R-squared and Adjusted R-squared questions

Use this order for any goodness-of-fit question in a vignette.

  1. 1Identify what is given: n, k, R², SSE, RSS, SST, or AIC and BIC values. Read the ANOVA table carefully.
  2. 2Check what the question asks: fit, comparison of models, or a flaw in the reasoning.
  3. 3If asked for R², use RSS ÷ SST or 1 − SSE ÷ SST. Find missing sums of squares from SST = RSS + SSE.
  4. 4If asked for adjusted R², compute n − 1 and n − k − 1 first, then apply the formula.
  5. 5If asked for SEE, divide SSE by n − k − 1 and take the square root.
  6. 6To compare models, remember R² always favors the bigger model. Use adjusted R², AIC or BIC instead, and pick the lowest AIC or BIC.
  7. 7Match the criterion to the goal: AIC for forecasting, BIC for parsimonious description.
  8. 8Sanity-check: adjusted R² ≤ R², and both lie below 1.

Quickest way: Compare models by ranking

When to use it: When the vignette gives a table of two or three models and asks which fits best.

  1. Scan the table for the criterion given.
  2. For adjusted R², pick the highest. For AIC or BIC, pick the lowest.
  3. Ignore plain R² when models have different numbers of variables.
  4. If the question states a purpose, choose AIC for prediction and BIC for description.
  5. Only compute by formula if a value is missing.

Common mistakes in Goodness of Fit: R-squared and Adjusted R-squared

  • Choosing the model with the highest R² when models have different numbers of variables.

    R² looks like a score where higher is always better.

    Fix: R² cannot fall when variables are added. Compare adjusted R², AIC or BIC.

  • Picking the highest AIC or BIC as best.

    Students link bigger numbers with better fit.

    Fix: Both are penalized error measures. Lower is better.

  • Using n − k instead of n − k − 1 in adjusted R² and SEE.

    Forgetting the intercept uses one degree of freedom.

    Fix: Always use n − k − 1 as the residual degrees of freedom.

  • Saying a high R² proves the model is correct or the coefficients are significant.

    R² is read as a quality stamp.

    Fix: R² only measures explained variation. Significance needs t-tests and F-tests, and validity needs checks for assumption violations.

  • Saying adjusted R² is always lower than R² by a fixed amount, or never negative.

    Overgeneralizing the usual pattern.

    Fix: It is at or below R² when k ≥ 1, and it can be negative when R² is very low relative to k.

  • Comparing AIC values across models with different dependent variables or data.

    The numbers look comparable.

    Fix: Compare only models fitted to the same dependent variable and sample.

Worked examples

Example 1

An analyst regresses a fund's monthly return on 3 factors using 60 observations. The ANOVA shows RSS = 180 and SSE = 120. (1) What is R²? (2) What is adjusted R²? (3) What is the SEE?

Show the solution
  1. SST = RSS + SSE = 180 + 120 = 300.
  2. R² = 180 ÷ 300 = 0.60.
  3. n − 1 = 59 and n − k − 1 = 60 − 3 − 1 = 56.
  4. Adjusted R² = 1 − (59 ÷ 56) × (1 − 0.60) = 1 − 1.05357 × 0.40 = 1 − 0.42143 = 0.5786.
  5. MSE = 120 ÷ 56 = 2.1429.
  6. SEE = √2.1429 = 1.464.

Answer: R² = 0.60; adjusted R² ≈ 0.579; SEE ≈ 1.46.

Example 2

An analyst fits three models to the same data, with n = 40. Model A (k = 2): AIC = 80.1, BIC = 85.2, adjusted R² = 0.52. Model B (k = 4): AIC = 76.4, BIC = 86.0, adjusted R² = 0.58. Model C (k = 6): AIC = 77.0, BIC = 90.1, adjusted R² = 0.57. (1) Which model is best for forecasting? (2) Which is best if the goal is the most parsimonious description? (3) Does Model C's R² exceed Model B's?

Show the solution
  1. Forecasting points to AIC. The lowest AIC is Model B at 76.4.
  2. A parsimonious description points to BIC. The lowest BIC is Model A at 85.2.
  3. Model C has more variables than B. R² never falls when variables are added, if the same data are used and the models are nested.
  4. Model C's adjusted R² is lower, but that reflects the penalty, not a fall in R².
  5. So C's R² is at least B's, assuming B's variables are a subset of C's.

Answer: (1) Model B. (2) Model A. (3) Yes, if the models are nested; C's R² is at least B's even though its adjusted R² is lower.

Exam tips

  • Expect a vignette with an ANOVA table. Find RSS, SSE and the degrees of freedom before reading the question.
  • Reading the question's purpose is the key. Forecasting means AIC; description means BIC.
  • Questions often test the direction: highest adjusted R², lowest AIC and BIC. Check this before computing.
  • Watch for a conclusion that treats R² as proof of a valid model. It is usually the wrong answer.
  • There is no penalty for wrong answers, so always answer every question.

Goodness of Fit: R-squared and Adjusted R-squared in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Goodness of Fit: R-squared and Adjusted R-squared: frequently asked questions

What is the difference between R-squared and adjusted R-squared?

R-squared is the share of variation in Y explained by the model, and it never falls when you add variables. Adjusted R-squared corrects for the number of variables, so it rises only if a new variable adds real explanatory power. Use adjusted R-squared to compare models with different numbers of variables.

How do I calculate adjusted R-squared?

Use 1 − [(n − 1) ÷ (n − k − 1)] × (1 − R²). Here n is the number of observations and k is the number of independent variables. Work out the two degree-of-freedom terms first to avoid slips.

What is the difference between AIC and BIC?

Both penalize extra parameters and a lower value is better. BIC uses a heavier penalty when n is 8 or more, so it tends to choose smaller models. AIC is preferred for forecasting and BIC for describing the data.

Can adjusted R-squared be negative?

Yes. It can be negative when R² is very low compared with the number of variables and observations. R² itself, for a regression with an intercept, is between 0 and 1.