CFA Level II Exam · Basics of Multiple Regression and Underlying Assumptions
ANOVA, F-Test, R-squared and Adjusted R-squared
Updated 7 October 2026 · Fact-checked
The ANOVA table splits total variation in Y into explained (regression) and unexplained (residual) parts. F = MSR ÷ MSE tests whether all slope coefficients are jointly zero. R-squared = SSR ÷ SST measures fit. Adjusted R-squared penalises extra variables, so it can fall when a useless variable is added.
Understand ANOVA, F-Test, R-squared and Adjusted R-squared
Multiple regression tries to explain variation in a dependent variable Y using k independent variables. The total variation in Y is the sum of squared deviations from its mean, called SST. The regression splits it into two pieces: SSR (the part the model explains) and SSE (the part left in the residuals). So SST = SSR + SSE.
The ANOVA table lays this out. For each source (regression, error, total) it gives degrees of freedom (df), sum of squares (SS) and mean square (MS). Regression df = k. Error df = n − k − 1. Total df = n − 1. A mean square is a sum of squares divided by its df. So MSR = SSR ÷ k and MSE = SSE ÷ (n − k − 1).
The F-test asks one question: do the slope coefficients together explain anything? H0: all slopes equal zero. Ha: at least one slope is not zero. F = MSR ÷ MSE, with k and n − k − 1 degrees of freedom. It is a one-tailed, right-tail test. Reject H0 if F is above the critical value. Use it for joint significance, not for single coefficients (that is the t-test's job).
R-squared = SSR ÷ SST is the share of variation in Y explained by the model. It never falls when you add a variable, even a useless one. So it rewards complexity. Adjusted R-squared corrects this by scaling with degrees of freedom. It rises only if the new variable improves fit enough to offset the lost df. It is never above R-squared in a model with k ≥ 1.
The standard error of estimate (SEE) = √MSE. It is the typical size of a residual, in the units of Y. Lower is better. For comparing models, AIC and BIC are also used. Both start from SSE and add a penalty for the number of parameters. Lower values are better. AIC suits prediction. BIC penalises extra parameters more heavily, so it suits finding the best-fitting parsimonious model. Both compare models with the same dependent variable and sample.
Key formulas to remember
- Variation decomposition
- SST = SSR + SSE
- SST is total, SSR is explained (regression), SSE is unexplained (residual).
- Degrees of freedom
- Regression = k; Error = n − k − 1; Total = n − 1
- k is the number of independent variables, n the number of observations.
- Mean squares
- MSR = SSR ÷ k; MSE = SSE ÷ (n − k − 1)
- Mean square = SS ÷ df.
- F-statistic
- F = MSR ÷ MSE, df = k and n − k − 1
- One-tailed right-tail test of H0: all slopes = 0.
- R-squared
- R² = SSR ÷ SST = 1 − SSE ÷ SST
- Never decreases when a variable is added.
- Adjusted R-squared
- Adj R² = 1 − [(n − 1) ÷ (n − k − 1)] × (1 − R²)
- Can fall when a variable is added. Not above R² when k ≥ 1.
- Standard error of estimate
- SEE = √MSE = √[SSE ÷ (n − k − 1)]
- In the units of the dependent variable.
- AIC
- AIC = n × ln(SSE ÷ n) + 2(k + 1)
- Lower is better. Compare models on the same data and dependent variable.
- BIC
- BIC = n × ln(SSE ÷ n) + ln(n) × (k + 1)
- Lower is better. Heavier penalty than AIC when n ≥ 8.
How to solve ANOVA, F-Test, R-squared and Adjusted R-squared questions
Most questions give an ANOVA table, partial numbers, or R-squared with n and k. Pull out what you have and fill in the rest.
- 1Read n and k from the vignette. If only df is given, use total df = n − 1 and regression df = k.
- 2Write the table skeleton: SSR, SSE, SST, with df k, n − k − 1, n − 1. Fill known cells and use SST = SSR + SSE to find missing ones.
- 3Compute mean squares: MSR = SSR ÷ k, MSE = SSE ÷ (n − k − 1).
- 4For joint significance, compute F = MSR ÷ MSE and compare with the critical F at k and n − k − 1 df. Reject H0 if F is larger.
- 5For fit, compute R² = SSR ÷ SST, then adjusted R² with the n − 1 and n − k − 1 terms.
- 6For SEE take √MSE. Do not use SSE or MSR by mistake.
- 7When asked to choose between models, pick the one with the lower AIC or BIC, or higher adjusted R², and check the criterion matches the purpose.
- 8Answer in the form asked: reject or fail to reject, a number, or a comparison. State the conclusion in context.
Quickest way: Table-fill shortcut
When to use it: Use when the vignette gives an ANOVA table with one or two blank cells, or gives R² with n and k.
- Write k and n − k − 1 beside the table first.
- Use SST = SSR + SSE to get the missing sum of squares.
- Divide each SS by its df, then divide MSR by MSE for F.
- For adjusted R², use 1 − (1 − R²) × (n − 1) ÷ (n − k − 1).
- Eliminate options that put adjusted R² above R² or give a negative F.
Common mistakes in ANOVA, F-Test, R-squared and Adjusted R-squared
Using n − k instead of n − k − 1 as the error degrees of freedom.
Simple regression uses n − 2, and students drift to a memorised pattern.
Fix: Always count the intercept. Error df = n − (k + 1). Check that regression df + error df = n − 1.
Treating F as a test of each coefficient.
Students link significance with any single test statistic.
Fix: F tests all slopes jointly. Use t-statistics for individual slopes. A significant F does not mean every slope is significant.
Taking SEE as MSE or SSE.
The names sound alike and the table shows MSE directly.
Fix: SEE = √MSE. Take the square root before answering.
Assuming a higher R-squared means a better model after adding variables.
R² always rises or stays the same when variables are added.
Fix: Compare adjusted R², AIC or BIC when models differ in the number of variables.
Choosing the higher AIC or BIC as better.
Students are used to higher being better, as with R².
Fix: For AIC and BIC, lower is better. Remember BIC penalises extra parameters more heavily.
Using a two-tailed critical value for F.
The t-test is usually two-tailed, so students carry over the habit.
Fix: The F-test is one-tailed. Only a large F rejects H0.
Worked examples
Example 1
An analyst regresses monthly returns of a global equity fund on 3 factors using 40 observations. The ANOVA table shows SSR = 90 and SSE = 72. (1) Compute R². (2) Compute F. (3) The critical F at 5% with 3 and 36 df is about 2.87. What is the conclusion?
Show the solution
- k = 3, n = 40. Error df = 40 − 3 − 1 = 36. SST = 90 + 72 = 162.
- R² = 90 ÷ 162 = 0.5556.
- MSR = 90 ÷ 3 = 30. MSE = 72 ÷ 36 = 2.
- F = 30 ÷ 2 = 15.
- 15 is greater than 2.87, so reject H0.
Answer: R² ≈ 55.6%, F = 15. Reject H0: at least one slope differs from zero.
Example 2
A researcher fits Model A with 2 variables and Model B with 3 variables on the same 31 observations. Model A has R² = 0.60. Model B has R² = 0.61. (1) Compute adjusted R² for each. (2) Which model does adjusted R² favour? (3) Compute SEE for Model B if SSE = 54.
Show the solution
- Model A: k = 2, n − 1 = 30, n − k − 1 = 28. Adj R² = 1 − (30 ÷ 28) × 0.40 = 1 − 0.4286 = 0.5714.
- Model B: k = 3, n − k − 1 = 27. Adj R² = 1 − (30 ÷ 27) × 0.39 = 1 − 0.4333 = 0.5667.
- Model A has the higher adjusted R² (0.5714 versus 0.5667), so it is favoured.
- Model B: MSE = 54 ÷ 27 = 2. SEE = √2 = 1.414.
Answer: Adjusted R² is 57.1% for A and 56.7% for B, so A is favoured. SEE for B ≈ 1.414.
Exam tips
- Expect a partly filled ANOVA table. Fill it using SST = SSR + SSE and the df rules before anything else.
- When R² rises but adjusted R² falls, the added variable is not worth its cost in degrees of freedom.
- For AIC and BIC questions, the answer is the lowest value. Check which criterion the question names.
- Read the hypothesis carefully. Joint tests use F. A single coefficient uses t. Do not mix them up.
- Scan the vignette for n and k before you start. Many errors come from miscounting the intercept.
ANOVA, F-Test, R-squared and Adjusted R-squared in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
ANOVA, F-Test, R-squared and Adjusted R-squared: frequently asked questions
How do I calculate the F statistic from an ANOVA table?
Divide SSR by k to get MSR. Divide SSE by n − k − 1 to get MSE. Then F = MSR ÷ MSE. Compare it with the critical F value at k and n − k − 1 degrees of freedom.
What is the difference between R-squared and adjusted R-squared?
R-squared is the share of variation explained and never falls when you add variables. Adjusted R-squared adjusts for degrees of freedom and falls if a new variable adds too little. Use adjusted R-squared to compare models with different numbers of variables.
Can adjusted R-squared be negative?
Yes. If R² is low relative to the number of variables, the adjustment can push adjusted R² below zero. It is never above R² when there is at least one independent variable.
AIC or BIC: which one should I use?
Both are lower-is-better. AIC is aimed at prediction and BIC at picking the best-fitting model with fewer parameters, since its penalty is larger. Use the one the question names or fits its stated goal.
If F is significant, are all coefficients significant?
No. A significant F only says at least one slope is not zero. Individual t-tests show which ones matter.