Skip to content

Actuarial Statistics · Linear regression models

Goodness of Fit and ANOVA for Regression: R² and the F Test

Updated 11 October 2026 · Fact-checked

Goodness of fit measures how much of the variation in y a regression explains. Split total variation SST into regression SSR and residual SSE. R² = SSR ÷ SST. The ANOVA F test, F = MSR ÷ MSE, tests whether the slope is zero. In simple regression, R² equals r².

Understand Goodness of Fit and ANOVA for Regression

A regression line never fits the data perfectly. You need a number that says how well it fits. The idea is to split the variation in y into two parts: the part the line explains and the part it leaves over.

Start with the total sum of squares, SST = Σ(yᵢ − ȳ)². This is the variation in y if you ignore x. Now fit the line. The residual sum of squares, SSE = Σ(yᵢ − ŷᵢ)², is the variation left after using x. The difference is the regression sum of squares, SSR = Σ(ŷᵢ − ȳ)². It is the variation the line explains. For a model with an intercept, SST = SSR + SSE.

The coefficient of determination is R² = SSR ÷ SST = 1 − SSE ÷ SST. It lies between 0 and 1. A value of 0.85 means 85% of the variation in y is explained by the model. It does not mean the model is correct or that x causes y.

R² describes fit, but it does not test anything. For a test, use the ANOVA table. Divide each sum of squares by its degrees of freedom to get mean squares. In simple linear regression with n points, SSR has 1 degree of freedom, SSE has n − 2 and SST has n − 1. Under H₀: β = 0, with normal errors, F = MSR ÷ MSE follows an F distribution with 1 and n − 2 degrees of freedom.

In simple regression this F test is the same as the two-sided t test on the slope, because F = t². Also, R² = r², where r is the sample correlation between x and y. The sign of r is the sign of the slope, but R² has no sign.

Key rules to remember

Sums of squares
SST = Σ(yᵢ − ȳ)² = Syy; SSR = Σ(ŷᵢ − ȳ)²; SSE = Σ(yᵢ − ŷᵢ)²
Syy is the usual notation for the corrected sum of squares of y.
Decomposition
SST = SSR + SSE
Holds when the model includes an intercept.
Coefficient of determination
R² = SSR ÷ SST = 1 − SSE ÷ SST
Proportion of variation in y explained by the model.
Simple regression shortcuts
SSR = Sxy² ÷ Sxx = β̂ × Sxy; SSE = Syy − Sxy² ÷ Sxx
Here Sxx = Σ(xᵢ − x̄)², Sxy = Σ(xᵢ − x̄)(yᵢ − ȳ) and β̂ = Sxy ÷ Sxx.
Link to correlation
r = Sxy ÷ √(Sxx × Syy); R² = r²
Valid for simple linear regression only.
Degrees of freedom
SSR: 1; SSE: n − 2; SST: n − 1
For multiple regression with k explanatory variables: k, n − k − 1, n − 1.
Mean squares and F statistic
MSR = SSR ÷ 1; MSE = SSE ÷ (n − 2) = σ̂²; F = MSR ÷ MSE ~ F(1, n − 2) under H₀: β = 0
Reject H₀ for large F. The test is one-sided in F but equals a two-sided t test.
F and t link
F = t², where t = β̂ ÷ se(β̂)
t has n − 2 degrees of freedom.
F in terms of R²
F = R² × (n − 2) ÷ (1 − R²)
Useful when only R² and n are given.

How to solve Goodness of Fit and ANOVA for Regression questions

Use this order for any question on sums of squares, R² or the ANOVA table in simple linear regression.

  1. 1Write down what you are given: n, Sxx, Sxy, Syy, or some of SST, SSR, SSE, R², r.
  2. 2Find the missing sums of squares. Use SST = SSR + SSE, or SSR = Sxy² ÷ Sxx and SST = Syy.
  3. 3Compute R² = SSR ÷ SST. State it in words as the proportion of variation in y explained by x.
  4. 4Set out the degrees of freedom: 1, n − 2 and n − 1. Check that 1 + (n − 2) = n − 1.
  5. 5Compute MSR, MSE and F = MSR ÷ MSE. Note that MSE is the estimate of σ².
  6. 6State H₀: β = 0 against H₁: β ≠ 0. Compare F with the F(1, n − 2) critical value from the Tables, and write the conclusion in context.
  7. 7If asked about correlation, use r = ±√R² with the sign of the slope, or check that F = t².

Quickest way: Fill the ANOVA table from three numbers

When to use it: Use when you are given n, Sxx, Sxy and Syy, or enough to get them. It is the fastest route to R² and F.

  1. Compute SSR = Sxy² ÷ Sxx.
  2. Compute SSE = Syy − SSR.
  3. R² = SSR ÷ Syy.
  4. F = SSR ÷ (SSE ÷ (n − 2)).
  5. If only R² and n are given, jump to F = R²(n − 2) ÷ (1 − R²).
  6. Compare F with the 5% point of F(1, n − 2) and write a one-line conclusion.

Common mistakes in Goodness of Fit and ANOVA for Regression

  • Using n − 1 degrees of freedom for SSE.

    n − 1 is familiar from the sample variance, so it gets used without thought.

    Fix: SSE has n − 2 degrees of freedom because two parameters, α and β, are estimated. n − 1 belongs to SST.

  • Saying R² = 0.9 means the model is correct or that x causes y.

    A high number feels like proof.

    Fix: Say only that 90% of the variation in y is explained by the fitted line. Check residual plots for the model assumptions and remember that correlation is not causation.

  • Treating R² and r as the same number.

    In simple regression R² = r², so they are mixed up.

    Fix: R² = r². If r = −0.8, then R² = 0.64. To go back, take r = ±√R² and use the sign of the slope.

  • Applying R² = r² to multiple regression.

    The simple regression result is remembered as a general rule.

    Fix: In multiple regression, R² is the squared correlation between observed y and fitted ŷ. It is not the square of any single x–y correlation.

  • Taking F = SSR ÷ SSE without dividing by degrees of freedom.

    The mean squares step is skipped to save time.

    Fix: Always divide first. F = MSR ÷ MSE, with MSR = SSR ÷ 1 and MSE = SSE ÷ (n − 2).

  • Believing a larger R² always means a better model.

    R² never falls when you add variables.

    Fix: In multiple regression, compare models with adjusted R² or F tests. Adjusted R² = 1 − [SSE ÷ (n − k − 1)] ÷ [SST ÷ (n − 1)].

Worked examples

Example 1

For 12 observations of (x, y): Sxx = 40, Sxy = 60, Syy = 120. Fit the simple linear regression of y on x. Find R² and complete the ANOVA F test at the 5% level. The 5% critical value of F(1, 10) is 4.965.

Show the solution
  1. β̂ = Sxy ÷ Sxx = 60 ÷ 40 = 1.5.
  2. SST = Syy = 120.
  3. SSR = Sxy² ÷ Sxx = 3,600 ÷ 40 = 90.
  4. SSE = 120 − 90 = 30.
  5. R² = 90 ÷ 120 = 0.75.
  6. Degrees of freedom: regression 1, residual n − 2 = 10, total 11.
  7. MSR = 90 ÷ 1 = 90. MSE = 30 ÷ 10 = 3.
  8. F = 90 ÷ 3 = 30.
  9. H₀: β = 0 against H₁: β ≠ 0. Since 30 > 4.965, reject H₀.

Answer: R² = 0.75, so 75% of the variation in y is explained by x. F = 30 on (1, 10) degrees of freedom, which exceeds 4.965. There is strong evidence at the 5% level that the slope is not zero.

Example 2

A simple linear regression on 22 data points gives R² = 0.64 and a negative fitted slope. (a) Find the sample correlation coefficient. (b) Find the F statistic. (c) Find the absolute value of the t statistic for the slope.

Show the solution
  1. (a) r = ±√R² = ±√0.64 = ±0.8. The slope is negative, so r = −0.8.
  2. (b) Here n − 2 = 20. F = R² × (n − 2) ÷ (1 − R²) = 0.64 × 20 ÷ 0.36.
  3. 0.64 × 20 = 12.8, and 12.8 ÷ 0.36 = 35.56 (to 2 decimal places).
  4. (c) Since F = t², |t| = √35.56 = 5.96 (to 2 decimal places).
  5. Check using r: t = r√(n − 2) ÷ √(1 − r²) = −0.8 × √20 ÷ 0.6 = −5.96, which agrees in size.

Answer: (a) r = −0.8. (b) F ≈ 35.56 on (1, 20) degrees of freedom. (c) |t| ≈ 5.96 on 20 degrees of freedom, which is significant at the 5% level.

Exam tips

  • In written questions, show the ANOVA table in full: source, sum of squares, degrees of freedom, mean square and F. Marks are given for each column.
  • Check your table: SSR + SSE must equal SST, and the degrees of freedom must add up to n − 1.
  • Interpret in words. Write what R² means in context, for example the proportion of variation in claim cost explained by the rating factor.
  • In MCQs, look for the shortcut: F = t² and R² = r² often give the answer without computing the full table.
  • In the computer-based paper, the R summary of lm() gives R², the F statistic and its p-value. Read them, then state the conclusion and its assumptions.

Practice questions from Linear regression models

Goodness of Fit and ANOVA for Regression in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Goodness of Fit and ANOVA for Regression: frequently asked questions

What is the difference between the correlation coefficient and R squared?

The correlation coefficient r measures the strength and direction of the linear relationship, and lies between −1 and 1. R² is the proportion of variation in y explained by the model, and lies between 0 and 1. In simple linear regression, R² = r².

What does the ANOVA F test check in simple linear regression?

It tests H₀: β = 0 against H₁: β ≠ 0, that is, whether x helps explain y at all. The statistic is F = MSR ÷ MSE on (1, n − 2) degrees of freedom. It gives the same conclusion as the two-sided t test on the slope.

Can R squared be negative?

Not for a least squares fit with an intercept, since SSR and SST are both non-negative and SSR ≤ SST. It lies between 0 and 1. Adjusted R² can be negative if the model fits very poorly.

Does a high R squared mean the regression is valid?

No. It shows only that the line explains a large share of the variation. You still need to check the assumptions, such as linearity, constant variance and normal errors, using residual plots.