Skip to content

Actuarial Statistics · Linear regression models

Least Squares Estimation of Slope and Intercept

Updated 11 October 2026 · Fact-checked

Least squares estimation picks the line y = α + βx that minimises the sum of squared vertical errors. Calculate Sxx, Sxy and Syy from the data. The slope is β̂ = Sxy ÷ Sxx and the intercept is α̂ = ȳ − β̂x̄. Under the model assumptions, both estimators are unbiased.

Understand Least Squares Estimation of Parameters

In simple linear regression you model a response as Yi = α + βxi + εi. The x values are fixed. The errors εi have mean 0 and constant variance σ², and are uncorrelated. You want estimates of α and β from the data.

The least squares idea is simple. For any trial line, each point has a vertical error (the residual). Square each error and add them up. Choose α and β that make this sum as small as possible. Call this sum S(α, β) = Σ(yi − α − βxi)².

To minimise S, set its partial derivatives with respect to α and β to zero. This gives the normal equations: Σy = nα + βΣx and Σxy = αΣx + βΣx². Solving them gives the formulas in terms of Sxx, Sxy and Syy. These are sums of squares and cross-products about the means. They are the building blocks of almost every regression calculation in the syllabus.

The fitted line always passes through the point (x̄, ȳ). That is why the intercept is ȳ − β̂x̄. The residuals add up to zero. This also acts as a quick check on your work.

The estimators have useful properties. Both are unbiased. Their variances depend on σ² and Sxx, so spread-out x values give a more precise slope. By the Gauss-Markov theorem, among all linear unbiased estimators they have the smallest variance (they are BLUE). If the errors are also normal, they are the maximum likelihood estimators of α and β and are themselves normally distributed.

Key rules to remember

Sxx
Sxx = Σ(xi − x̄)² = Σxi² − n x̄²
Use the second form for calculations. It needs only Σx and Σx².
Sxy
Sxy = Σ(xi − x̄)(yi − ȳ) = Σxiyi − n x̄ ȳ
Sign of Sxy is the sign of the slope.
Syy
Syy = Σ(yi − ȳ)² = Σyi² − n ȳ²
Needed for σ̂², SSE and R².
Normal equations
Σy = nα + βΣx; Σxy = αΣx + βΣx²
Obtained by setting ∂S/∂α = 0 and ∂S/∂β = 0.
Least squares slope
β̂ = Sxy ÷ Sxx
Also equal to Σ(xi − x̄)yi ÷ Sxx.
Least squares intercept
α̂ = ȳ − β̂ x̄
Fitted line passes through (x̄, ȳ).
Mean of estimators
E(β̂) = β and E(α̂) = α
Both are unbiased under the model assumptions.
Variance of slope
Var(β̂) = σ² ÷ Sxx
Larger spread in x gives smaller variance.
Variance of intercept
Var(α̂) = σ² (1/n + x̄²/Sxx)
Smallest when x̄ = 0.
Covariance
Cov(α̂, β̂) = −σ² x̄ ÷ Sxx
Zero only if x̄ = 0.
Residual sum of squares
SSE = Syy − Sxy² ÷ Sxx
Equals Σ(yi − ŷi)².
Unbiased estimator of σ²
σ̂² = SSE ÷ (n − 2)
Divide by n − 2 because two parameters are estimated.

How to solve Least Squares Estimation of Parameters questions

Use this order for any question that asks you to fit a line, derive the estimators or state their properties.

  1. 1Write down the model and assumptions: Yi = α + βxi + εi, with E(εi) = 0, Var(εi) = σ², uncorrelated errors.
  2. 2List n, Σx, Σy, Σx², Σxy and, if needed, Σy². Compute x̄ and ȳ.
  3. 3Compute Sxx = Σx² − n x̄² and Sxy = Σxy − n x̄ ȳ. Compute Syy if the question needs SSE or σ̂².
  4. 4Calculate β̂ = Sxy ÷ Sxx, then α̂ = ȳ − β̂ x̄. State the fitted line ŷ = α̂ + β̂x.
  5. 5If asked for σ̂², compute SSE = Syy − Sxy² ÷ Sxx, then divide by n − 2.
  6. 6For standard errors, substitute σ̂² into Var(β̂) = σ²/Sxx or Var(α̂) = σ²(1/n + x̄²/Sxx), then take the square root.
  7. 7For a derivation or proof, write β̂ as a linear combination Σ ci Yi with ci = (xi − x̄)/Sxx, then use Σci = 0 and Σci xi = 1.
  8. 8Check: sense of the sign of the slope, and that the line passes through (x̄, ȳ). Give the answer with units and an interpretation.

Quickest way: Summation-table shortcut

When to use it: Use when you are given raw data or the totals Σx, Σy, Σx², Σxy and need the line quickly.

  1. Get x̄ = Σx ÷ n and ȳ = Σy ÷ n.
  2. Compute Sxx = Σx² − n x̄² and Sxy = Σxy − n x̄ ȳ. Do not subtract means from every data point.
  3. Slope = Sxy ÷ Sxx. Intercept = ȳ − slope × x̄.
  4. Check that the fitted line gives ȳ when you plug in x̄.
  5. For SSE use Syy − Sxy²/Sxx. Do not compute residuals one by one.
  6. Keep full decimals until the end. Round only the final answer.

Common mistakes in Least Squares Estimation of Parameters

  • Dividing by n − 1 or n instead of n − 2 when estimating σ².

    Students copy the sample variance rule, which loses only one degree of freedom.

    Fix: Regression estimates two parameters, α and β. So σ̂² = SSE ÷ (n − 2) is the unbiased estimator.

  • Using Σx² in place of Sxx, or Σxy in place of Sxy.

    The raw sums look similar to the corrected sums and the subtraction of n x̄² or n x̄ ȳ is forgotten.

    Fix: Always write Sxx = Σx² − n x̄² and Sxy = Σxy − n x̄ ȳ before computing the slope.

  • Mixing up the roles of x and y, giving Sxy ÷ Syy as the slope.

    Students forget which variable is the response. Regressing x on y gives a different line.

    Fix: The slope of y on x is Sxy ÷ Sxx. The denominator is always the sum of squares of the explanatory variable.

  • Writing Var(α̂) = σ²/n or Var(α̂) = σ²/Sxx.

    Students half-remember the variance of ȳ and of β̂.

    Fix: Remember Var(α̂) = σ²(1/n + x̄²/Sxx). It is the variance of ȳ plus the extra uncertainty from β̂.

  • Saying the estimators are unbiased only if the errors are normal.

    Confusing the unbiasedness result with the distribution result.

    Fix: Unbiasedness, the variance formulas and Gauss-Markov need only zero mean errors, constant variance and uncorrelated errors. Normality is needed for exact distributions, t tests and the MLE link.

  • Treating the estimate as exact when predicting far outside the data.

    The formula gives a number for any x.

    Fix: State that extrapolation beyond the observed range of x relies on the linear model still holding, which the data cannot confirm.

Worked examples

Example 1

Five policyholders have x = years of driving experience and y = annual claim cost (₹ thousands): (1, 3), (2, 5), (3, 4), (4, 7), (5, 8). Fit the least squares line of y on x, predict y at x = 6, and estimate σ².

Show the solution
  1. n = 5. Σx = 15, Σy = 27, Σx² = 1 + 4 + 9 + 16 + 25 = 55, Σy² = 9 + 25 + 16 + 49 + 64 = 163.
  2. Σxy = 3 + 10 + 12 + 28 + 40 = 93. x̄ = 3 and ȳ = 5.4.
  3. Sxx = 55 − 5 × 3² = 55 − 45 = 10.
  4. Sxy = 93 − 5 × 3 × 5.4 = 93 − 81 = 12.
  5. β̂ = 12 ÷ 10 = 1.2. α̂ = 5.4 − 1.2 × 3 = 1.8. Fitted line: ŷ = 1.8 + 1.2x.
  6. Prediction at x = 6: 1.8 + 1.2 × 6 = 9.
  7. Syy = 163 − 5 × 5.4² = 163 − 145.8 = 17.2.
  8. SSE = Syy − Sxy²/Sxx = 17.2 − 144/10 = 17.2 − 14.4 = 2.8.
  9. σ̂² = 2.8 ÷ (5 − 2) = 0.9333 (to 4 decimal places).

Answer: ŷ = 1.8 + 1.2x. The predicted value at x = 6 is 9 (₹9 thousand). σ̂² ≈ 0.9333.

Example 2

For the model Yi = α + βxi + εi with E(εi) = 0 and Var(εi) = σ², uncorrelated, show that the least squares estimator β̂ = Sxy/Sxx is unbiased and has variance σ²/Sxx.

Show the solution
  1. Since Σ(xi − x̄)ȳ = ȳ Σ(xi − x̄) = 0, we have Sxy = Σ(xi − x̄)Yi. So β̂ = Σ ci Yi with ci = (xi − x̄)/Sxx.
  2. Useful facts: Σci = 0, and Σci xi = Σ(xi − x̄)xi / Sxx = Sxx/Sxx = 1. Also Σci² = Sxx/Sxx² = 1/Sxx.
  3. E(β̂) = Σ ci E(Yi) = Σ ci (α + βxi) = α Σci + β Σci xi = 0 + β = β.
  4. The Yi are uncorrelated with variance σ², so Var(β̂) = Σ ci² Var(Yi) = σ² Σ ci² = σ²/Sxx.

Answer: E(β̂) = β, so β̂ is unbiased, and Var(β̂) = σ²/Sxx.

Exam tips

  • Show the three sums Sxx, Sxy and Syy clearly. Marks are often given for each, even if the final arithmetic slips.
  • For derivation questions, set up β̂ as Σ ci Yi with ci = (xi − x̄)/Sxx. This makes unbiasedness and variance short.
  • State which assumptions you use at each step. Examiners separate results that need only zero mean and constant variance from those that need normality.
  • Use SSE = Syy − Sxy²/Sxx rather than computing residuals, and divide by n − 2.
  • In the computer-based paper, check that your R or Excel output matches your hand values for β̂ and α̂. Name the function you used and interpret the coefficients in context.

Practice questions from Linear regression models

Least Squares Estimation of Parameters in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Least Squares Estimation of Parameters: frequently asked questions

Why do we divide Sxy by Sxx to get the slope?

Setting the derivative of the sum of squared errors to zero gives the normal equations. Solving them leaves β̂ = Sxy/Sxx. Sxy measures how x and y move together. Sxx measures how much x varies.

Why is σ̂² divided by n − 2?

Two parameters, α and β, are estimated from the data, so two degrees of freedom are used. Dividing SSE by n − 2 makes σ̂² an unbiased estimator of σ².

Do I need normal errors for least squares estimates to be unbiased?

No. Unbiasedness and the variance formulas need only zero mean, constant variance and uncorrelated errors. Normality is needed for exact t and F results and for least squares to equal maximum likelihood.

What does the Gauss-Markov theorem say?

Under the standard assumptions, the least squares estimators have the smallest variance among all linear unbiased estimators. They are the best linear unbiased estimators, or BLUE.

Does the regression line always pass through (x̄, ȳ)?

Yes, when the model includes an intercept. It follows directly from α̂ = ȳ − β̂x̄. You can use it as a quick check.