Skip to content

CFA Level I · CFA Level I Exam

Applications of Simple Linear Regression in Finance: formula sheet

Full chapter guide

Key formulas

Population regression model
Yi = b0 + b1 Xi + εi
Y is dependent, X is independent, ε is the error term. i = 1, ..., n.
Estimated regression line
Ŷi = b̂0 + b̂1 Xi
Ŷ is the predicted value from the fitted line.
Residual
ei = Yi − Ŷi
OLS chooses b̂0 and b̂1 to minimize Σ ei².
Slope estimate
b̂1 = Cov(X, Y) ÷ Var(X)
Sample covariance and variance must use the same denominator (n − 1).
Intercept estimate
b̂0 = Ȳ − b̂1 X̄
The fitted line always passes through the point (X̄, Ȳ).
Core assumptions
Linearity; homoskedasticity; independence of errors; normality of errors
Errors have an expected value of zero. X is not constant and is uncorrelated with ε.
OLS objective
Minimise Σ(Yi − b0 − b1Xi)²
The sum of squared residuals. OLS chooses b0 and b1 to make this as small as possible.
Slope coefficient
b1 = Cov(X, Y) ÷ Var(X) = Σ(Xi − X̄)(Yi − Ȳ) ÷ Σ(Xi − X̄)²
X is the independent variable and goes in the denominator. If you use sample covariance and variance, use the same divisor (n − 1) for both. The divisor then cancels.
Intercept
b0 = Ȳ − b1 × X̄
Calculate the slope first. This forces the line through the point of means.
Slope from correlation
b1 = r × (s_Y ÷ s_X)
Useful when you are given the correlation and the two standard deviations.
Residual
ei = Yi − Ŷi = Yi − (b0 + b1Xi)
OLS residuals sum to zero when the model includes an intercept.
Fitted value
Ŷ = b0 + b1X
Use it to predict Y for a given X.
Sum of squares decomposition
SST = SSR + SSE
SST = Σ(Yi − Ȳ)²; SSR = Σ(Ŷi − Ȳ)²; SSE = Σ(Yi − Ŷi)².
Coefficient of determination
R² = SSR ÷ SST = 1 − SSE ÷ SST
Lies between 0 and 1. In simple regression, R² = r².
Mean square regression
MSR = SSR ÷ k, with k = 1
k is the number of independent variables.
Mean square error
MSE = SSE ÷ (n − 2)
Degrees of freedom are n − 2 in simple regression.
Standard error of estimate
SEE = √MSE = √[SSE ÷ (n − 2)]
In the units of the dependent variable.
F-statistic
F = MSR ÷ MSE, df = 1 and n − 2
Tests H0: slope = 0. One-tailed, reject if F is above the critical value. F = t² for the slope.
Correlation from R²
r = ±√R²
Take the sign of the slope coefficient.
t-statistic for a coefficient
t = (b̂1 − B1) ÷ s_b1
B1 is the hypothesized slope, often 0. For the intercept use b̂0, B0 and s_b0. Degrees of freedom = n − 2.
Standard error of the slope
s_b1 = s_e ÷ √Σ(Xi − X̄)²
s_e is the standard error of the estimate. A larger s_e raises the slope's standard error. More spread in X lowers it.
Confidence interval for the slope
b̂1 ± t_c × s_b1
t_c is the critical value for n − 2 degrees of freedom. If the hypothesized value lies outside the interval, reject H0 at the matching significance level.
Decision rule (critical value)
Reject H0 if |t| > t_c (two-tailed)
For a one-tailed test, use the one-tailed critical value and check the sign of t.
Decision rule (p-value)
Reject H0 if p-value < α
The p-value is the smallest significance level at which H0 can be rejected.
t-test for correlation
t = r√(n − 2) ÷ √(1 − r²)
Tests H0: ρ = 0 with n − 2 degrees of freedom. It equals the slope t-statistic in simple regression.
Link to ANOVA
F = t² (slope, simple regression)
The F-test of the slope and the two-tailed t-test give the same conclusion. Also R² = r².
Predicted value
Ŷ = b0 + b1 × X_f
X_f is the given value of the independent variable. Use unrounded b0 and b1 if given.
Variance of the forecast
s_f² = s_e² × [1 + 1/n + (X_f − X̄)² ÷ ((n − 1) × s_x²)]
s_x² is the sample variance of X. (n − 1) × s_x² equals Σ(Xi − X̄)².
Standard error of forecast
s_f = √(s_f²)
Always greater than s_e. Smallest when X_f = X̄.
Prediction interval
Ŷ ± t(α/2, n − 2) × s_f
Two-tailed critical t with n − 2 degrees of freedom. Simple regression loses two degrees of freedom.
Standard error of estimate
s_e = √(SSE ÷ (n − 2))
Often given in the question or taken from the ANOVA table as √MSE.
Log-lin model
ln(Y) = b0 + b1X + ε
Y is logged, X is not. A 1-unit rise in X changes Y by about 100 × b1 percent (relative change). Used for constant growth with X as time.
Lin-log model
Y = b0 + b1 ln(X) + ε
X is logged, Y is not. A 1% change in X changes Y by about b1 ÷ 100 units. Slope is an absolute change in Y.
Log-log model
ln(Y) = b0 + b1 ln(X) + ε
Both logged. A 1% change in X changes Y by about b1 percent. b1 is the elasticity of Y with respect to X.
Predicting Y from a log-dependent model
Ŷ = e^(predicted ln Y)
If the dependent variable is ln(Y), compute the fitted ln(Y) first, then exponentiate to get Y in original units.
Constant growth as log-lin
ln(Y_t) = b0 + b1 t, growth rate ≈ b1 per period (continuously compounded)
The equivalent periodic growth rate is e^b1 − 1.
Choosing a form
Pick the form whose residuals show no pattern and that fits well
Check the scatter plot, residual plot, and R² or standard error only among models with the same dependent variable.

Quick revision

  • Model: Y = b0 + b1X + ε, with one independent variable.
  • Slope b1 = Cov(X,Y) ÷ Var(X); intercept b0 = Ȳ − b1X̄.
  • The fitted line always passes through the point (X̄, Ȳ).
  • SST = SSR + SSE; R² = SSR ÷ SST.
  • In simple regression, R² equals the square of the correlation between X and Y.
  • Standard error of estimate = √(SSE ÷ (n − 2)); a lower value means a better fit.
  • Degrees of freedom for the slope t-test are n − 2.
  • Slope t-statistic = (b1 − hypothesised value) ÷ s(b1); reject H0 if |t| exceeds the critical value.
  • In simple regression, the F-statistic = MSR ÷ MSE equals t² for the slope.
  • A prediction interval is Ŷ ± tc × sf, and it is wider than the interval for the mean response.
  • Log-lin: ln Y on X; log-log: ln Y on ln X, where the slope is an elasticity-type measure.
  • Check residual plots for heteroskedasticity and non-normality before trusting the results.

Common mistakes

  • Swapping dependent and independent variables. Fix: Ask which variable is being explained or predicted. That is Y, whatever the order in the sentence.
  • Interpreting the intercept as always meaningful. Fix: State it as the expected Y when X = 0, and treat it as an extrapolation if X = 0 is not realistic.
  • Swapping X and Y, so the slope is Cov(X, Y) ÷ Var(Y). Fix: The variable in the denominator is always the independent variable. Underline Y and X in the stem before calculating.
  • Computing the intercept as Ȳ + b1 × X̄ or X̄ − b1 × Ȳ. Fix: Rearrange from Ŷ = b0 + b1X at the means: Ȳ = b0 + b1X̄, so b0 = Ȳ − b1X̄.
  • Dividing SSE by n instead of n − 2 when computing SEE. Fix: Two parameters are estimated in simple regression, so always use n − 2.
  • Forgetting the square root and reporting MSE as the SEE. Fix: SEE = √MSE. Check that the units match Y, not Y squared.
  • Using n − 1 degrees of freedom instead of n − 2. Fix: In simple linear regression two parameters are estimated, so df = n − 2 for coefficient and correlation tests.
  • Testing against zero when the question hypothesizes another value, such as a beta of 1. Fix: Always subtract the hypothesized value from the estimate before dividing by the standard error.
  • Using s_e directly as the margin of error instead of s_f. Fix: s_e only measures residual noise. The interval needs s_f, which adds the 1/n and (X_f − X̄)² terms.
  • Using n − 1 degrees of freedom for the t value. Fix: Simple regression estimates two parameters, so df = n − 2.

Exam tips

  • Know the four assumptions by name and by the picture each violation gives in a residual plot.
  • Expect questions that give a fitted equation and ask for interpretation of the slope or intercept in the units stated.
  • On three-option items, eliminate options that reverse X and Y or attach the wrong assumption to a pattern.
  • Remember the normality assumption applies to the errors, not to X or Y.
  • For quick checks, the line passes through (X̄, Ȳ), so b̂0 = Ȳ − b̂1X̄.
  • Read which variable is the dependent one first. A swapped denominator is the most common trap in these items.
  • Questions often give Cov(X, Y), Var(X) and the means. Do two lines of arithmetic: slope, then intercept.
  • When a slope option equals Ȳ, X̄ or the product b1 × X̄, treat it as a distractor and check what it represents.