Skip to content

CFA Level I Exam · Applications of Simple Linear Regression in Finance

Simple Linear Regression Model and Its Assumptions for CFA Level 1

Updated 7 October 2026 · Fact-checked

Simple linear regression models a dependent variable Y as a straight-line function of one independent variable X plus an error term: Y = b0 + b1X + ε. The slope is the expected change in Y per one-unit change in X. Key assumptions: linearity, homoskedasticity, independence of errors, and normally distributed errors.

Understand Simple Linear Regression Model and Assumptions

Simple linear regression describes how one variable changes with another. The dependent variable (Y, also called the explained or predicted variable) is the one you want to explain. The independent variable (X, also called the explanatory or predictor variable) is the one you use to explain it. Example: explain a stock's return (Y) using the market's return (X).

The population model is Y = b0 + b1X + ε. The intercept b0 is the expected value of Y when X is zero. The slope b1 is the expected change in Y for a one-unit change in X. The error term ε captures everything that X does not explain. For each observation i, the error is the gap between the actual Yi and the value the line gives.

The regression gives a line of best fit, estimated by ordinary least squares, which minimizes the sum of squared errors. The estimated line is Ŷ = b̂0 + b̂1X. The fitted values Ŷ are predictions. The differences Y − Ŷ are residuals.

For the estimates and the tests on them to be reliable, the exam expects four assumptions. Linearity: the relationship between X and Y is linear in the parameters. Homoskedasticity: the variance of the errors is the same for all observations. Independence: the observations are independent of each other, so the errors are not correlated (no serial correlation). Normality: the errors are normally distributed, with an expected value of zero. Also, X must not be a constant, and X is assumed to be uncorrelated with the error.

When an assumption fails, the line may still be drawn, but conclusions suffer. Heteroskedasticity (non-constant error variance) makes standard errors unreliable, so t-tests and confidence intervals mislead. Correlated errors do the same. A curved relationship makes a straight line a poor fit. You can spot these problems in a scatter plot of the data or a plot of residuals against X or over time.

Key formulas to remember

Population regression model
Yi = b0 + b1 Xi + εi
Y is dependent, X is independent, ε is the error term. i = 1, ..., n.
Estimated regression line
Ŷi = b̂0 + b̂1 Xi
Ŷ is the predicted value from the fitted line.
Residual
ei = Yi − Ŷi
OLS chooses b̂0 and b̂1 to minimize Σ ei².
Slope estimate
b̂1 = Cov(X, Y) ÷ Var(X)
Sample covariance and variance must use the same denominator (n − 1).
Intercept estimate
b̂0 = Ȳ − b̂1 X̄
The fitted line always passes through the point (X̄, Ȳ).
Core assumptions
Linearity; homoskedasticity; independence of errors; normality of errors
Errors have an expected value of zero. X is not constant and is uncorrelated with ε.

How to solve Simple Linear Regression Model and Assumptions questions

Use this routine for any question on the model, its terms or its assumptions.

  1. 1Identify Y and X. Y is what is being explained or predicted; X is what is used to explain it.
  2. 2Write the equation in the form Y = b0 + b1X + ε and match each given number to b0 or b1.
  3. 3Interpret the slope as the expected change in Y per one-unit change in X, using the units in the question.
  4. 4Interpret the intercept as the expected Y when X = 0, and check whether X = 0 is meaningful.
  5. 5If asked for a prediction, substitute X into the fitted line and compute Ŷ. If asked for the residual, subtract Ŷ from actual Y.
  6. 6If the question describes a pattern, match it to an assumption: curved pattern means linearity; widening spread means homoskedasticity fails; errors in a time series tied to the previous one means independence fails; skewed or fat-tailed errors means normality fails.
  7. 7State the consequence of the violation: unreliable standard errors and tests, or a poor fit.
  8. 8Pick the option that fits and eliminate the two others by checking units, direction and the wrong assumption.

Quickest way: Label, interpret, match

When to use it: Use under time pressure on standalone three-option questions about definitions, interpretation or assumption violations.

  1. Underline Y and X in the stem. The variable being explained is Y.
  2. Read the slope as 'Y changes by b1 units per 1 unit of X'. Check the sign and units.
  3. For an assumption question, name the pattern: spread changes means heteroskedasticity; errors linked over time means serial correlation; curve means nonlinearity.
  4. Remove any option that reverses X and Y or misstates the assumption, then choose the remaining one.

Common mistakes in Simple Linear Regression Model and Assumptions

  • Swapping dependent and independent variables.

    Students assume the first variable named in the stem is Y.

    Fix: Ask which variable is being explained or predicted. That is Y, whatever the order in the sentence.

  • Interpreting the intercept as always meaningful.

    The intercept is defined at X = 0, which may be outside the data range or impossible.

    Fix: State it as the expected Y when X = 0, and treat it as an extrapolation if X = 0 is not realistic.

  • Confusing homoskedasticity with heteroskedasticity.

    The names are similar and long.

    Fix: Homo means same: constant error variance. Hetero means different: error variance changes with X.

  • Saying the slope shows that X causes Y.

    A fitted line looks like a cause-and-effect statement.

    Fix: The slope describes an association in the data. Regression alone does not prove causation.

  • Thinking the model assumes X or Y is normally distributed.

    Students remember 'normality' without the object.

    Fix: The assumption is about the errors, not the variables themselves.

  • Using the wrong sign or direction for the residual.

    Students compute predicted minus actual.

    Fix: Residual = actual Y − predicted Ŷ. A positive residual means the line under-predicted.

Worked examples

Example 1

An analyst regresses a fund's monthly return (Y, in %) on the market's monthly return (X, in %) and gets Ŷ = 0.40 + 1.20X. The market returns 5% in a month. Which is the predicted fund return? A) 5.0%, B) 6.4%, C) 7.2%

Show the solution
  1. Y is the fund return and X is the market return.
  2. Substitute X = 5: Ŷ = 0.40 + 1.20 × 5.
  3. 1.20 × 5 = 6.00, and 6.00 + 0.40 = 6.40.
  4. Check: 5.0% ignores the slope and intercept; 7.2% would result from 1.20 × 6, which is not the given X.

Answer: B) 6.4%. The slope of 1.20 also means the fund is expected to move 1.2 percentage points per 1 point of market return.

Example 2

A residual plot from a regression of company expenses on revenue shows that the spread of residuals gets wider as revenue increases. Which assumption is most likely violated, and what is the main consequence? A) Independence; the slope estimate is biased upward, B) Homoskedasticity; the standard errors are unreliable, C) Normality; the intercept cannot be estimated

Show the solution
  1. Widening spread of residuals means the error variance is not constant.
  2. Constant error variance is homoskedasticity, so it is violated: this is heteroskedasticity.
  3. The effect is that standard errors are unreliable, so t-tests and confidence intervals on the coefficients can mislead.
  4. Option A names independence, which concerns correlation between errors, not their spread. Option C names normality and claims something false about the intercept.

Answer: B) Homoskedasticity is violated, so standard errors and tests on the coefficients are unreliable.

Exam tips

  • Know the four assumptions by name and by the picture each violation gives in a residual plot.
  • Expect questions that give a fitted equation and ask for interpretation of the slope or intercept in the units stated.
  • On three-option items, eliminate options that reverse X and Y or attach the wrong assumption to a pattern.
  • Remember the normality assumption applies to the errors, not to X or Y.
  • For quick checks, the line passes through (X̄, Ȳ), so b̂0 = Ȳ − b̂1X̄.

Practice questions from Applications of Simple Linear Regression in Finance

Simple Linear Regression Model and Assumptions in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Simple Linear Regression Model and Assumptions: frequently asked questions

What are the assumptions of simple linear regression for CFA Level I?

The main ones are a linear relationship, constant variance of errors (homoskedasticity), independent errors, and normally distributed errors. The error term is also assumed to have an expected value of zero. X should not be constant and should be uncorrelated with the error.

What is the dependent and independent variable in regression?

The dependent variable (Y) is the one you are trying to explain or predict. The independent variable (X) is the one used to explain it. In a regression of stock return on market return, the stock return is dependent.

How do I interpret the slope and intercept?

The slope is the expected change in Y for a one-unit increase in X. The intercept is the expected value of Y when X equals zero. Always use the units given in the question, and be careful if X = 0 is outside the data range.

What happens if homoskedasticity is violated?

The error variance changes across observations. The coefficient estimates can still be computed, but the standard errors are unreliable. That makes hypothesis tests and confidence intervals on the coefficients misleading.