CFA Level I Exam · Applications of Simple Linear Regression in Finance
Prediction Using Regression and Prediction Intervals
Updated 7 October 2026 · Fact-checked
A regression forecast plugs a given X into Ŷ = b0 + b1X. A prediction interval puts a range around that forecast: Ŷ ± t × s_f, using n − 2 degrees of freedom. The standard error of forecast, s_f, is larger than the standard error of estimate and grows as X moves farther from the mean of X.
Understand Prediction Using Regression and Prediction Intervals
A simple linear regression gives you a line: Ŷ = b0 + b1X. If you are given a value of the independent variable X, you put it into the line and read off the predicted value of the dependent variable Y. This point forecast is your best single guess, but the actual outcome will almost never equal it.
The gap between a real Y and the line comes from the error term. The standard error of estimate (s_e) measures the typical size of that gap in the sample. A prediction interval uses s_e, and adds the uncertainty from having estimated b0 and b1 from a limited sample. The result is a range in which a single new observation of Y is expected to fall, at a stated confidence level.
The measure of that total uncertainty is the standard error of the forecast (s_f). It is always larger than s_e. It has three parts: the residual noise (the 1), the uncertainty about the intercept (the 1/n term), and the uncertainty about the slope, which depends on how far your X is from the sample mean of X.
That last part gives the key intuition. The line is best pinned down near X̄. Forecasts far from X̄ are extrapolations, and the interval widens. A larger sample (bigger n), a larger spread in X, and a smaller s_e all give a narrower interval.
Do not confuse this with a confidence interval for the mean response, which estimates the average Y at a given X and leaves out the 1 inside the bracket. A prediction interval is for one individual outcome, so it is always wider.
Key formulas to remember
- Predicted value
- Ŷ = b0 + b1 × X_f
- X_f is the given value of the independent variable. Use unrounded b0 and b1 if given.
- Variance of the forecast
- s_f² = s_e² × [1 + 1/n + (X_f − X̄)² ÷ ((n − 1) × s_x²)]
- s_x² is the sample variance of X. (n − 1) × s_x² equals Σ(Xi − X̄)².
- Standard error of forecast
- s_f = √(s_f²)
- Always greater than s_e. Smallest when X_f = X̄.
- Prediction interval
- Ŷ ± t(α/2, n − 2) × s_f
- Two-tailed critical t with n − 2 degrees of freedom. Simple regression loses two degrees of freedom.
- Standard error of estimate
- s_e = √(SSE ÷ (n − 2))
- Often given in the question or taken from the ANOVA table as √MSE.
How to solve Prediction Using Regression and Prediction Intervals questions
Use this order for any question on forecasting or prediction intervals in simple linear regression.
- 1Write the regression equation and the given X_f. Compute Ŷ = b0 + b1 × X_f.
- 2List what is given: n, s_e, X̄, and either s_x² or Σ(Xi − X̄)². Check whether the variance of X or the sum of squares is given, so you use the right denominator.
- 3Compute the bracket: 1 + 1/n + (X_f − X̄)² ÷ ((n − 1) × s_x²).
- 4Multiply the bracket by s_e² and take the square root to get s_f. Check that s_f is larger than s_e.
- 5Find the critical t value for n − 2 degrees of freedom and the stated confidence level (two-tailed).
- 6Compute the interval Ŷ ± t × s_f. State the lower and upper limits in the units of Y.
- 7Sanity-check: the interval must be centred on Ŷ, and Ŷ must lie inside it.
Quickest way: Shortcut: compare, don't compute
When to use it: Use when the question asks which forecast is more precise, how the interval changes, or which option is plausible, rather than asking for a full calculation.
- Compute Ŷ first. Many options can be eliminated by the centre of the interval alone.
- Remember s_f is always greater than s_e. Any option with s_f below s_e is wrong.
- If X_f equals X̄, the last term is zero and s_f = s_e × √(1 + 1/n). That is the minimum s_f.
- If the question changes X_f, only the (X_f − X̄)² term changes. Farther from X̄ means a wider interval.
- For numbers, use the calculator: BA II Plus, enter the bracket value × s_e² then 2nd √x. HP 12C, enter the value then g √x.
Common mistakes in Prediction Using Regression and Prediction Intervals
Using s_e directly as the margin of error instead of s_f.
s_e is given in the question and looks like the standard error to use.
Fix: s_e only measures residual noise. The interval needs s_f, which adds the 1/n and (X_f − X̄)² terms.
Using n − 1 degrees of freedom for the t value.
Students carry over the rule from a confidence interval for a mean.
Fix: Simple regression estimates two parameters, so df = n − 2.
Dropping the '1' from the bracket and so producing the interval for the mean response.
The two formulas look alike and the mean-response version is simpler.
Fix: A prediction interval for a single new observation always includes the 1. Without it you get a narrower interval for the average Y.
Mixing up s_x² with Σ(Xi − X̄)².
The denominator is (n − 1) × s_x², and the question may give either quantity.
Fix: If you are given the sample variance of X, multiply by (n − 1). If you are given the sum of squared deviations, use it as it is.
Forgetting to square (X_f − X̄) or to square s_e before taking the root.
Rushing under time pressure, especially when working with a calculator.
Fix: Work in variances: s_e² × bracket, then one square root at the end.
Believing a forecast far outside the sample range is just as reliable.
The regression line extends indefinitely, so any X can be plugged in.
Fix: The (X_f − X̄)² term widens the interval away from X̄, and the linear relationship may not hold outside the sample range.
Worked examples
Example 1
An analyst regresses a fund's monthly return (Y, %) on a benchmark's monthly return (X, %) using n = 32 observations. The estimated equation is Ŷ = 1.0 + 0.8X. The standard error of estimate is 2.0, the mean of X is 3.0, and the sample variance of X is 4.0. The critical t value for 30 degrees of freedom at 95% confidence (two-tailed) is 2.042. Construct a 95% prediction interval for the fund's return if the benchmark returns 5%.
Show the solution
- Point forecast: Ŷ = 1.0 + 0.8 × 5 = 5.0%.
- Deviation term: (5 − 3)² = 4. Denominator: (n − 1) × s_x² = 31 × 4 = 124. Ratio = 4 ÷ 124 = 0.03226.
- Bracket: 1 + 1/32 + 0.03226 = 1 + 0.03125 + 0.03226 = 1.06351.
- s_f² = 2.0² × 1.06351 = 4 × 1.06351 = 4.25403. s_f = √4.25403 = 2.0625.
- Margin = 2.042 × 2.0625 = 4.2116.
- Interval = 5.0 ± 4.2116 = 0.79% to 9.21% (rounded).
Answer: The 95% prediction interval is approximately 0.79% to 9.21%.
Example 2
Using the same regression (n = 32, s_e = 2.0, X̄ = 3.0), what is the standard error of the forecast when the benchmark return is X_f = 3.0%? A. 2.00 B. 2.03 C. 2.06
Show the solution
- X_f equals X̄, so (X_f − X̄)² = 0 and the slope-uncertainty term drops out.
- Bracket: 1 + 1/32 = 1.03125.
- s_f² = 4 × 1.03125 = 4.125.
- s_f = √4.125 = 2.031. Calculator: BA II Plus, 4.125 then 2nd √x. HP 12C, 4.125 then g √x.
- Eliminate A: 2.00 is just s_e, and s_f must exceed s_e. Eliminate C: 2.06 is the value at X_f = 5, a point farther from the mean, so it is too large for X_f = X̄.
Answer: B. 2.03
Exam tips
- Questions are three-option MCQs. You can often eliminate two options by checking that s_f > s_e and that the interval is centred on Ŷ.
- Read the question for whether it gives s_x² or Σ(Xi − X̄)². This decides the denominator and is a common trap.
- Conceptual items ask what makes the interval wider or narrower: a larger s_e, smaller n, smaller spread of X, or X_f farther from X̄ all widen it.
- Use df = n − 2 and a two-tailed critical value that matches the confidence level. A 95% interval uses the 2.5% upper tail.
- With about 90 seconds per question, do the conceptual comparison first and calculate only if the options cannot be separated otherwise.
Practice questions from Applications of Simple Linear Regression in Finance
- In a simple linear regression of a stock's monthly excess returns on the market's monthly excess returns, the independent variable is most l…
- In a simple linear regression, the standard error of estimate is most likely:
- In a simple linear regression with 27 observations, the estimated slope is 0.80 and its standard error is 0.30. The two-tailed critical t-va…
- A regression of a company's earnings growth on GDP growth was estimated using GDP growth between 1% and 4%. An analyst uses it to forecast e…
- Which statement about the ordinary least squares (OLS) estimation of a simple linear regression is most accurate?
Prediction Using Regression and Prediction Intervals in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Prediction Using Regression and Prediction Intervals: frequently asked questions
What is the difference between a confidence interval and a prediction interval in regression?
A confidence interval for the mean response estimates the average Y at a given X. A prediction interval estimates where one new observation of Y will fall. The prediction interval adds the extra 1 inside the bracket, so it is always wider.
What is the standard error of forecast formula in CFA Level I?
s_f² = s_e² × [1 + 1/n + (X_f − X̄)² ÷ ((n − 1) × s_x²)]. Take the square root to get s_f. Then the prediction interval is Ŷ ± t × s_f with n − 2 degrees of freedom.
Why does the prediction interval get wider far from the mean of X?
The slope is estimated with error, and that error has a bigger effect on Ŷ the farther X_f is from X̄. The (X_f − X̄)² term captures this. At X_f = X̄ the interval is at its narrowest.
How do I forecast the dependent variable from a regression equation?
Substitute the given X value into Ŷ = b0 + b1X. Use the intercept and slope exactly as given. This gives the point forecast, which is the centre of any prediction interval.