Skip to content

CFA Level II Exam · Basics of Multiple Regression and Underlying Assumptions

Dummy Variables and Predicting with Regression

Updated 7 October 2026 · Fact-checked

A dummy variable is an indicator that equals 1 when a qualitative condition is true and 0 otherwise. Its coefficient is the difference in the dependent variable versus the omitted category, holding other variables constant. To predict, substitute the given values (1 or 0 for dummies) into the estimated equation and add up.

Understand Dummy Variables and Predicting with Regression

Some factors are not numbers. Examples are whether a quarter is the fourth quarter, whether a firm is in the technology sector, or whether a country is in a crisis. A dummy variable (also called an indicator variable) turns such a factor into a number: 1 if the condition is true, 0 if not.

A dummy can shift the intercept. In the model Y = b0 + b1X + b2D + ε, the intercept is b0 when D = 0 and b0 + b2 when D = 1. So b2 is the average difference in Y between the two groups, holding X constant. The slope on X is assumed to be the same for both groups.

If a factor has n categories, use n - 1 dummies. The category left out is the base (omitted) category, and it is absorbed in the intercept. Each dummy coefficient measures the difference from that base category. Using n dummies together with an intercept creates perfect multicollinearity, called the dummy variable trap, and the regression cannot be estimated.

To test whether a group really differs, run a t-test on the dummy's coefficient. The null is that the coefficient equals zero, meaning no difference from the base category.

Predicting is plain substitution. Take the estimated equation, plug in the given values of each independent variable, and compute the predicted Y. Dummies take 1 for the category that applies and 0 for the others. Predictions are most reliable when the inputs sit inside the range of the sample data. Predicting well outside that range is risky.

Key formulas to remember

Regression with a dummy
Y = b0 + b1X1 + b2D + ε
D = 1 if the condition is true, 0 otherwise. b2 is the shift in the intercept.
Intercept by group
D = 0: intercept = b0; D = 1: intercept = b0 + b2
Slope b1 is the same in both groups unless you add an interaction term.
Number of dummies
Dummies needed = n - 1 (for n categories, with an intercept)
The omitted category is the base. Using n dummies causes perfect multicollinearity.
Predicted value
Ŷ = b0 + b1X1 + b2X2 + ... + bkXk
Substitute the given values. Use 1 or 0 for each dummy.
t-test for a coefficient
t = (bj - 0) ÷ SE(bj), with n - k - 1 degrees of freedom
k = number of independent variables. Tests whether the group difference is zero.
Interaction dummy (slope shift)
Y = b0 + b1X + b2D + b3(D × X) + ε
When D = 1 the slope is b1 + b3 and the intercept is b0 + b2.

How to solve Dummy Variables and Predicting with Regression questions

Use this method for any vignette question on dummies or forecasting from a regression.

  1. 1Read the regression output in the exhibit and note which variable is the dummy and what it equals when 1 and when 0.
  2. 2Identify the base category: the group where all dummies are 0. Interpret every dummy coefficient as the difference from that base, other variables held constant.
  3. 3If a question asks about the number of dummies, count categories and subtract one. Check whether the model includes an intercept.
  4. 4For a significance question, compute t = coefficient ÷ standard error, or read the reported t-statistic, and compare with the critical value or p-value.
  5. 5For a prediction, write the equation and substitute each given value. Set dummies to 1 for the category that applies and 0 for all others.
  6. 6Do the arithmetic in order, paying attention to the signs and units (for example percent versus decimal).
  7. 7Check the answer for reasonableness and whether the inputs lie within the sample range. Then pick the option.

Quickest way: Substitute and compare

When to use it: When the question gives a fitted equation and asks for a forecast or the gap between two groups.

  1. For a gap between groups, the answer is just the dummy coefficient (if no interaction term).
  2. For a forecast, plug values in and add term by term; do not re-derive anything.
  3. With several dummies, only one dummy per factor is 1; the rest are 0, and the base category has all zeros.
  4. Eliminate options with the wrong sign or size using a rough estimate before computing exactly.

Common mistakes in Dummy Variables and Predicting with Regression

  • Using n dummies for n categories while keeping the intercept.

    It feels natural to give every category its own variable.

    Fix: Use n - 1 dummies. The omitted category is the base and is captured by the intercept.

  • Reading a dummy coefficient as the absolute level of the group rather than a difference.

    Students forget the base category sits in the intercept.

    Fix: State it as: the group's value differs from the base by the coefficient, holding other variables constant.

  • Setting the dummy to 1 for every category in a forecast.

    Confusion when several dummies are in the equation.

    Fix: Set only the dummy for the applicable category to 1. All others are 0. The base category has all dummies 0.

  • Ignoring significance when comparing groups.

    A large coefficient looks meaningful on its own.

    Fix: Check the t-statistic or p-value. A coefficient that is not significantly different from zero gives no evidence of a group difference.

  • Mixing units, such as plugging 5 for 5% when the model uses decimals.

    Exhibits do not always state the units clearly.

    Fix: Check the variable definitions in the vignette and use the same units as in the estimation.

  • Assuming a dummy changes the slope.

    Students think the group has a different response to X.

    Fix: An intercept dummy only shifts the intercept. A slope change needs an interaction term D × X.

Worked examples

Example 1

An analyst regresses quarterly revenue growth (%) on GDP growth (%) and a dummy, Q4, equal to 1 for fourth quarters and 0 otherwise. Estimated equation: Revenue growth = 1.20 + 0.80 × GDP growth + 2.50 × Q4. The standard error of the Q4 coefficient is 1.00. Questions: (1) What is the predicted revenue growth in a fourth quarter when GDP growth is 2.0%? (2) What is the predicted growth in a non-fourth quarter with the same GDP growth? (3) Using a critical t-value of 2.0, is the Q4 effect significant?

Show the solution
  1. (1) Q4 = 1. Growth = 1.20 + 0.80 × 2.0 + 2.50 × 1 = 1.20 + 1.60 + 2.50 = 5.30%.
  2. (2) Q4 = 0. Growth = 1.20 + 1.60 + 0 = 2.80%.
  3. (3) t = 2.50 ÷ 1.00 = 2.50. Since 2.50 > 2.0, reject the null that the coefficient is zero.

Answer: (1) 5.30%; (2) 2.80%; (3) Yes, the Q4 effect is significant. Fourth quarters have growth 2.50 percentage points higher, holding GDP growth constant.

Example 2

A researcher models a stock's monthly excess return using the market excess return (MKT, %) and seasonal dummies. Each month falls in one of four seasons: Winter, Spring, Summer or Autumn. Winter is the base category. Equation: Excess return = 0.40 + 1.10 × MKT + 0.90 × Spring - 0.60 × Summer + 0.20 × Autumn. Questions: (1) How many dummies are in the model, and why? (2) What is the predicted excess return in a Summer month when MKT = 3.0%? (3) By how much does the predicted Spring return exceed the predicted Summer return for the same MKT?

Show the solution
  1. (1) Four seasons need 4 - 1 = 3 dummies (Spring, Summer, Autumn). Winter is the base, with all three dummies equal to 0.
  2. (2) Summer = 1, Spring = 0, Autumn = 0. Return = 0.40 + 1.10 × 3.0 - 0.60 = 0.40 + 3.30 - 0.60 = 3.10%.
  3. (3) Spring minus Summer = 0.90 - (-0.60) = 1.50 percentage points. The market term cancels because the slope is the same.

Answer: (1) 3 dummies, because there are four seasons and one is the base; (2) 3.10%; (3) 1.50 percentage points.

Exam tips

  • Always find the base category first. Most interpretation questions are really asking for a difference from it.
  • In a vignette, check the dummy's definition (what equals 1) before interpreting the sign of its coefficient.
  • Expect a trap option that counts n dummies instead of n - 1, or that treats the dummy coefficient as a level.
  • For forecasts, write out the full substitution line. It prevents sign and unit errors under time pressure.
  • A significance question needs the t-statistic: coefficient ÷ standard error, compared with the critical value given in the exhibit.

Dummy Variables and Predicting with Regression in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Dummy Variables and Predicting with Regression: frequently asked questions

How do I interpret a dummy variable coefficient?

It is the average difference in the dependent variable between the group coded 1 and the base group coded 0, holding the other independent variables constant. For example, a coefficient of 2.5 means the group coded 1 is 2.5 units higher. Check its t-statistic to see if the difference is significant.

What is the dummy variable trap and why use n - 1 dummies?

If you include a dummy for every one of n categories along with an intercept, the dummies add up to the intercept column. That is perfect multicollinearity and the model cannot be estimated. Using n - 1 dummies avoids it, and the omitted category becomes the base.

How do I forecast using a multiple regression equation?

Substitute the given values of the independent variables into the estimated equation and sum the terms. Use 1 for a dummy whose condition holds and 0 otherwise. The result is the predicted value of the dependent variable.

Does a dummy variable change the slope of the regression?

An ordinary dummy changes only the intercept. To let the slope differ by group, you add an interaction term, the dummy multiplied by the independent variable. Then the slope for the group coded 1 is the original slope plus the interaction coefficient.