Skip to content

CFA Level II Exam · Evaluating Regression Model Fit and Interpreting Model Results

Dummy Variables and Model Misspecification in Regression

Updated 7 October 2026 · Fact-checked

A dummy variable is an independent variable coded 1 or 0 to show a category. With n categories you use n − 1 dummies. Each coefficient is the difference from the omitted category. Misspecification means the model form is wrong, such as omitted variables or wrong functional form. Logistic regression models a 0/1 outcome.

Understand Dummy Variables and Misspecification

A dummy variable (indicator variable) lets you put qualitative information into a regression. It takes the value 1 if an observation has a feature and 0 if it does not. Examples: a stock is in the technology sector, a month is January, a bond is rated investment grade.

If you have n categories, use n − 1 dummies. The omitted category is the base (reference) group, and it is captured by the intercept. Using n dummies together with an intercept creates perfect multicollinearity, called the dummy variable trap.

How to read the coefficients: an intercept dummy shifts the intercept. In Y = b0 + b1D + b2X, b0 is the intercept for the base group and b0 + b1 is the intercept for the D = 1 group. The slope on X is the same for both. Each dummy coefficient is the average difference in Y from the base group, holding other variables constant. Test it with a t-test like any other coefficient. A slope dummy (interaction term) uses D × X, so the slope on X differs by group. Slope for D = 1 is b2 + b3 in Y = b0 + b1D + b2X + b3(D × X).

Model misspecification means the regression is set up wrongly, so coefficients may be biased or inconsistent. Main types: (1) omitting a variable, (2) wrong functional form, such as not transforming a variable, pooling data from different samples, or mis-scaled variables, (3) variables correlated with the error term, such as including a lagged dependent variable with serially correlated errors, measurement error in the independent variable, or a forecasting variable that uses the dependent variable, (4) other time-series misspecification, such as nonstationarity. Misspecification is more serious than heteroskedasticity or serial correlation alone, because it can bias the coefficients themselves.

Logistic (logit) regression is used when the dependent variable is qualitative, such as default or no default (coded 0/1). Linear regression can give fitted values below 0 or above 1, which makes no sense for a probability. Logit transforms the probability so it stays between 0 and 1. It models the log odds, ln[P ÷ (1 − P)], as a linear function of the independent variables. It is estimated by maximum likelihood, not least squares. Slope coefficients are changes in log odds, not changes in probability. Probit uses the normal distribution instead of the logistic one. Discriminant analysis produces a score used to classify observations into groups.

Key formulas to remember

Number of dummies
Dummies needed = n − 1 (n = number of categories)
Using n dummies plus an intercept causes perfect multicollinearity.
Intercept dummy
Y = b0 + b1D + b2X + ε
Base group intercept = b0. D = 1 group intercept = b0 + b1. Slope on X is the same.
Slope dummy (interaction)
Y = b0 + b1D + b2X + b3(D × X) + ε
Slope when D = 0 is b2. Slope when D = 1 is b2 + b3.
Logit model
ln[P ÷ (1 − P)] = b0 + b1X1 + … + bkXk
Left side is the log odds. Estimated by maximum likelihood.
Probability from log odds
P = 1 ÷ (1 + e^(−z)), where z = b0 + b1X1 + …
Gives a probability between 0 and 1. Odds = P ÷ (1 − P) = e^z.

How to solve Dummy Variables and Misspecification questions

Use this method for any item-set question on dummy variables, misspecification or logit.

  1. 1Read the vignette and list the dependent variable, each independent variable and how each is coded. Note which are 0/1 dummies.
  2. 2Identify the base group: the category with every dummy equal to 0. Its value is captured by the intercept.
  3. 3For a dummy coefficient, state it as the difference from the base group, holding other variables constant. Check the t-statistic or p-value in the exhibit for significance.
  4. 4For predictions, substitute 1 or 0 for each dummy and the given values for other variables. For slope dummies, include the interaction term.
  5. 5For misspecification, match the symptom to the type: missing variable, wrong form (e.g. needs a log), pooled samples, or regressor correlated with the error.
  6. 6For logit, decide whether the exhibit gives log odds. Convert to odds with e^z, then to probability with P = odds ÷ (1 + odds).
  7. 7Check that the answer makes sense: probabilities are between 0 and 1, and the sign matches the story.

Quickest way: Plug in 0 and 1, then compare

When to use it: When the exhibit gives a regression table with a dummy and you must interpret or predict.

  1. Write the equation with the coefficients from the exhibit.
  2. Set the dummy to 0 to get the base group, then to 1 to get the other group.
  3. The difference between the two predictions is the dummy coefficient (plus any interaction term times X).
  4. For misspecification questions, eliminate options that describe a different problem, such as heteroskedasticity when the issue is a missing variable.

Common mistakes in Dummy Variables and Misspecification

  • Using n dummies for n categories along with an intercept.

    It feels natural to give every category its own variable.

    Fix: Use n − 1 dummies. The omitted group is the base and sits in the intercept.

  • Interpreting a dummy coefficient as the group's average level.

    Students forget the intercept carries the base group.

    Fix: The coefficient is the difference from the base group. The D = 1 group's intercept is b0 + b1.

  • Ignoring the interaction term when predicting for the D = 1 group.

    The slope dummy is easy to overlook in a long table.

    Fix: When D = 1, the slope on X is b2 + b3, not b2.

  • Reading logit coefficients as changes in probability.

    Linear regression habits carry over.

    Fix: Logit slopes are changes in log odds. Compute z, then convert to probability.

  • Treating misspecification as the same as heteroskedasticity or serial correlation.

    All are called model problems.

    Fix: Misspecification (omitted variable, wrong form, regressor correlated with error) can bias coefficients. The others mainly affect standard errors.

Worked examples

Example 1

An analyst regresses monthly return (%) on market return (X) and a January dummy (D = 1 in January, 0 otherwise): Return = 0.40 + 1.50D + 0.90X. (1) What is the expected return in January if the market return is 2%? (2) What is it in other months for the same market return? (3) How is the 1.50 interpreted?

Show the solution
  1. January: 0.40 + 1.50(1) + 0.90(2) = 0.40 + 1.50 + 1.80 = 3.70.
  2. Other months: 0.40 + 0 + 1.80 = 2.20.
  3. The difference is 3.70 − 2.20 = 1.50, which equals the dummy coefficient.

Answer: (1) 3.70%. (2) 2.20%. (3) January returns are on average 1.50 percentage points higher than in other months, holding market return constant.

Example 2

A logit model estimates default: ln[P ÷ (1 − P)] = −3.00 + 0.50 × Leverage, where leverage is debt to equity. A firm has leverage of 4.0. (1) What is the log odds of default? (2) What is the probability of default? (3) Why is logit preferred to linear regression here?

Show the solution
  1. Log odds z = −3.00 + 0.50 × 4.0 = −1.00.
  2. Odds = e^(−1) = 0.3679.
  3. P = odds ÷ (1 + odds) = 0.3679 ÷ 1.3679 = 0.269.
  4. Linear regression on a 0/1 outcome can give fitted values outside 0 to 1. Logit keeps probabilities between 0 and 1.

Answer: (1) −1.00. (2) About 26.9%. (3) Logit constrains the predicted probability to between 0 and 1 and suits a qualitative dependent variable.

Exam tips

  • Always find the base group first. Most dummy questions turn on this.
  • If a vignette says the slope differs between groups, look for an interaction term D × X.
  • For misspecification, name the cause from the vignette clue: a left-out variable, a curved relationship fitted with a straight line, or data from different regimes pooled together.
  • For logit, remember the output is log odds. Convert it to a probability before answering if asked.
  • There is no penalty for wrong answers, so never leave an item blank.

Dummy Variables and Misspecification in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Dummy Variables and Misspecification: frequently asked questions

How do you interpret a dummy variable coefficient?

It is the average difference in the dependent variable between the group coded 1 and the base group coded 0, holding other variables constant. Check its t-statistic to see if the difference is statistically significant.

Why use n − 1 dummy variables?

With n dummies and an intercept, the dummies add up to the intercept column. This is perfect multicollinearity and the regression cannot be estimated. Dropping one gives a base group.

What are the main types of model misspecification?

Omitted variables, incorrect functional form, regressors correlated with the error term, and time-series problems such as nonstationarity. These can bias coefficients and make results unreliable.

How does logistic regression differ from linear regression?

Logistic regression has a 0/1 dependent variable and models the log odds of the outcome. It is estimated by maximum likelihood and gives probabilities between 0 and 1. Linear regression uses a continuous dependent variable and least squares.