CFA Level II Exam · Extensions of Multiple Regression
Model Misspecification in Multiple Regression
Updated 7 October 2026 · Fact-checked
Model misspecification means the regression equation does not match the true relationship in the data. Common causes are omitted variables, wrong functional form, badly transformed variables, pooled data that differ, and nonstationary time series. Find the error in the vignette, name its consequence (usually biased, inconsistent estimates), then apply the matching fix.
Understand Model Misspecification
A regression model is a claim about how variables relate. Model misspecification means that claim is wrong in some way, so the estimated coefficients and tests cannot be trusted. It is different from a plain assumption violation such as heteroskedasticity, although the two overlap.
The CFA curriculum groups specification errors into a few families. Omitted variables: a relevant variable is left out. If it is correlated with the included variables, their coefficients pick up its effect and become biased and inconsistent. Wrong functional form: you fit a straight line to a curved relationship, or leave out an interaction. Residuals then show a pattern and the errors are often heteroskedastic or serially correlated.
Variable transformation errors include not scaling variables, using the wrong transformation, or pooling data from different samples that have different relationships. Fixes include taking logs, using ratios or per-unit values, adding a squared term, adding an interaction term, or splitting the sample.
Time-series misspecification covers specification errors that commonly arise in time-series settings. These are: a lagged dependent variable used together with serially correlated errors, a function of the dependent variable used as a regressor, an independent variable measured with error, and nonstationary series (those with a unit root or trend). Regressing one trend-driven series on another is the nonstationarity problem. It can produce a spurious regression: a high R-squared and significant t-statistics with no real link. The usual fix for nonstationarity is to difference the data or test for a unit root and, if appropriate, cointegration.
A good specification is grounded in economic reasoning, parsimonious, tested out of sample, and checked for functional form. Consequences matter more than names: biased and inconsistent coefficients, unreliable standard errors, and invalid hypothesis tests.
Key formulas to remember
- Principles of a good model
- Economic reasoning + parsimony + good out-of-sample performance + appropriate functional form + no violated assumptions
- Use this as a checklist when a vignette asks whether a model is well specified.
- Omitted variable bias (direction)
- Bias in included coefficient has the sign of [corr(omitted, included) × true effect of omitted]
- Bias exists only if the omitted variable affects Y and is correlated with an included regressor. If uncorrelated, slope estimates stay unbiased.
- Log-linear (log-log) form
- ln(Y) = b0 + b1·ln(X) + ε
- Slope b1 is an elasticity: a 1% change in X is associated with about b1% change in Y.
- Quadratic term
- Y = b0 + b1·X + b2·X² + ε
- Use when residuals against X show a curve. The marginal effect of X is b1 + 2·b2·X, so it changes with X.
- Interaction term
- Y = b0 + b1·X1 + b2·X2 + b3·(X1·X2) + ε
- Marginal effect of X1 is b1 + b3·X2. Omitting a real interaction is a form error.
- Remedy for a unit root: first differences
- ΔY = b0 + b1·ΔX + ε
- Differencing removes a unit root. It is a remedy, not a test. Diagnose with unit root tests and a cointegration test (Engle-Granger). If the levels are cointegrated, the levels regression can still be valid.
How to solve Model Misspecification questions
Use the same sequence for any specification question in an item set.
- 1Read the question stem first so you know whether it asks for the error, its consequence, or the fix.
- 2Scan the vignette and exhibits for clues: a missing relevant variable, a curved residual plot, pooled samples, trending series, or a unit root test result.
- 3Name the error type: omitted variable, wrong form, transformation or pooling issue, or time-series problem.
- 4State the consequence: biased and inconsistent coefficients, invalid t-tests, or a spurious high R-squared.
- 5Check conditions. Omitted variable bias needs correlation with an included regressor. Failing to reject the unit root null suggests, but does not prove, that the series has a unit root.
- 6Pick the fix that matches the error: add the variable, change to log or add a squared or interaction term, split the sample, or difference the data.
- 7Eliminate options that fix a different problem, such as using robust standard errors for an omitted variable.
Quickest way: Match symptom to error to fix
When to use it: When time is short and the vignette gives a clear symptom.
- Missing relevant variable correlated with the others: biased slopes, add the variable.
- Curved residual pattern against a regressor: wrong form, use logs, squared or interaction terms.
- Different groups or regimes pooled together: split the sample or add dummies.
- Trending series with high R-squared and low Durbin-Watson: this pattern suggests spurious regression. Test for unit roots and cointegration, then difference if the series are not cointegrated.
- Choose the option that names both the correct error and the correct fix.
Common mistakes in Model Misspecification
Saying omitted variable bias always occurs when a variable is left out.
Students remember the rule but not its condition.
Fix: Check two things: the omitted variable affects Y, and it is correlated with an included regressor. Otherwise slopes are not biased.
Treating a high R-squared as proof the model is well specified.
High fit looks like quality.
Fix: With trending nonstationary series, a high R-squared can be spurious. Look at unit root tests and residual behaviour.
Using robust standard errors to fix an omitted variable or wrong form.
Confusing misspecification with heteroskedasticity.
Fix: Robust errors correct standard errors only. They do not remove bias from a wrong model. Fix the model itself.
Interpreting a log-log slope as a unit change.
Forgetting that both variables are in logs.
Fix: Read it as an elasticity: a percentage change in X goes with a percentage change in Y.
Differencing every time series regardless of the test.
Students memorise differencing as the cure.
Fix: Difference only when a unit root is present. If the series are cointegrated, a levels regression can be valid.
Worked examples
Example 1
An analyst regresses a bank's loan growth on GDP growth. Exhibit: the analyst omitted the policy interest rate, which affects loan growth and is negatively correlated with GDP growth. The estimated GDP slope is 1.40. (1) What is the likely direction of bias in the GDP coefficient? (2) Is the estimate consistent? (3) What is the fix?
Show the solution
- Interest rate affects loan growth negatively: true effect is negative.
- Interest rate is negatively correlated with GDP growth: correlation is negative.
- Bias sign = negative × negative = positive, so the GDP slope is likely overstated.
- Because the omitted variable is relevant and correlated with an included regressor, the bias persists in large samples, so the estimator is inconsistent.
- The fix is to add the interest rate to the model.
Answer: (1) Upward bias, GDP coefficient overstated. (2) Not consistent. (3) Include the policy rate as a regressor.
Example 2
An analyst regresses the level of a stock index on the level of a commodity price over 30 years. R-squared is 0.94, the Durbin-Watson statistic is 0.3, and unit root tests fail to reject for both series. The residual test for cointegration (null: no cointegration) also fails to reject. (1) What is the problem? (2) Is the t-statistic reliable? (3) What should the analyst do?
Show the solution
- Failing to reject the unit root null suggests, but does not prove, that both series have a unit root and are nonstationary.
- A very low Durbin-Watson statistic points to strong positive serial correlation in residuals.
- High R-squared plus low DW with apparently nonstationary series is the pattern of a spurious regression.
- The null of the cointegration test is no cointegration. Failing to reject it means there is no evidence the series are cointegrated, so the levels regression is likely spurious and its t-statistics unreliable.
- The remedy is to regress first differences, ΔY on ΔX.
Answer: (1) Likely spurious regression from apparently nonstationary data with no evidence of cointegration. (2) No, the t-statistic is unreliable. (3) Use first differences.
Exam tips
- Read the vignette for the clue word: omitted, curved residuals, pooled, trending, unit root. Each points to one error.
- Always check the conditions in the omitted variable rule before choosing biased.
- If an option says robust standard errors fix the problem, ask whether the problem is about standard errors or about the model.
- For time-series items, link high R-squared with low Durbin-Watson and unit root results to spot spurious regression.
- Read log-log coefficients as elasticities and be careful with semi-log forms.
Model Misspecification: frequently asked questions
What is omitted variable bias in CFA Level II?
It is bias in the included coefficients when a relevant variable is left out and that variable is correlated with the included regressors. The estimates are also inconsistent. If the omitted variable is uncorrelated with the regressors, the slopes are not biased.
How do I correct an incorrect functional form?
Look at the residuals against each regressor. If a curve appears, try logs, a squared term, or an interaction term. Then recheck the residuals and use out-of-sample performance to compare models.
How are nonstationary data a specification problem?
Regressing one trending or unit root series on another can give a high R-squared and significant coefficients with no real relationship. This is a spurious regression. Test for unit roots, and difference the data unless the series are cointegrated.
Is misspecification the same as heteroskedasticity or multicollinearity?
No. Those are separate assumption violations with their own fixes. Misspecification is about the model's form or variables being wrong, and it often causes those other symptoms.