Skip to content

CFA Level II Exam · Model Misspecification

Principles of Model Specification in Regression

Updated 7 October 2026 · Fact-checked

A correctly specified regression is grounded in economic reasoning, is parsimonious, performs well out of sample, has an appropriate functional form, and does not violate regression assumptions. To answer exam questions, read the vignette, check each principle in turn, and name the one that fails and its likely consequence.

Understand Principles of Model Specification

Model specification means choosing which variables go into a regression and what form they take. A model is correctly specified when it captures the real relationship between the dependent variable and its drivers, and when the regression assumptions hold well enough for the estimates and tests to be trusted.

The curriculum lists several features of a good model. It is based on sound economic reasoning, so each variable has a logical reason to be included. It is parsimonious, meaning it uses few variables to explain the data rather than piling on many. It performs well when tested on out-of-sample data, not just the data used to fit it. Its functional form is appropriate, for example variables are transformed (such as logs) when the relationship is not linear. Finally, it does not violate the regression assumptions, such as homoskedasticity, no serial correlation and no perfect multicollinearity. Perfect multicollinearity violates the assumptions and prevents the model from being estimated. High (but imperfect) multicollinearity leaves the estimates unbiased but inflates standard errors, which can produce insignificant t-statistics despite a high R-squared.

The main specification failures fall into groups. Functional form errors include omitting an important variable, failing to transform a variable, using the wrong scale, pooling data from different samples that should be separate, and using the wrong form of the data. Another group is correlation between the error term and the independent variables. It can arise from lagged dependent variables used with serial correlation in the errors, from a function of the dependent variable used as an independent variable, from independent variables measured with error, and from other timing issues. The two groups overlap. Omitted variables, inappropriate form or scaling, and pooling of different samples can themselves cause regressor-error correlation. Whenever the independent variables are correlated with the error, the estimates are biased and inconsistent, so tests and forecasts cannot be relied on.

Why does this matter in an item set? The vignette will usually show a regression output and some analyst comments. Your job is to spot which principle is broken. An omitted variable, for instance, can bias the coefficients of the variables that remain. A missing log transform can leave a pattern in the residuals. Adding many weak variables may raise R-squared but hurt out-of-sample performance.

Think of specification as the first check, before you read t-statistics. If the model is wrong, precise-looking statistics can be misleading.

Key formulas to remember

Features of a correctly specified model
Economic reasoning + parsimony + good out-of-sample fit + appropriate functional form + no assumption violations
Use this as a checklist when a vignette asks whether a model is well specified.
Log transformation for a non-linear relationship
ln(Y) = b0 + b1X + ε or Y = b0 + b1 ln(X) + ε
Use when the relationship is curved or when the variable grows proportionally. Check the residuals for a pattern first.
Consequence of regressors correlated with the error
Cov(Xj, ε) ≠ 0 ⇒ coefficient estimates are biased and inconsistent
Hypothesis tests and confidence intervals are then unreliable. This is the most serious type of misspecification.
Adjusted R-squared
Adjusted R² = 1 − [(n − 1) ÷ (n − k − 1)] × (1 − R²)
Penalises extra variables, so it supports parsimony. R² never falls when you add a variable.

How to solve Principles of Model Specification questions

Use the same sequence for any question about whether a model is well specified or what went wrong.

  1. 1Read the question first so you know whether it asks for the flaw, the consequence, or the fix.
  2. 2Scan the vignette for the model description: variables chosen, transformations, data used, and how it was tested.
  3. 3Check economic logic: does each variable have a sensible reason to be there, and is any important driver missing?
  4. 4Check functional form: is the relationship likely curved or proportional, and was a transformation applied? Look for residual patterns, mixed samples or mismatched scales.
  5. 5Check whether any independent variable could be correlated with the error, such as a lagged dependent variable with serial correlation or a variable measured with error.
  6. 6Check parsimony and out-of-sample fit: many variables, high in-sample R-squared and weak out-of-sample results point to overfitting.
  7. 7State the flaw and its consequence in one line, then match it to the answer option. Biased and inconsistent estimates mean tests and forecasts are unreliable.

Quickest way: Four-question specification scan

When to use it: Use when time is short and the vignette contains a regression description with analyst comments.

  1. Ask: is something important missing or wrongly transformed? That is a functional form problem, and it can also cause regressor-error correlation.
  2. Ask: could any regressor be tied to the error term? That means biased and inconsistent estimates.
  3. Ask: are there too many variables for the sample? That signals overfitting and weak out-of-sample results.
  4. Ask: is the problem in the error variance or error correlation? Assumption violations such as heteroskedasticity are a separate feature of a good model, but they interact with misspecification. For example, serial correlation combined with a lagged dependent variable is a specification error.
  5. Pick the option that names the principle broken and the right consequence.

Common mistakes in Principles of Model Specification

  • Choosing the model with the highest R-squared as the best specified.

    R-squared looks like a score for the model, and it always rises when you add variables.

    Fix: Remember that a good model also needs economic logic, parsimony and out-of-sample performance. Use adjusted R-squared and out-of-sample results to judge.

  • Treating all misspecification as having the same consequence.

    Students memorise a list of errors without linking each to its effect.

    Fix: Link each error to its effect. Any misspecification that makes a regressor correlated with the error, including omitted variables, wrong form or scale, and pooling, gives biased and inconsistent estimates, so you cannot trust tests.

  • Confusing misspecification with assumption violations such as heteroskedasticity.

    Both are described as problems with the regression, and the chapter covers them close together.

    Fix: Specification is mainly about variables and form. Heteroskedasticity and serial correlation are about error behaviour and are a separate feature of a good model, but they can interact with misspecification, as with serial correlation and a lagged dependent variable. The curriculum lists absence of assumption violations as one feature of a good model.

  • Assuming more variables always improve the model.

    Adding variables lifts in-sample fit, which feels like progress.

    Fix: Look for economic justification for each variable. Extra weak variables risk overfitting and poor out-of-sample results.

  • Ignoring the data when judging a model.

    Candidates focus on the variables and miss comments about pooled samples or mismatched data.

    Fix: Check whether data from different regimes or populations were combined, since this can distort the estimated relationship.

Worked examples

Example 1

An analyst regresses a stock's monthly return on 14 candidate variables using 40 observations. The in-sample R-squared is very high. When the model is used on the following 12 months of data, forecasts are poor. Several variables have no clear economic link to returns. (1) Which principle of specification is most clearly violated? (2) What is the most likely explanation for the poor forecasts?

Show the solution
  1. Count: 14 variables on 40 observations is a large number of regressors for the sample size.
  2. The vignette says several variables lack an economic link. That breaks the economic reasoning principle and also parsimony.
  3. High in-sample fit with poor out-of-sample forecasts is the signature of overfitting.

Answer: (1) Parsimony and economic reasoning are violated. (2) The model is overfitted to the sample, so it performs badly out of sample.

Example 2

A researcher estimates a time-series regression of a firm's sales on its lagged sales and other variables. The residuals are serially correlated. The analyst notes the coefficient on lagged sales is statistically significant and says the estimate can be trusted. (1) Is the analyst correct? (2) What is the consequence for the estimates?

Show the solution
  1. Lagged sales is a lagged dependent variable used as an independent variable.
  2. Lagged sales depends on the previous period's error. If the errors are serially correlated, the previous error is correlated with the current error, so lagged sales is correlated with the current error term.
  3. When a regressor is correlated with the error, the estimates are biased and inconsistent.
  4. Significance tests rest on those estimates, so a significant t-statistic does not make the estimate reliable.

Answer: (1) No. (2) The coefficient estimates are biased and inconsistent, so the significance test cannot be trusted.

Exam tips

  • Read each analyst statement in the vignette as a possible specification claim and test it against the checklist.
  • Do not choose an option just because it cites a higher R-squared; look for out-of-sample or economic support.
  • Know which failures make estimates biased and inconsistent, since options often hinge on that wording.
  • Keep specification errors and error-assumption violations distinct, but remember they can interact; options often blur them deliberately.
  • Remember that perfect multicollinearity violates the assumption and prevents estimation. High multicollinearity leaves estimates unbiased but inflates standard errors, so t-statistics can be insignificant despite a high R-squared.
  • If a question asks for a fix, match it: transform for curvature, drop or justify weak variables for parsimony, and remove the correlation between regressor and error.

Principles of Model Specification: frequently asked questions

What is a correctly specified regression model?

It is a model built on sound economic reasoning that is parsimonious, performs well out of sample, uses an appropriate functional form and does not violate the regression assumptions. Such a model gives estimates and tests you can rely on.

Why is parsimony important in model specification?

A parsimonious model uses few variables to explain the data. Too many variables can fit noise in the sample, which raises in-sample fit but hurts forecasts on new data.

What happens when an independent variable is correlated with the error term?

The regression coefficient estimates become biased and inconsistent. As a result, hypothesis tests and confidence intervals are unreliable, and forecasts may be misleading.

How do I avoid model misspecification in regression?

Choose variables with a sound economic rationale, apply transformations where the relationship is non-linear, avoid mixing unlike data samples, and keep the model simple. Test it out of sample and check the residuals for assumption violations.