CFA Level II Exam · Model Misspecification
Misspecified Functional Form in Regression Models
Updated 7 October 2026 · Fact-checked
Misspecified functional form means the regression equation does not match the true relationship. Causes are omitted variables, variables that need a transformation, pooling data from different regimes, and incorrect scaling. These errors can bias coefficients and cause heteroskedasticity or serial correlation. In the exam, name the cause, then state its effect on estimates and tests.
Understand Misspecified Functional Form
A regression model is a claim about how variables relate. Functional form is the shape of that claim: which variables are in, how they enter (level, log, ratio, squared), and whether one equation fits all observations. If the shape is wrong, the model is misspecified.
The damage is usually worse than a poor fit. Misspecification in general can lead to biased and inconsistent coefficient estimates, so more data does not fix it. It may also show up as residual patterns such as heteroskedasticity or serial correlation. The standard errors, t-statistics and forecasts then cannot be trusted.
The curriculum groups the problems into four types:
- Omitted variables. An important independent variable is left out. If it is correlated with an included variable, the included variable's coefficient picks up part of the omitted variable's effect. This is omitted variable bias. The error term now contains the missing variable, so it is correlated with the regressor. If the omitted variable is uncorrelated with the included ones, the slope estimates are not biased by it, but the model still fits worse.
- Variable should be transformed. The relationship is not linear in the raw variable. Examples: a log is needed for a variable growing exponentially, or a squared or other nonlinear term for a curved relationship. Using the raw variable gives a misfit, and the residuals may show patterns such as heteroskedasticity.
- Inappropriate pooling of data. You combine samples from different regimes or groups, such as before and after a policy change, or two sectors with different relationships. Each subsample may have its own slope. The pooled regression gives an average that may describe neither. The residuals may also show heteroskedasticity or serial correlation.
- Inappropriate scaling. Variables are in different units or magnitudes, for example total revenue for firms of very different size, when the relationship needs common-size or per-unit data. Differences in size can leave residual variance that differs across observations, which is heteroskedasticity. Fix it by using ratios, common-size statements or per-share values.
The exam gives you a vignette describing an analyst's model and asks you to identify which error it contains and what follows. Match the description to the type first. Then recall the consequence.
Key formulas to remember
- Omitted variable bias condition
- Bias in b₁ ≠ 0 when the omitted variable affects Y and is correlated with X₁
- Both conditions are needed. If the omitted variable is uncorrelated with the included regressors, no bias in their slopes arises from it.
- Direction of bias (general econometric rule of thumb)
- Sign of bias = sign(effect of omitted variable on Y) × sign(correlation of omitted variable with X₁)
- This is a general econometric rule of thumb for one omitted variable, not a formula from the CFA curriculum. Same signs overstate the included coefficient; opposite signs understate it. It does not carry over cleanly when several variables are omitted.
- Four types of functional form error
- Omitted variable | Wrong transformation | Inappropriate pooling | Inappropriate scaling
- Learn these as a checklist. Most questions ask you to name one.
- Typical consequence
- Misspecification → biased, inconsistent coefficients; unreliable standard errors and tests
- Adding more observations does not cure it.
How to solve Misspecified Functional Form questions
Use this method on any item set question about a model that looks wrong.
- 1Read the vignette for how the model was built: which variables, what units, what sample period or groups.
- 2Find the clue. A missing driver suggests omission. A curved or exponential pattern suggests transformation. Mixed periods or groups suggest pooling. Mixed magnitudes or units suggest scaling.
- 3Name the error using the curriculum term.
- 4State the effect: biased and inconsistent coefficients for omission, poor fit and often heteroskedasticity for wrong form, an averaged and misleading relationship for pooling.
- 5For omission, check whether the omitted variable is correlated with an included regressor. If not, bias in slopes is not expected.
- 6Choose the fix: add the variable, transform it (log or square), split the sample or run separate regressions, or use ratios, common-size and per-unit data.
- 7Check the answer options for the one that matches both the diagnosis and the consequence.
Quickest way: Clue-to-error matching
When to use it: When time is short and the question asks which misspecification is present or what it causes.
- Underline the one phrase in the vignette that describes the data or model choice.
- Map it: left out variable = omission; growth or curve = transformation; different periods or groups combined = pooling; mixed units or sizes = scaling.
- Pick the option that names that error and its consequence (bias, heteroskedasticity, or misleading average).
- Eliminate options that blame multicollinearity or outliers unless the vignette says so.
Common mistakes in Misspecified Functional Form
Saying any omitted variable causes bias.
Students remember omission equals bias and skip the condition.
Fix: Check that the omitted variable affects Y and is correlated with an included regressor. If not, slope bias from it is not expected.
Confusing pooling with omission.
Both can produce a poor fit and similar symptoms.
Fix: Pooling is about combining different regimes or groups in one equation. Omission is about a missing explanatory variable.
Confusing scaling with transformation.
Both involve changing how a variable enters the model.
Fix: Scaling concerns units or size differences, fixed by ratios, per-share or common-size data. Transformation concerns a nonlinear relationship, fixed by logs or squared or other nonlinear terms.
Believing a larger sample cures misspecification.
Students carry over the idea that more data reduces error.
Fix: Misspecification makes the estimators inconsistent, so the bias persists as the sample grows. The model itself must change.
Treating a high R² as proof the form is right.
Fit statistics feel like a final check.
Fix: A model can fit well and still omit a variable or pool regimes. Judge the specification from the design and residual patterns.
Getting the direction of omitted variable bias wrong.
Students multiply the wrong signs.
Fix: Use the product of the omitted variable's effect on Y and its correlation with the included X. Positive product overstates; negative product understates.
Worked examples
Example 1
An analyst regresses monthly returns of 40 mid-cap equity funds on their expense ratios. Fund manager skill is not in the model. Better managers tend to charge higher fees and also produce higher returns. The expense ratio coefficient comes out positive. (1) Which functional form error is present? (2) Is the coefficient biased, and in which direction? (3) Does adding more funds fix it?
Show the solution
- Manager skill is a driver of returns that is missing, so this is an omitted variable.
- Skill affects returns positively and is positively correlated with expense ratio.
- The product of the signs is positive, so the expense ratio coefficient is overstated (biased upward).
- The error is inconsistent, so more funds with the same specification do not remove it.
Answer: (1) Omitted variable. (2) Yes, biased upward. (3) No, the model must include a skill proxy.
Example 2
An analyst fits one regression of sales growth on advertising spend using 20 years of data. Midway through, the firm switched from a domestic to a global business model, and the advertising effect changed. Separately, the advertising variable is entered in raw rupees for a small division and a very large division together, without scaling. (1) Name both errors. (2) What does each cause? (3) What is the fix for each?
Show the solution
- One equation spans two business regimes with different relationships, which is inappropriate pooling. The clue is the change in the advertising effect midway through the sample.
- Raw rupee amounts for a small and a very large division, with no adjustment for size, are a separate problem: inappropriate scaling. The clue is the mixed magnitudes.
- Pooling gives a fitted slope that is an average and may describe neither regime.
- Scaling leaves residual variance that likely differs with division size, which is heteroskedasticity and makes standard errors and tests unreliable.
- Fix pooling by estimating separate regressions for each regime. Fix scaling by using ratios such as advertising as a percentage of sales.
Answer: (1) Inappropriate pooling and, separately, inappropriate scaling. (2) Pooling gives a misleading averaged slope; scaling can cause heteroskedasticity and unreliable tests. (3) Split the sample by regime; use ratio or common-size variables.
Exam tips
- Most questions ask you to name the error from a short description, so memorize the four types and their clues.
- For omitted variable questions, always test the correlation condition before claiming bias.
- If the vignette shows residual patterns, remember that misspecification can look like heteroskedasticity or serial correlation, and the true cause may be the form.
- Fixes are also tested: add variable, transform, split sample, or rescale. Match the fix to the error.
- No penalty for wrong answers, so never leave a question blank.
Misspecified Functional Form: frequently asked questions
What is misspecified functional form in regression?
It means the equation does not match the true relationship between the variables. The curriculum lists omitted variables, variables needing transformation, inappropriate pooling, and inappropriate scaling. Misspecification in general can lead to biased, inconsistent estimates and may appear as heteroskedasticity or serial correlation, which makes tests unreliable.
What is the difference between omitted variable bias and inappropriate scaling?
Omitted variable bias arises because an important explanatory variable is missing and is correlated with an included one. Inappropriate scaling arises because variables are in different units or sizes and are not put on a common basis. The fix for the first is adding the variable; for the second, using ratios or common-size data.
Can you give an example of inappropriate pooling of data?
Estimating one regression across two periods when a policy change altered the relationship, or across two industries with different behaviour. The pooled slope averages the groups and may describe neither. Estimate separate regressions instead.
Does omitting a variable always bias the other coefficients?
No. Bias in the included slopes arises when the omitted variable influences the dependent variable and is correlated with an included regressor. If it is uncorrelated with them, the slopes are not biased by its omission.