CFA Level II Exam · Basics of Multiple Regression and Underlying Assumptions
Assumptions of the Multiple Linear Regression Model
Updated 7 October 2026 · Fact-checked
The multiple linear regression model assumes a linear relationship between the dependent and independent variables, independent variables that are not random and have no exact linear relationship, errors with zero mean, constant variance (homoskedasticity), no serial correlation, and normality. In an item set, you match each violation to its symptom and consequence.
Understand Assumptions of the Multiple Linear Regression Model
A multiple regression estimates a dependent variable Y from several independent variables: Y = b0 + b1X1 + b2X2 + ... + bkXk + ε. The error term ε is what the model does not explain. Every test you run on the coefficients, such as t-tests and the F-test, is only reliable if the assumptions about this setup hold.
There are six assumptions to know. Linearity: the relationship between Y and the X variables is linear in the coefficients. No exact collinearity: no independent variable is an exact linear combination of the others, and the X variables are not random. Zero mean errors: the expected value of the error, given the X variables, is zero. Homoskedasticity: the variance of the errors is the same across all observations. Independence: the errors are not correlated with each other, so there is no serial correlation. Normality: the errors are normally distributed.
The key link is between each assumption and its failure. Unequal error variance is heteroskedasticity. Correlated errors are serial correlation. Highly (but not perfectly) correlated X variables are multicollinearity. A wrong functional form is misspecification. Exact collinearity is different: the model cannot be estimated at all.
Why do these matter? Heteroskedasticity and serial correlation usually leave the coefficient estimates unbiased but make the standard errors unreliable. Then t-statistics and p-values mislead you. Multicollinearity inflates standard errors, so coefficients look insignificant even when the model fits well. Normality matters mainly for small samples, because with large samples test statistics are approximately valid anyway.
Level II questions give you a vignette with regression output or residual plots. You must name the assumption in doubt and say what it does to inference. Memorize the pairs, not just the list.
Key formulas to remember
- Multiple regression model
- Yi = b0 + b1X1i + b2X2i + ... + bkXki + εi
- Linear in the coefficients. Variables can be transformed (for example, a log) and still satisfy linearity.
- Zero conditional mean of errors
- E(ε | X1, ..., Xk) = 0
- If errors are related to an X variable, coefficient estimates are biased. This often signals an omitted variable or misspecification.
- Homoskedasticity
- Var(εi) = σ² for all i
- Constant error variance. Violation is heteroskedasticity.
- Independence of errors
- Cov(εi, εj) = 0 for i ≠ j
- Violation is serial correlation, common in time series.
- Normality of errors
- ε ~ N(0, σ²)
- Matters most for small samples.
- No exact collinearity
- No Xj is an exact linear combination of the other X variables
- Exact collinearity makes estimation impossible. High but imperfect correlation is multicollinearity.
How to solve Assumptions of the Multiple Linear Regression Model questions
Use this method for any question on regression assumptions in an item set.
- 1Read the question first to see which assumption or consequence is being tested.
- 2Find the evidence in the vignette: a residual plot, a test statistic, a correlation table, or a description of the data.
- 3Match the evidence to the assumption. Fanning residuals point to heteroskedasticity. Residual patterns over time point to serial correlation. High X-variable correlation with a significant F-test but insignificant t-tests points to multicollinearity. A curved residual pattern points to misspecified functional form.
- 4State the effect. Ask whether coefficients are biased, whether standard errors are unreliable, and whether t-tests are affected.
- 5Check the exact wording. Linear in coefficients is not the same as linear in variables, and exact collinearity is not the same as multicollinearity.
- 6Choose the option that names both the violation and its correct consequence, and reject options that get one right and one wrong.
Quickest way: Symptom to violation to consequence
When to use it: Use when you have little time and the vignette gives a clear symptom.
- Underline the symptom: fanning residuals, residuals correlated over time, high X correlations, or a curved residual pattern.
- Name the violation: heteroskedasticity, serial correlation, multicollinearity, or misspecification.
- Recall the consequence: unreliable standard errors for the first two, inflated standard errors for multicollinearity, and biased coefficients for misspecification that leaves out relevant variables.
- Pick the option that matches, and skip options that describe a different violation.
Common mistakes in Assumptions of the Multiple Linear Regression Model
Saying linearity means the X variables must enter without transformation.
The word linear sounds like it applies to the variables.
Fix: The model must be linear in the coefficients. A log of X or X squared can be used and the model is still linear in the parameters.
Treating multicollinearity as an exact-collinearity problem that stops estimation.
Both terms sound alike.
Fix: Exact collinearity means estimation fails. Multicollinearity means high but imperfect correlation, which inflates standard errors but still gives estimates.
Claiming heteroskedasticity biases the coefficient estimates.
Students assume any violation makes everything wrong.
Fix: With conditional heteroskedasticity the coefficients are usually still unbiased, but the standard errors are unreliable, so t-tests mislead.
Missing that serial correlation is mainly a time-series issue.
Students apply every assumption to every data type.
Fix: Look for time-ordered data. For cross-sectional data, heteroskedasticity is the more likely concern.
Assuming normality is critical for every sample size.
The assumption is listed with the others as equally important.
Fix: Normality matters most in small samples. With large samples, test statistics are approximately valid.
Reading multicollinearity as a poor overall fit.
Insignificant t-statistics look like a weak model.
Fix: Multicollinearity often shows as a high R² and significant F-test with insignificant individual coefficients. The model can still fit well.
Worked examples
Example 1
Vignette: An analyst regresses monthly fund returns on market return, a size factor and a value factor. A plot of residuals against the market return shows a funnel shape: residuals are small at low market returns and spread widely at high market returns. The coefficients look sensible. Q1: Which assumption is most likely violated? Q2: What is the main consequence for inference? Q3: Is the coefficient estimate for the market factor likely biased because of this violation?
Show the solution
- Q1: A funnel shape in residuals means the error variance changes with an X variable. That breaks homoskedasticity, so the violation is heteroskedasticity.
- Q2: The standard errors are unreliable, so t-statistics and p-values cannot be trusted, and hypothesis tests may be wrong.
- Q3: With this kind of heteroskedasticity the coefficient estimates are generally still unbiased. The problem lies in the standard errors.
Answer: Q1: Heteroskedasticity (constant variance violated). Q2: Unreliable standard errors and therefore misleading t-tests. Q3: No, the coefficients are generally still unbiased.
Example 2
Vignette: An analyst models a company's quarterly sales using advertising spend, a price index and a second price index built from the same data. The regression output shows R² of 0.91 and a significant F-statistic, but none of the three slope coefficients is significant on its own. The correlation between the two price indexes is 0.97. Q1: What is the most likely problem? Q2: How does it affect the standard errors? Q3: Would the problem be called exact collinearity?
Show the solution
- Q1: A high R² and significant F-test with insignificant t-tests, plus a 0.97 correlation between two X variables, is the classic sign of multicollinearity.
- Q2: Multicollinearity inflates the standard errors of the affected coefficients, which lowers the t-statistics and makes variables look insignificant.
- Q3: A correlation of 0.97 is high but not 1.0, so one variable is not an exact linear function of the other. This is multicollinearity, not exact collinearity. Exact collinearity would prevent estimation.
Answer: Q1: Multicollinearity. Q2: Standard errors are inflated. Q3: No, the correlation is high but not perfect, so it is not exact collinearity.
Exam tips
- Expect to name the violation from a symptom, then state its effect on coefficients or standard errors.
- Learn the pairs: heteroskedasticity and serial correlation affect standard errors, multicollinearity inflates them, and omitted variables or wrong functional form can bias coefficients.
- Watch for the phrase linear in the coefficients when an option mentions transformed variables.
- If an option says the model cannot be estimated, check whether the vignette describes exact or merely high correlation.
- There is no penalty for wrong answers, so never leave an item blank.
Assumptions of the Multiple Linear Regression Model in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Assumptions of the Multiple Linear Regression Model: frequently asked questions
What are the assumptions of multiple linear regression in CFA Level II?
The model is linear in the coefficients, the X variables are not random and have no exact linear relationship, errors have zero mean and constant variance, errors are uncorrelated with each other, and errors are normally distributed. Know the name of each violation and its effect.
What does linearity mean in regression?
It means the model is linear in the parameters. You can still use transformed variables such as logs or squares. What you cannot have is a parameter that enters in a non-linear way.
What is the difference between multicollinearity and exact collinearity?
Exact collinearity means one X variable is a perfect linear combination of others, and the model cannot be estimated. Multicollinearity means the variables are highly but not perfectly correlated, which inflates standard errors.
Which violations bias the regression coefficients?
Heteroskedasticity and serial correlation usually leave the coefficients unbiased but make the standard errors unreliable. Omitted variables and a wrong functional form can bias coefficients, because the errors become related to the included X variables.