CMA Foundation · Fundamentals of Business Mathematics and Statistics
Correlation and Regression: formula sheet
Key formulas
- Range of the correlation coefficient
- -1 ≤ r ≤ +1
- r is the coefficient of correlation. It can never be less than -1 or more than +1.
- Perfect positive correlation
- r = +1
- All points lie on a straight line rising from left to right.
- Perfect negative correlation
- r = -1
- All points lie on a straight line falling from left to right.
- Zero correlation
- r = 0
- No linear relationship. A curved relationship can still exist, so r = 0 does not rule out non-linear dependence.
- Sign and direction
- r > 0: same direction; r < 0: opposite direction
- The sign gives direction. The size of r (ignoring the sign) gives strength.
- Definition of r
- r = Cov(X, Y) ÷ (σx × σy)
- Cov(X, Y) = Σ(X − X̄)(Y − Ȳ) ÷ n = ΣXY ÷ n − X̄Ȳ.
- Deviation-from-mean form
- r = Σ(X − X̄)(Y − Ȳ) ÷ √[Σ(X − X̄)² × Σ(Y − Ȳ)²]
- Use when the means are whole numbers. The factor n cancels out.
- Direct method
- r = [nΣXY − ΣXΣY] ÷ √{[nΣX² − (ΣX)²][nΣY² − (ΣY)²]}
- Best when the values are small. n is the number of pairs.
- Assumed mean method
- r = [nΣdxdy − ΣdxΣdy] ÷ √{[nΣdx² − (Σdx)²][nΣdy² − (Σdy)²]}
- dx = X − A and dy = Y − B, where A and B are assumed means. The same formula as direct, with d replacing X and Y.
- Step-deviation method
- u = (X − A) ÷ h, v = (Y − B) ÷ k; r = [nΣuv − ΣuΣv] ÷ √{[nΣu² − (Σu)²][nΣv² − (Σv)²]}
- If h and k are both positive, r is the same as for X and Y. If exactly one of h or k is negative, r changes sign.
- Limits of r
- −1 ≤ r ≤ +1
- r = ±1 means perfect linear correlation. r = 0 means no linear correlation.
- Change of origin and scale
- If u = (X − a) ÷ h and v = (Y − b) ÷ k with h, k > 0, then r(u, v) = r(X, Y)
- r does not depend on origin or on positive scale.
- Spearman's rank correlation (no ties)
- r = 1 − 6Σd² ÷ [n(n² − 1)]
- d = difference between the two ranks of an item; n = number of items (pairs).
- With repeated ranks (tie correction)
- r = 1 − 6[Σd² + Σ m(m² − 1) ÷ 12] ÷ [n(n² − 1)]
- m = number of items in a tied group. Add one correction for every tied group, in both series.
- Rank for tied items
- Average of the ranks the tied items would have taken
- Example: three items tied for 2nd, 3rd and 4th each get 3.
- Range and checks
- −1 ≤ r ≤ +1; Σd = 0
- If Σd is not zero, a rank or difference is wrong. r = +1 means identical ranks; r = −1 means fully reversed ranks.
- Correction value for common ties
- m = 2 gives 0.5; m = 3 gives 2
- Calculated from m(m² − 1) ÷ 12: 2×3÷12 = 0.5 and 3×8÷12 = 2.
- Probable error of r
- PE = 0.6745 × (1 − r²) ÷ √N
- N is the number of pairs of observations. Use r², not r.
- Standard error of r
- SE = (1 − r²) ÷ √N
- PE = 0.6745 × SE.
- Limits for population correlation
- r − PE to r + PE
- Range in which the population correlation is likely to lie.
- Significance rule of thumb
- If r > 6 × PE, r is significant. If r < 6 × PE, r is not significant.
- Also, if r is not significant it gives no evidence of correlation. Use |r| for negative values.
- Coefficient of determination
- r² = (explained variation) ÷ (total variation)
- Usually quoted as a percentage, r² × 100.
- Coefficient of non-determination
- 1 − r²
- Share of variation left unexplained.
- Regression line of Y on X
- Y − Ȳ = byx (X − X̄)
- Use to estimate Y when X is given.
- Regression line of X on Y
- X − X̄ = bxy (Y − Ȳ)
- Use to estimate X when Y is given.
- byx from raw data
- byx = [n ΣXY − ΣX ΣY] ÷ [n ΣX² − (ΣX)²]
- The denominator uses the X column.
- bxy from raw data
- bxy = [n ΣXY − ΣX ΣY] ÷ [n ΣY² − (ΣY)²]
- The denominator uses the Y column. The numerator is the same as for byx.
- Coefficients using r and standard deviations
- byx = r × σy ÷ σx ; bxy = r × σx ÷ σy
- Use when r and the standard deviations are given.
- Link with correlation
- r² = byx × bxy ; r = ±√(byx × bxy)
- r takes the common sign of byx and bxy. Product of the coefficients cannot exceed 1.
- Normal equations for Y = a + bX
- ΣY = na + bΣX ; ΣXY = aΣX + bΣX²
- Solve the two equations for a and b. This is the least squares method.
- Intersection of lines
- Both lines pass through (X̄, Ȳ)
- Solve the two line equations together to get the means.
- Regression coefficient of y on x
- byx = r × σy ÷ σx = Σ(x − x̄)(y − ȳ) ÷ Σ(x − x̄)²
- This is the slope of the line of y on x. The denominator is the sum of squared deviations of x. If you divide the top and bottom by n, it becomes covariance ÷ variance of x.
- Regression coefficient of x on y
- bxy = r × σx ÷ σy = Σ(x − x̄)(y − ȳ) ÷ Σ(y − ȳ)²
- This is the slope of the line of x on y. The denominator is the sum of squared deviations of y. If you divide the top and bottom by n, it becomes covariance ÷ variance of y.
- Link with correlation
- r² = byx × bxy, so r = ±√(byx × bxy)
- Give r the common sign of byx and bxy.
- Sign property
- byx, bxy and r always have the same sign
- If one is negative, the other is too.
- Size property
- byx × bxy ≤ 1, and |byx + bxy| ÷ 2 ≥ |r|
- The two coefficients always have the same sign. So the AM of their absolute values is at least |r|. This applies to absolute values, because for negative coefficients the plain AM is negative. If the question does not label which line is y on x and a product comes out above 1, you have assigned the equations wrongly, so swap them. If the question does label the lines and the product is above 1, the data are inconsistent. Swapping is not the fix in that case.
- Regression lines
- y − ȳ = byx (x − x̄) and x − x̄ = bxy (y − ȳ)
- Both lines pass through (x̄, ȳ).
- Slope from an equation ax + by + c = 0
- As y on x: byx = −a ÷ b. As x on y: bxy = −b ÷ a
- This holds only when the equation is written as ax + by + c = 0, with everything on one side.
- Change of origin and scale
- If u = (x − a) ÷ h and v = (y − c) ÷ k, then byx = (k ÷ h) × bvu and bxy = (h ÷ k) × buv
- Regression coefficients do not change with the origin. They change with the scale.
Quick revision
- r always lies between -1 and +1, and it has no unit.
- r = +1 is perfect positive, r = -1 is perfect negative, and r = 0 means no linear relationship.
- Pearson's r = Cov(X, Y) ÷ (σx × σy).
- Correlation does not prove cause and effect.
- Spearman's R = 1 - 6Σd² ÷ [n(n² - 1)], where d is the difference in ranks.
- For tied ranks, add a correction of (m³ - m) ÷ 12 to Σd² for each tied group of m items.
- Probable error = 0.6745 × (1 - r²) ÷ √n.
- Coefficient of determination = r², and it shows the share of variation explained.
- Regression of Y on X: Y - Ȳ = byx (X - X̄), where byx = r × σy ÷ σx.
- Regression of X on Y: X - X̄ = bxy (Y - Ȳ), where bxy = r × σx ÷ σy.
- r = ±√(bxy × byx), and the sign is the same as that of both b values.
- Both regression lines pass through the point (X̄, Ȳ).
Common mistakes
- Treating correlation as proof of cause and effect. Fix: Remember correlation only shows association. Both may be driven by a third factor.
- Thinking negative correlation means a weak relationship. Fix: The sign is only direction. r = -0.9 is a stronger relationship than r = +0.4.
- Writing (ΣX)² as ΣX² (or the reverse). Fix: Keep separate columns. ΣX² comes from squaring each X and then adding. (ΣX)² comes from adding first and then squaring.
- Forgetting to multiply ΣXY by n in the numerator. Fix: In the sums formula, every first term has n: nΣXY, nΣX² and nΣY². Write the formula before you start.
- Ranking one series from highest and the other from lowest Fix: Choose one direction before you start and use it for both series.
- Giving tied items the same rank as the first position, such as 3 and 3 Fix: Use the average of the positions. Two items tied for 3rd and 4th both get 3.5, and the next item gets 5.
- Using r instead of r² inside the PE formula Fix: Always write (1 − r²). Square r before subtracting from 1.
- Treating r as the percentage of variation explained Fix: Square r first. r = 0.7 means 49% explained, not 70%.
- Using the Y on X line to estimate X (or the reverse). Fix: Estimate Y from X with the line of Y on X, and X from Y with the line of X on Y. Choose the line by the dependent variable.
- Using ΣX² in the denominator of bxy. Fix: byx has the X column in the denominator. bxy has the Y column. Say it as 'divide by the variable you are predicting from'.
Exam tips
- Questions on this topic are mostly conceptual, so learn one clear example for each type.
- Memorise -1 ≤ r ≤ +1 and use it to remove wrong options quickly.
- For scatter diagram questions, read direction first and tightness second.
- Beware of options claiming that correlation proves causation. They are usually wrong.
- Distinguish the sign (direction) from the magnitude (strength) in every question.
- Check the sign of the numerator first. If it is negative, r is negative, and you can drop all positive options at once.
- Learn the property questions: r lies between −1 and +1, has no unit, and is unchanged by positive change of origin and scale.
- If the question gives Cov(X, Y), σx and σy, just compute Cov ÷ (σx × σy). No table is needed.