Skip to content

Financial Management and Business Data Analytics · Data Analysis and Modelling

Correlation and Regression Modelling for CMA Inter

Updated 10 October 2026 · Fact-checked

Correlation measures the strength and direction of a linear relationship between two variables, shown by r from −1 to +1. Regression builds an equation, such as Y = a + bX, to predict one variable from another. To solve: tabulate sums, find b, then a, then predict.

Understand Correlation and Regression Modelling

Businesses want to know if two things move together. Does higher advertising go with higher sales? Does a rise in interest rates go with a fall in loan demand? Correlation answers this. It gives one number, r, that shows how strongly and in which direction two variables move together.

r lies between −1 and +1. A value near +1 means both rise together. A value near −1 means one rises as the other falls. A value near 0 means there is little linear relationship. Correlation shows association only. It does not prove that one variable causes the other.

Regression goes one step further. It fits a line through the data so you can predict. In simple regression, Y = a + bX. Y is the dependent variable you want to predict. X is the independent variable you use to predict it. The slope b tells you how much Y changes for a one-unit change in X. The intercept a is the value of Y when X is zero.

Multiple regression uses more than one independent variable, for example Y = a + b1X1 + b2X2. Sales may depend on advertising and on the number of outlets. Each coefficient shows the effect of its variable while the others are held constant. In the exam you mostly interpret computer output rather than compute by hand.

The line is fitted by least squares. It makes the sum of squared gaps between actual and predicted Y as small as possible. R² shows the share of the variation in Y that the model explains.

Key rules to remember

Karl Pearson's correlation coefficient
r = [nΣXY − ΣX·ΣY] ÷ √{[nΣX² − (ΣX)²] × [nΣY² − (ΣY)²]}
Use when you have raw paired data. r is always between −1 and +1.
Correlation using covariance
r = Cov(X, Y) ÷ (σx × σy)
Cov(X, Y) = ΣXY ÷ n − X̄·Ȳ. Use when the question gives covariance and standard deviations.
Regression line of Y on X
Y = a + bX, where b = [nΣXY − ΣX·ΣY] ÷ [nΣX² − (ΣX)²] and a = Ȳ − b·X̄
Use this to predict Y from X. The line always passes through (X̄, Ȳ).
Regression coefficients using r
b(yx) = r × σy ÷ σx ; b(xy) = r × σx ÷ σy
b(yx) is for Y on X. b(xy) is for X on Y. Both have the same sign as r.
Link between r and the two coefficients
r² = b(yx) × b(xy)
r = √(b(yx) × b(xy)), with the sign of the coefficients. Both coefficients must have the same sign.
Coefficient of determination
R² = 1 − SSE ÷ SST
SSE is the sum of squared errors. SST is the total sum of squares. In simple regression, R² = r².
Adjusted R²
Adjusted R² = 1 − (1 − R²) × (n − 1) ÷ (n − k − 1)
k is the number of independent variables. It penalises adding variables that add little.
Spearman's rank correlation (no tied ranks)
ρ = 1 − 6Σd² ÷ [n(n² − 1)]
d is the difference between the two ranks of each item. Use when data are ranks.

How to solve Correlation and Regression Modelling questions

Use this order for any question on correlation or regression. It keeps your working clean and earns step marks.

  1. 1Identify which variable is dependent (Y) and which is independent (X). The one to be predicted is Y.
  2. 2Check what is given. If you have raw pairs, build a table. If you have means, standard deviations and r, use the b = r × σ ratio formulas.
  3. 3For raw data, make columns for X, Y, XY, X² and Y². Total each column and write n.
  4. 4Compute b first, then a = Ȳ − b·X̄. Write the equation Y = a + bX.
  5. 5Compute r if asked, and check that its sign matches the sign of b.
  6. 6Predict by putting the given X into the equation. Say whether the X value is within the range of the data.
  7. 7Interpret in words: the direction and strength of the relationship, what b means in business terms, and what R² explains. State that correlation does not prove causation.

Quickest way: Deviation-free shortcut with totals

When to use it: Use when you get raw paired data and the numbers are small, or when the question gives means, standard deviations and r.

  1. With raw data, find only ΣX, ΣY, ΣXY, ΣX² and, if r is needed, ΣY². Do not build extra columns.
  2. Compute the numerator nΣXY − ΣXΣY once. It is common to both b and r, so reuse it.
  3. Find b by dividing it by nΣX² − (ΣX)². Find a using the means.
  4. If the question gives r, σx and σy, write b(yx) = r·σy ÷ σx directly and use Y − Ȳ = b(yx)(X − X̄).
  5. Check with r² = b(yx) × b(xy). A product above 1 means you have made an error.
  6. If X and Y are large numbers, you may subtract a common value from each. This does not change b or r, but you must adjust a at the end.

Common mistakes in Correlation and Regression Modelling

  • Writing that a high correlation proves that X causes Y.

    Students read a strong r as a cause-and-effect link.

    Fix: Say that r shows association only. Two variables may move together because of a third factor or by chance.

  • Regressing in the wrong direction, using X on Y when Y on X is asked.

    Students use the same b for both lines.

    Fix: Decide which variable is to be predicted. The two lines differ, with b(yx) = r·σy ÷ σx and b(xy) = r·σx ÷ σy.

  • Confusing ΣX² with (ΣX)² in the denominator.

    Both look similar and are rushed in the table.

    Fix: Square each X first and then add for ΣX². Add all X first and then square for (ΣX)². Write both clearly.

  • Giving r a positive sign when b is negative, or finding r > 1.

    Students take the square root and forget the sign, or make an arithmetic slip.

    Fix: The sign of r is the sign of b. r must lie between −1 and +1. If not, recheck the sums.

  • Predicting far outside the data range.

    Students trust the equation for any X.

    Fix: Predict only within or close to the range of observed X. State that extrapolation is unreliable.

  • Treating a higher R² as always better in multiple regression.

    R² never falls when you add a variable.

    Fix: Compare adjusted R², and check that each added variable makes business sense.

Worked examples

Example 1

A firm records advertising spend (X, ₹ lakh) and sales (Y, ₹ lakh) for five months: X = 2, 4, 6, 8, 10 and Y = 20, 30, 35, 45, 55. (a) Find the regression line of Y on X. (b) Find r. (c) Estimate sales when advertising is ₹12 lakh.

Show the solution
  1. n = 5. ΣX = 30, ΣY = 185. ΣXY = 40 + 120 + 210 + 360 + 550 = 1,280. ΣX² = 4 + 16 + 36 + 64 + 100 = 220. ΣY² = 400 + 900 + 1,225 + 2,025 + 3,025 = 7,575.
  2. Numerator: nΣXY − ΣXΣY = 5 × 1,280 − 30 × 185 = 6,400 − 5,550 = 850.
  3. Denominator for b: nΣX² − (ΣX)² = 5 × 220 − 900 = 200. So b = 850 ÷ 200 = 4.25.
  4. X̄ = 30 ÷ 5 = 6 and Ȳ = 185 ÷ 5 = 37. So a = 37 − 4.25 × 6 = 37 − 25.5 = 11.5.
  5. Regression line: Y = 11.5 + 4.25X.
  6. For r: nΣY² − (ΣY)² = 5 × 7,575 − 185² = 37,875 − 34,225 = 3,650. r = 850 ÷ √(200 × 3,650) = 850 ÷ √730,000 = 850 ÷ 854.4 = 0.995 (approx).
  7. Estimate at X = 12: Y = 11.5 + 4.25 × 12 = 11.5 + 51 = 62.5.

Answer: (a) Y = 11.5 + 4.25X. (b) r ≈ 0.995, a very strong positive correlation, so R² ≈ 0.99. (c) Estimated sales = ₹62.5 lakh. Each extra ₹1 lakh of advertising goes with about ₹4.25 lakh more sales. The ₹12 lakh figure lies outside the observed range of 2 to 10, so treat it with caution.

Example 2

For two variables X and Y: X̄ = 40, Ȳ = 60, σx = 5, σy = 8 and r = 0.6. (a) Find the regression equations of Y on X and of X on Y. (b) Estimate Y when X = 50. (c) Verify the link between the coefficients and r.

Show the solution
  1. b(yx) = r × σy ÷ σx = 0.6 × 8 ÷ 5 = 0.96.
  2. Y on X: Y − 60 = 0.96(X − 40). So Y = 0.96X − 38.4 + 60 = 0.96X + 21.6.
  3. b(xy) = r × σx ÷ σy = 0.6 × 5 ÷ 8 = 0.375.
  4. X on Y: X − 40 = 0.375(Y − 60). So X = 0.375Y − 22.5 + 40 = 0.375Y + 17.5.
  5. Estimate Y at X = 50: Y = 0.96 × 50 + 21.6 = 48 + 21.6 = 69.6.
  6. Check: b(yx) × b(xy) = 0.96 × 0.375 = 0.36, and r² = 0.6² = 0.36. The two agree.

Answer: (a) Y = 0.96X + 21.6 and X = 0.375Y + 17.5. (b) Estimated Y = 69.6 when X = 50. (c) b(yx) × b(xy) = 0.36 = r², so the figures are consistent.

Exam tips

  • Section A may ask for the meaning of r, r², or the sign of b. Learn what each value of r says and that r² = b(yx) × b(xy).
  • In written answers, show the table with totals and write the formula before substituting. Step marks come from the working, not just the final equation.
  • Always end with a one-line interpretation in business words. Say what b means and whether the relationship is strong, and avoid claiming causation.
  • For multiple regression, expect to read output and interpret coefficients, R² and adjusted R². Practise explaining each in one sentence.
  • Read carefully which line is asked: Y on X or X on Y. Using the wrong one is a common way to lose marks.

Practice questions from Data Analysis and Modelling

Correlation and Regression Modelling: frequently asked questions

What is the difference between correlation and regression?

Correlation measures how strongly two variables move together and gives one number, r. Regression gives an equation that predicts one variable from another. Correlation treats both variables alike, while regression separates a dependent variable from an independent one.

How do I calculate the regression line step by step?

Make a table of X, Y, XY and X², and total each column. Find b = (nΣXY − ΣXΣY) ÷ (nΣX² − (ΣX)²). Then find a = Ȳ − b·X̄ and write Y = a + bX.

What does R² tell me?

R² is the share of the variation in the dependent variable that the model explains. An R² of 0.80 means the model explains 80% of the variation in Y. In simple regression it equals r².

Why are there two regression lines?

One line minimises errors in predicting Y from X. The other minimises errors in predicting X from Y. They are the same line only when r is +1 or −1.

What is the use of linear regression in business analytics?

It is a predictive technique. Firms use it to forecast sales, costs or demand from drivers such as price or advertising, and to see which drivers matter most. Decisions on budgets and targets then rest on data.