Skip to content

Quantitative Aptitude · Correlation and Regression

Karl Pearson's Coefficient of Correlation (CA Foundation)

Updated 1 October 2026 · Fact-checked

Karl Pearson's coefficient of correlation, r, measures the strength and direction of a linear relationship between two variables. It lies from −1 to +1. Calculate it as covariance divided by the product of the two standard deviations, or use the sums formula with ΣX, ΣY, ΣXY, ΣX² and ΣY².

Understand Karl Pearson's Coefficient of Correlation

Correlation tells you whether two variables move together. When one rises, does the other tend to rise (positive), fall (negative), or show no pattern (zero)? Karl Pearson's r puts a number on this, but only for linear relationships.

Start with covariance. It is the average of the products of deviations from the means: Cov(X, Y) = Σ(X − X̄)(Y − Ȳ) ÷ n. If big X values go with big Y values, the products are mostly positive and covariance is positive. If big X goes with small Y, covariance is negative. The problem is that covariance depends on units. Rupees against kilograms gives a different size than rupees against grams.

To remove units, divide covariance by σx × σy. That gives r. It is a pure number between −1 and +1. A value near +1 means a strong positive linear relation. A value near −1 means a strong negative linear relation. A value near 0 means little or no linear relation.

r is not changed by shifting the origin (subtracting a constant) or by rescaling with a positive constant. That is why the assumed mean and step deviation methods work. They make the numbers smaller without changing r.

A high r shows association, not cause. Also, r = 0 only says there is no linear relation. A curved relation can still exist.

Key formulas to remember

Definition of r
r = Cov(X, Y) ÷ (σx × σy)
σx and σy are population standard deviations (divide by n). Use the same divisor in all three terms.
Covariance
Cov(X, Y) = Σ(X − X̄)(Y − Ȳ) ÷ n = ΣXY ÷ n − X̄ × Ȳ
The second form is faster when ΣXY, X̄ and Ȳ are known.
Direct method (sums)
r = [nΣXY − ΣX ΣY] ÷ [√(nΣX² − (ΣX)²) × √(nΣY² − (ΣY)²)]
Needs only five sums and n. No need to find the means.
Deviations from actual means
r = Σxy ÷ √(Σx² × Σy²), where x = X − X̄ and y = Y − Ȳ
Best when the means are whole numbers.
Assumed mean method
r = [nΣdxdy − ΣdxΣdy] ÷ [√(nΣdx² − (Σdx)²) × √(nΣdy² − (Σdy)²)], where dx = X − A and dy = Y − B
A and B are assumed means. Do not confuse the Σdx² terms with (Σdx)².
Step deviation method
u = (X − A) ÷ h, v = (Y − B) ÷ k, then r = r(u, v) using the same formula on u and v
Valid when h and k are both positive. If h and k have opposite signs, r changes sign.
Range and key properties
−1 ≤ r ≤ +1; r(X, Y) = r(Y, X); r is unchanged by change of origin and positive change of scale
If X and Y are independent, r = 0. The converse is not always true.
Coefficient of determination
r² = proportion of variation in one variable explained by a linear relation with the other
For example, r = 0.8 gives r² = 0.64.
Probable error
PE = 0.6745 × (1 − r²) ÷ √n
Commonly used rule: if r < PE, r is not significant. If r > 6 × PE, r is significant.

How to solve Karl Pearson's Coefficient of Correlation questions

Use this method for any Pearson's r question, whether the data are raw values or summary sums.

  1. 1Read what is given. If you get raw data, note n. If you get sums, check that you have ΣX, ΣY, ΣXY, ΣX² and ΣY².
  2. 2If the numbers are large, choose an assumed mean near the middle. If the values share a common factor, divide by it (step deviation).
  3. 3Make a table with columns for X, Y, their reduced forms, their squares and the product. Fill it row by row.
  4. 4Total each column. Check the totals once, because one addition error spoils the answer.
  5. 5Substitute into the sums formula. Work out the numerator and each square-root term separately.
  6. 6Simplify. Look for perfect squares in the denominator before using decimals.
  7. 7Check that your answer lies between −1 and +1 and that its sign matches the numerator.
  8. 8If asked, interpret r or compute r² or the probable error.

Quickest way: Shrink the numbers, then use the sums formula

When to use it: Use this for raw-data MCQs with five to ten pairs, or when the options are far enough apart that you do not need many decimals.

  1. Before any arithmetic, check the sign. If Y mostly rises as X rises, r must be positive. This often removes two options at once.
  2. Subtract a convenient constant from X and Y. If both columns share a common factor, divide by it. r stays the same.
  3. Use nΣuv − ΣuΣv for the numerator. If both Σu and Σv are 0, the numerator is just nΣuv.
  4. Keep the denominator as √(nΣu² − (Σu)²) × √(nΣv² − (Σv)²). Look for perfect squares such as 50 × 50 or 400.
  5. If the options are close and the arithmetic is heavy, check whether any option can be removed using the sign and the range −1 to +1. Skip the question if you cannot narrow it to two options, because each wrong answer costs 0.25 marks.

Common mistakes in Karl Pearson's Coefficient of Correlation

  • Writing Σdx² when the formula needs (Σdx)², or mixing the two up.

    The symbols look similar, and students rush the denominator.

    Fix: Compute them separately and label them: 'sum of squares' and 'square of sum'. In the denominator, use n × (sum of squares) − (square of sum).

  • Using the wrong n, such as the number of columns or n − 1, or forgetting n in the formula.

    Students copy the formula from memory and drop the n factors.

    Fix: Count the pairs of observations. In the sums formula, n multiplies ΣXY, ΣX² and ΣY² only, not the other term in each pair.

  • Thinking that r changes when each value is increased by a constant or multiplied by a positive constant.

    Students confuse correlation with covariance, which does change with scale.

    Fix: Remember that r is unaffected by change of origin and positive change of scale. Covariance changes with scale, but r does not.

  • Missing the sign flip when X is multiplied by a negative number.

    Students remember that scale does not matter and forget that a negative scale reverses direction.

    Fix: If exactly one of the two multipliers is negative, r changes sign. If both are negative or both positive, r stays the same.

  • Concluding that r = 0 means the variables are independent, or that a high r proves cause.

    Students read r as a measure of every kind of relation.

    Fix: r measures only linear association. Independence implies r = 0, but r = 0 does not imply independence. Correlation does not prove cause.

  • Accepting an answer like 1.2 or −1.5 without checking.

    An arithmetic slip goes unnoticed.

    Fix: Always check that |r| ≤ 1. If not, recheck your sums.

Worked examples

Example 1

For 10 pairs of observations, ΣX = 40, ΣY = 50, ΣXY = 230, ΣX² = 200 and ΣY² = 290. Karl Pearson's coefficient of correlation is: (a) 0.25 (b) 0.60 (c) 0.75 (d) 0.90

Show the solution
  1. Numerator = nΣXY − ΣXΣY = 10 × 230 − 40 × 50 = 2,300 − 2,000 = 300.
  2. First root: nΣX² − (ΣX)² = 10 × 200 − 1,600 = 400, so the root is 20.
  3. Second root: nΣY² − (ΣY)² = 10 × 290 − 2,500 = 400, so the root is 20.
  4. r = 300 ÷ (20 × 20) = 300 ÷ 400 = 0.75.

Answer: (c) 0.75

Example 2

X takes values 10, 20, 30, 40, 50 and Y takes values 20, 25, 30, 40, 35 respectively. Using u = (X − 30) ÷ 10 and v = (Y − 30) ÷ 5, the coefficient of correlation is: (a) 0.45 (b) 0.80 (c) 0.90 (d) −0.90

Show the solution
  1. u values: −2, −1, 0, 1, 2. v values: −2, −1, 0, 2, 1.
  2. Σu = 0 and Σv = 0.
  3. Σuv = (−2)(−2) + (−1)(−1) + 0 + (1)(2) + (2)(1) = 4 + 1 + 0 + 2 + 2 = 9.
  4. Σu² = 4 + 1 + 0 + 1 + 4 = 10. Σv² = 4 + 1 + 0 + 4 + 1 = 10.
  5. r = (5 × 9 − 0) ÷ (√(5 × 10) × √(5 × 10)) = 45 ÷ 50 = 0.90.
  6. h = 10 and k = 5 are both positive, so r(X, Y) = r(u, v) = 0.90.

Answer: (c) 0.90

Example 3

From 25 pairs of observations, r = 0.8. The probable error of r is approximately: (a) 0.0243 (b) 0.0486 (c) 0.0972 (d) 0.2430

Show the solution
  1. PE = 0.6745 × (1 − r²) ÷ √n.
  2. r² = 0.64, so 1 − r² = 0.36.
  3. √25 = 5.
  4. PE = 0.6745 × 0.36 ÷ 5 = 0.6745 × 0.072 = 0.048564, which is about 0.0486.
  5. Check: 6 × PE ≈ 0.29, and r = 0.8 is more than that, so r is significant.

Answer: (b) 0.0486

Exam tips

  • Questions often give summary sums, so practise the sums formula until the numerator and denominator are quick.
  • Property questions ask what happens to r after a change like U = aX + b. Positive a: r unchanged. Negative a: sign flips. Practise this logic, because it takes seconds.
  • Use the sign of the numerator and the range −1 to +1 to remove options before doing long arithmetic.
  • If a question gives regression coefficients, remember r = ±√(bxy × byx), with the sign matching the regression coefficients.
  • Interpretation questions use words like 'high positive' or 'perfect negative'. Know that r = +1 or −1 means all points lie on a straight line.

Practice questions from Correlation and Regression

Karl Pearson's Coefficient of Correlation: frequently asked questions

What is the formula for Karl Pearson's coefficient of correlation?

r = Cov(X, Y) ÷ (σx × σy). For raw sums, use r = [nΣXY − ΣXΣY] ÷ [√(nΣX² − (ΣX)²) × √(nΣY² − (ΣY)²)]. Both give the same value.

Does r change if I add a constant to every X value?

No. r is independent of change of origin. It is also independent of change of scale when the multiplier is positive. If the multiplier is negative, the sign of r flips.

What do the values of r mean?

r = +1 is perfect positive linear correlation, r = −1 is perfect negative, and r = 0 means no linear correlation. Values near ±1 show a strong linear relation. Values near 0 show a weak one.

How do I use probable error?

Calculate PE = 0.6745 × (1 − r²) ÷ √n. If r is less than PE, it is not significant. If r is more than 6 times PE, it is considered significant.

When should I use the step deviation method?

Use it when the values are large and share a common factor, such as 10, 20, 30. Dividing by that factor and subtracting an assumed mean gives small numbers, and r stays the same when both factors are positive.