Strategic Cost Management · Business Forecasting Models - Time Series and Regression Analysis
Correlation and Standard Error of Estimate Explained
Updated 11 October 2026 · Fact-checked
Correlation (r) measures the strength and direction of a linear relationship between two variables, from -1 to +1. The coefficient of determination (r²) shows the share of variation in Y explained by X. The standard error of estimate measures the typical gap between actual and regression-predicted Y. Smaller values mean more reliable forecasts.
Understand Correlation and Standard Error of Estimate
Correlation tells you whether two variables move together in a straight-line pattern. The Karl Pearson coefficient, r, ranges from -1 to +1. A value near +1 means that when X rises, Y tends to rise. A value near -1 means Y tends to fall. A value near 0 means there is no linear relationship. It does not prove that X causes Y.
The coefficient of determination, r², is the square of r. It tells you what proportion of the total variation in Y is explained by the regression on X. If r = 0.8, then r² = 0.64. So 64% of the variation in Y is explained by X and 36% is unexplained. Because r² is a square, it is never negative and it loses the direction of the relationship.
The regression line gives a forecast, but actual points do not sit exactly on it. The standard error of estimate (Se) measures how far actual Y values scatter around the fitted line. It is like a standard deviation of the residuals, where a residual = actual Y - forecast Y. It is in the same units as Y, for example ₹ lakh. A small Se relative to the level of Y means forecasts are reliable.
The three ideas are linked. A high |r| goes with a high r² and a low Se. The slope of the regression line is also linked to r: byx = r × σy ÷ σx. The two regression coefficients and r always have the same sign, and r is the geometric mean of byx and bxy.
In Strategic Cost Management you use these measures to judge whether a regression-based cost or sales forecast can be trusted before you use it in a decision.
Key rules to remember
- Karl Pearson correlation coefficient
- r = [nΣXY − ΣXΣY] ÷ √{[nΣX² − (ΣX)²] × [nΣY² − (ΣY)²]}
- Use when you have raw data. The value always lies between -1 and +1.
- Covariance form of r
- r = Cov(X, Y) ÷ (σx × σy), where Cov(X, Y) = ΣXY ÷ n − X̄ × Ȳ
- Use when means and standard deviations are given or easy to find. Use the same divisor (n) throughout.
- Coefficient of determination
- r² = Explained variation ÷ Total variation = 1 − (Unexplained variation ÷ Total variation)
- Express as a percentage when interpreting. It lies between 0 and 1.
- Link between r and regression coefficients
- r = ±√(byx × bxy); byx = r × σy ÷ σx; bxy = r × σx ÷ σy
- Take the sign of the regression coefficients. Both have the same sign as r. Their product, r², cannot exceed 1.
- Standard error of estimate (definition)
- Se = √[Σ(Y − Ŷ)² ÷ (n − 2)]
- Divisor n − 2 is used for a simple linear regression fitted from sample data. Some questions divide by n; follow the question's instruction.
- Standard error of estimate (computational)
- Se = √[(ΣY² − aΣY − bΣXY) ÷ (n − 2)]
- Use for Y = a + bX when the sums are given and residuals are not worked out.
- Se from r
- Unexplained variation = Σ(Y − Ȳ)² × (1 − r²); Se = √(Unexplained variation ÷ (n − 2))
- Quick route when r and the total variation in Y are given. If using σy with divisor n, then Se = σy × √(1 − r²).
- Approximate forecast range
- Ŷ ± 1 Se ≈ 68% range; Ŷ ± 2 Se ≈ 95% range
- A rough rule that assumes residuals are roughly normal and the sample is reasonably large. Use it only when the question asks for it.
How to solve Correlation and Standard Error of Estimate questions
Use this order for any question on correlation, r² or the standard error of estimate. It works for raw data and for summary data.
- 1Read what is asked: r, r², regression line, Se, or a forecast with its reliability. Note which variable is X (the independent one) and which is Y.
- 2List what is given: raw pairs, or summary figures such as n, ΣX, ΣY, ΣXY, ΣX², ΣY². Build a small table with X, Y, XY, X² and Y² if you have raw data.
- 3Total the columns and write n. Check one total again, because one slip spoils every later step.
- 4Compute r with the formula. Check that it lies between -1 and +1 and that the sign matches the pattern of the data.
- 5Square r to get r². Interpret it in a sentence: 'x% of the variation in Y is explained by X.'
- 6Find the regression line Y = a + bX if needed: b = [nΣXY − ΣXΣY] ÷ [nΣX² − (ΣX)²] and a = Ȳ − bX̄.
- 7Compute Se. Either find the residuals, square and sum them, then divide by n − 2, or use the computational formula or Unexplained variation = Total variation × (1 − r²).
- 8State a conclusion: forecast value, Se, and whether the fit is strong enough to rely on. Give units and a clear recommendation.
Quickest way: Shortcut using total variation and r²
When to use it: Use when r (or r²) and Σ(Y − Ȳ)² or σy are given, or when you can find them quickly. It saves you from computing every residual.
- Find the total variation in Y: Σ(Y − Ȳ)² = ΣY² − n × Ȳ².
- Unexplained variation = Total variation × (1 − r²).
- Divide the unexplained variation by n − 2 and take the square root to get Se.
- Cross-check: explained variation = Total variation × r². Explained plus unexplained must equal total.
- For r, use the shifted-origin idea: subtracting or dividing all X or Y by a common number does not change r, so reduce large figures before you compute.
Common mistakes in Correlation and Standard Error of Estimate
Treating r = 0.8 as '80% explained'.
Students mix up r and r².
Fix: Explained variation is r², not r. If r = 0.8, r² = 0.64, so 64% is explained.
Concluding that a high correlation proves that X causes Y.
Strong numbers feel like proof.
Fix: Say that the variables move together. Cause needs reasoning beyond the number. Two variables may both be driven by a third, such as inflation.
Dividing the squared residuals by n instead of n − 2 for Se.
Students copy the standard deviation formula.
Fix: For a sample-fitted simple regression, use n − 2, since two parameters (a and b) are estimated. Divide by n only if the question or its formula says so.
Giving r a positive sign when the regression coefficients are negative.
Taking the square root of byx × bxy gives a plain positive number.
Fix: The sign of r is the sign of byx and bxy. If both are negative, r is negative.
Comparing Se across different units or different levels of Y without care.
Se looks like a pure number, so students judge it on its own.
Fix: Se is in the units of Y. Judge it against the size of Y or against the mean. A Se of ₹5 lakh is small for sales of ₹5 crore and large for sales of ₹20 lakh.
Arithmetic slips in the sums, so r exceeds 1 or the answer is nonsense.
Large totals and a long table under time pressure.
Fix: Always check that -1 ≤ r ≤ 1. Reduce the data by subtracting a common figure and re-total a column if the result looks wrong.
Worked examples
Example 1
A firm records advertising spend (X, ₹ lakh) and sales (Y, ₹ crore) for five months: X = 1, 2, 3, 4, 5 and Y = 2, 4, 5, 4, 5. Compute (a) the Karl Pearson correlation coefficient, (b) r² with its meaning, (c) the regression line of Y on X, and (d) the standard error of estimate.
Show the solution
- Totals: n = 5; ΣX = 15; ΣY = 20; ΣXY = 2 + 8 + 15 + 16 + 25 = 66; ΣX² = 1 + 4 + 9 + 16 + 25 = 55; ΣY² = 4 + 16 + 25 + 16 + 25 = 86.
- Numerator of r: nΣXY − ΣXΣY = 5 × 66 − 15 × 20 = 330 − 300 = 30.
- Denominator: [5 × 55 − 15²] = 275 − 225 = 50; [5 × 86 − 20²] = 430 − 400 = 30; √(50 × 30) = √1,500 = 38.73.
- r = 30 ÷ 38.73 = 0.775 (approximately).
- r² = 30² ÷ (50 × 30) = 900 ÷ 1,500 = 0.60. So 60% of the variation in sales is explained by advertising spend.
- Regression: b = 30 ÷ 50 = 0.6. X̄ = 3 and Ȳ = 4, so a = 4 − 0.6 × 3 = 2.2. Line: Ŷ = 2.2 + 0.6X.
- Fitted values for X = 1 to 5: 2.8, 3.4, 4.0, 4.6, 5.2. Residuals (Y − Ŷ): -0.8, 0.6, 1.0, -0.6, -0.2. Squares: 0.64, 0.36, 1.00, 0.36, 0.04; total = 2.40.
- Se = √(2.40 ÷ (5 − 2)) = √0.8 = 0.894 (approximately).
- Check: total variation = 86 − 5 × 4² = 6. Unexplained = 6 × (1 − 0.6) = 2.4. This matches.
Answer: r ≈ 0.775; r² = 0.60 (60% of sales variation explained by advertising); Ŷ = 2.2 + 0.6X; Se ≈ ₹0.89 crore. The fit is moderate, so the forecast is usable but should be treated with caution on five data points.
Example 2
For 12 pairs of observations of X (units produced) and Y (overhead cost, ₹ lakh), r = 0.8, the total variation in Y, Σ(Y − Ȳ)², is 500, and the regression coefficient of Y on X is byx = 0.5. Find (a) the explained and unexplained variation, (b) the standard error of estimate, and (c) the regression coefficient of X on Y.
Show the solution
- r² = 0.8² = 0.64.
- Explained variation = 500 × 0.64 = 320.
- Unexplained variation = 500 × (1 − 0.64) = 500 × 0.36 = 180. Check: 320 + 180 = 500.
- Se = √(180 ÷ (12 − 2)) = √18 = 4.243 (approximately).
- Use r² = byx × bxy. So bxy = 0.64 ÷ 0.5 = 1.28.
- Check the sign: r is positive, and both 0.5 and 1.28 are positive. Consistent.
Answer: Explained variation = 320; unexplained variation = 180; Se ≈ ₹4.24 lakh; bxy = 1.28. Since 64% of the variation is explained, overhead forecasts from this line are reasonably reliable.
Exam tips
- In MCQs, test the interpretation first. Check whether the question asks about r or r², and watch the sign and the range -1 to +1. Many options can be rejected in seconds.
- In written answers, always add one interpretation sentence after the number. Marks are often given for the conclusion, such as 'the forecast is reliable because r² is high and Se is small'.
- State the divisor you use for Se (n − 2 or n) and follow the question's wording. If it is silent, use n − 2 and say so.
- Use r = ±√(byx × bxy) to save time when both regression coefficients are given. Take the sign from the coefficients and check that the product is not above 1.
- In a forecasting question, give the point forecast first, then comment on reliability using r² and Se. Add that forecasting far outside the observed range of X is risky.
Practice questions from Business Forecasting Models - Time Series and Regression Analysis
- Using exponential smoothing with alpha = 0.3, a Kolkata manufacturer forecast sales of 200 units for March. Actual March sales were 240 unit…
- Using simple exponential smoothing with α = 0.4, a Kochi firm forecast sales of 500 units for May and actual sales were 600. What is the Jun…
- Sales (units) for four years are 100, 120, 140, 160 with t = 1 to 4. Exponential smoothing with alpha 0.5 starts with the forecast for Year …
- Exponential smoothing with alpha = 0.4 is used by a Chennai retailer. The forecast for May was 500 units and actual May sales were 560 units…
- A Pune retailer fits a linear trend to quarterly sales (in Rs lakh) using coded time t, where t = 0 at the middle of the series. The fitted …
Correlation and Standard Error of Estimate: frequently asked questions
What is the formula for the standard error of estimate in regression?
Se = √[Σ(Y − Ŷ)² ÷ (n − 2)], where Ŷ is the value predicted by the regression line. A computational form is √[(ΣY² − aΣY − bΣXY) ÷ (n − 2)]. Some questions divide by n instead, so follow the wording of the question.
How do you interpret r squared (coefficient of determination)?
r² is the proportion of the total variation in Y that the regression on X explains. If r² = 0.75, then 75% of the variation in Y is explained by X and 25% is due to other factors. It lies between 0 and 1 and does not show direction.
What is the relationship between regression coefficients and the correlation coefficient?
r is the geometric mean of the two regression coefficients: r = ±√(byx × bxy), with the sign of the coefficients. Also byx = r × σy ÷ σx and bxy = r × σx ÷ σy. Both regression coefficients always carry the same sign.
Does a high correlation mean the regression forecast is accurate?
A high |r| means a strong linear pattern, which usually goes with a small Se. But accuracy depends on the size of Se compared with Y, on the number of observations, and on whether you forecast within the range of the data. Correlation alone does not prove causation.