Management Accounting · Summarising and analysing data
Correlation, Regression and Linear Programming Basics for ACCA MA
Updated 11 October 2026 · Fact-checked
Correlation measures how strongly two variables move together, using r from −1 to +1. Regression fits a straight line, y = a + bx, to data so you can estimate costs or forecast. To solve questions, total the data, calculate b, then a, then substitute x. Square r to get r².
Understand Correlation, Regression and Linear Programming Basics
Management accountants often need to know whether one thing drives another. Does cost rise with output? Do sales rise with advertising? A scatter diagram plots paired data, with the independent variable (x, the driver) on the horizontal axis and the dependent variable (y, the result) on the vertical axis. If the points form a rough line, there is a relationship.
The correlation coefficient (r) puts a number on that. It always lies between −1 and +1. Near +1 means strong positive correlation: as x rises, y rises. Near −1 means strong negative correlation: as x rises, y falls. Near 0 means little or no linear correlation. Correlation shows a linear association only. It does not prove that x causes y.
The coefficient of determination (r²) is r squared. It is the proportion of the variation in y that is explained by variation in x. If r = 0.9, then r² = 0.81, so 81% of the variation in y is explained by x. The other 19% is due to other factors or chance. Because it is squared, r² is never negative.
Linear regression finds the line of best fit, y = a + bx, using the least squares method. It chooses the line that makes the sum of the squared vertical gaps between the points and the line as small as possible. Here a is the intercept (for costs, the fixed cost) and b is the slope (for costs, the variable cost per unit). You then use the line to estimate cost or forecast y for a given x.
The high-low method is a simpler way to get a cost line. It uses only two data points, the highest and lowest activity levels. Least squares uses every data point, so it is usually more reliable. This page focuses on correlation and regression. Linear programming is a separate technique and is not covered here.
Key formulas to remember
- Correlation coefficient
- r = (nΣxy − ΣxΣy) ÷ √[(nΣx² − (Σx)²) × (nΣy² − (Σy)²)]
- n is the number of pairs. The answer must lie between −1 and +1. If it does not, you have made an arithmetic error.
- Coefficient of determination
- r² = r × r
- Gives the proportion (or percentage) of variation in y explained by x. Never negative.
- Regression line
- y = a + bx
- y is the dependent variable, x the independent variable. In cost estimation, a is fixed cost and b is variable cost per unit.
- Slope (b)
- b = (nΣxy − ΣxΣy) ÷ (nΣx² − (Σx)²)
- The numerator is the same as in the r formula. Calculate it once and reuse it.
- Intercept (a)
- a = (Σy − bΣx) ÷ n, or a = ȳ − b x̄
- Calculate b first. ȳ and x̄ are the mean values of y and x.
- High-low variable cost
- b = (cost at highest activity − cost at lowest activity) ÷ (highest activity − lowest activity)
- Then fixed cost = total cost at either point − (b × activity at that point). Use the highest and lowest activity levels, not the highest and lowest costs.
How to solve Correlation, Regression and Linear Programming Basics questions
Use this method for any question that gives paired data and asks for r, r², the regression line or a forecast.
- 1Identify x (the driver, such as activity level) and y (the result, such as total cost). Mixing them up changes the line.
- 2Count n, the number of pairs. Build columns for x, y, xy, x² and, if r is needed, y².
- 3Total each column to get Σx, Σy, Σxy, Σx² and Σy².
- 4Calculate the numerator nΣxy − ΣxΣy. Use it for both b and r.
- 5Calculate b = numerator ÷ (nΣx² − (Σx)²). Then calculate a = (Σy − bΣx) ÷ n.
- 6If asked, calculate r with the full formula, then square it to get r². Interpret it in words: direction, strength and the percentage explained.
- 7For a forecast, substitute x into y = a + bx. Check the units of x and y (for example, thousands) and state the answer in the right units.
- 8Comment on reliability: how strong r is, how many data points there are, and whether x lies inside the data range (interpolation) or outside it (extrapolation).
Quickest way: Calculator-first shortcut for objective test questions
When to use it: Use this when the question gives you the totals (Σx, Σy, Σxy, Σx²) or a short data set, and you only need b, a, r or a forecast.
- Check whether the question already gives the sums. If so, skip the table.
- Work out nΣxy − ΣxΣy once and write it down.
- Divide by nΣx² − (Σx)² to get b. Then get a from (Σy − bΣx) ÷ n.
- Use the on-screen calculator memory to avoid retyping long numbers.
- Sense-check: b should have the same sign as r, and a forecast inside the data range should sit between the observed y values.
- For r² questions, you do not need the data at all. Just square r and convert to a percentage.
Common mistakes in Correlation, Regression and Linear Programming Basics
Saying r = 0.8 means 80% of variation is explained.
Students mix up r and r².
Fix: Square r first. r = 0.8 gives r² = 0.64, so 64% is explained.
Ignoring the sign of r when interpreting.
Students focus on size and forget direction.
Fix: State both. A negative r means y falls as x rises. r = −0.9 is a strong negative relationship, not a weak one.
Concluding that high correlation proves x causes y.
A strong number feels like proof.
Fix: Say correlation shows association only. Other factors, or coincidence, may be involved.
Swapping x and y, or using total cost as x.
Students do not stop to identify the driver.
Fix: Activity or the cause is x. Cost or the result is y. Write this down before calculating.
Calculating a before b, or using the wrong sign.
Rushing through the formulas.
Fix: Always find b first. Then a = (Σy − bΣx) ÷ n. If b is negative, keep the minus sign in the calculation.
Using the high-low method with the highest and lowest costs rather than activity levels.
Students pick the biggest and smallest numbers in the table.
Fix: Choose the points with highest and lowest activity, and take the costs that go with them.
Worked examples
Example 1
A company records output (x, in thousands of units) and total cost (y, in $000) over five months: (1, 12), (2, 15), (3, 19), (4, 22), (5, 27). Calculate the regression line y = a + bx, and estimate the total cost for 3,500 units. Also calculate r.
Show the solution
- n = 5. Σx = 1+2+3+4+5 = 15. Σy = 12+15+19+22+27 = 95.
- Σxy = 12 + 30 + 57 + 88 + 135 = 322. Σx² = 1+4+9+16+25 = 55. Σy² = 144+225+361+484+729 = 1,943.
- Numerator: nΣxy − ΣxΣy = (5 × 322) − (15 × 95) = 1,610 − 1,425 = 185.
- b = 185 ÷ (5 × 55 − 15²) = 185 ÷ (275 − 225) = 185 ÷ 50 = 3.7.
- a = (95 − 3.7 × 15) ÷ 5 = (95 − 55.5) ÷ 5 = 7.9.
- Line: y = 7.9 + 3.7x. For 3,500 units, x = 3.5: y = 7.9 + 12.95 = 20.85.
- For r: nΣy² − (Σy)² = (5 × 1,943) − 9,025 = 9,715 − 9,025 = 690. r = 185 ÷ √(50 × 690) = 185 ÷ √34,500 = 185 ÷ 185.74 = 0.996.
- r² is about 0.99, so about 99% of the variation in cost is explained by output. This is a very strong positive relationship.
Answer: y = 7.9 + 3.7x, which means fixed cost of $7,900 and variable cost of $3.70 per unit. Estimated cost for 3,500 units is $20,850. r ≈ 0.996.
Example 2
A study of advertising spend and sales finds r = −0.7. Interpret this, and state how much of the variation in sales is explained by advertising spend. Separately, using the high-low method, total cost was $46,000 at 8,000 units (the highest activity) and $26,000 at 3,000 units (the lowest). Find the variable cost per unit and the fixed cost.
Show the solution
- The sign is negative, so as advertising spend rises, sales tend to fall. A value of 0.7 in size is a fairly strong linear relationship.
- r² = (−0.7)² = 0.49. So 49% of the variation in sales is explained by advertising spend. The other 51% is due to other factors.
- Correlation does not prove cause, so you cannot say advertising reduces sales.
- High-low: variable cost per unit = (46,000 − 26,000) ÷ (8,000 − 3,000) = 20,000 ÷ 5,000 = $4.
- Fixed cost = 46,000 − (4 × 8,000) = 46,000 − 32,000 = $14,000.
- Check with the low point: 14,000 + (4 × 3,000) = 26,000. Correct.
Answer: r = −0.7 shows a fairly strong negative relationship, and 49% of the variation in sales is explained by advertising. High-low gives a variable cost of $4 per unit and fixed cost of $14,000.
Exam tips
- In multiple choice questions, check the sign and the range of r first. Any option with r above 1 or below −1 is wrong immediately.
- For number entry questions, check the units and rounding instruction. If x is in thousands, a forecast of 20.85 may need to be entered as 20,850.
- If a question asks about r², give the percentage, not the decimal, unless told otherwise. Remember it is explained variation in y.
- Treat forecasts outside the data range with caution. Examiners like the comment that extrapolation is less reliable than interpolation.
- Know the difference between methods. Least squares uses all the data and is usually more reliable. High-low uses only two points and can be distorted by unusual values.
Practice questions from Summarising and analysing data
- A cost index rose from 200 to 230 between two years. What was the percentage increase in cost over that period?
- A regression of total monthly maintenance cost (y, in $) on machine hours (x) gives y = 4,200 + 3.5x. Machine hours next month are forecast …
- A management accountant calculates the correlation coefficient between advertising spend and monthly sales and obtains r = -0.92. Which stat…
- A price index for raw materials was 125 in Year 1 (base year Year 0 = 100) and 150 in Year 2. A material cost $4,000 per tonne in Year 1. As…
- Monthly overhead cost data of Kirov Co are grouped into a frequency table with class intervals 0 to under 10, 10 to under 20, 20 to under 40…
Correlation, Regression and Linear Programming Basics in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Correlation, Regression and Linear Programming Basics: frequently asked questions
How do I calculate the correlation coefficient in ACCA MA?
Build totals for Σx, Σy, Σxy, Σx² and Σy², then apply r = (nΣxy − ΣxΣy) ÷ √[(nΣx² − (Σx)²)(nΣy² − (Σy)²)]. The result must lie between −1 and +1. In the exam, you are often given the formula sheet values or the sums, so practise the substitution.
What is the difference between the least squares method and the high-low method?
Least squares uses all the data points and finds the line that minimises the sum of squared gaps. High-low uses only the highest and lowest activity levels. Least squares is usually more accurate, but high-low is quicker to do by hand.
How do I interpret r squared?
r² is the proportion of variation in y explained by variation in x. For example, r² = 0.64 means 64% is explained and 36% is due to other factors. A higher r² means the regression line is a more reliable basis for forecasting.
Does a high correlation mean one variable causes the other?
No. Correlation shows that two variables move together in a linear way. It does not prove cause. A third factor may drive both, or the link may be coincidence.