Skip to content

ACCA Applied Knowledge · Management Accounting

Summarising and analysing data: formula sheet

Full chapter guide

Key formulas

Primary vs secondary data
Primary = collected first-hand for this purpose; Secondary = already collected by someone else
The test is who collected the data and why, not whether it is internal or external.
Quantitative vs qualitative data
Quantitative = numerical and measurable; Qualitative = descriptive, based on opinion or quality
Ratings scored from 1 to 5 are still based on opinion, but the scores are numerical.
Discrete vs continuous data
Discrete = counted, separate values; Continuous = measured, any value in a range
Ask: can the value sensibly be a fraction or decimal? If yes, it is continuous.
Internal vs external source
Internal = from within the organisation; External = from outside it
Internal data can be primary or secondary. Check the wording of the question.
Systematic sampling interval
Interval (n) = population size ÷ sample size
Pick a random start within the first n items, then take every nth item. Round sensibly if the answer is not whole.
Proportionate stratified sample
Sample from a stratum = (stratum size ÷ population size) × total sample size
Use this to allocate the sample across strata. Round so the total still equals the sample size.
Stratified vs cluster rule
Stratified: sample from every group. Cluster: sample all items from some groups.
Strata should differ from each other and be similar inside. Clusters should each resemble the whole population.
Class width
Class width = upper boundary − lower boundary
For 10 – under 20 the width is 10. Use boundaries, not stated limits, when classes have gaps.
Frequency density
Frequency density = frequency ÷ class width
Plot this on the vertical axis of a histogram when class widths are unequal.
Standard-width adjustment
Adjusted height = frequency ÷ (class width ÷ standard width)
An alternative to frequency density. A class twice the standard width has its frequency halved.
Pie chart angle
Angle = (category value ÷ total) × 360°
Angles must add to 360°.
Cumulative frequency
Cumulative frequency = running total of frequencies
Plot against the upper class boundary for an ogive.
Relative frequency
Relative frequency = frequency ÷ total frequency
Can be shown as a fraction, decimal or percentage.
Median position on an ogive
Read the value at cumulative frequency = n ÷ 2
Quartiles are read at n ÷ 4 and 3n ÷ 4.
Arithmetic mean (raw data)
x̄ = Σx ÷ n
Add all values, divide by the number of values.
Mean (frequency or grouped data)
x̄ = Σfx ÷ Σf
For grouped data, x is the class mid-point. The result is an estimate.
Class mid-point
Mid-point = (lower limit + upper limit) ÷ 2
Check class boundaries are continuous before using this.
Median position (raw data, ordered)
Position = (n + 1) ÷ 2
If the position is a half, take the average of the two middle values.
Median (grouped data)
Median = L + [(n ÷ 2 − cf) ÷ f] × w
L = lower boundary of median class, cf = cumulative frequency before it, f = its frequency, w = class width.
Mode (grouped data)
Mode = L + [d1 ÷ (d1 + d2)] × w
L = lower boundary of modal class, d1 = modal frequency minus previous frequency, d2 = modal frequency minus next frequency. Equal class widths assumed.
Weighted average
Weighted mean = Σwx ÷ Σw
w is the weight, such as units, hours or proportion.
Range
Range = highest value − lowest value
Uses only two values, so it is affected by extreme figures.
Interquartile range
IQR = Q3 − Q1
Measures the spread of the middle 50% of the data. Q1 and Q3 are the lower and upper quartiles.
Variance (population, ungrouped)
σ² = Σ(x − x̄)² ÷ n
Use n when the data is treated as the whole population, as is usual in MA questions. Check whether the question says sample.
Variance (shortcut form)
σ² = (Σx² ÷ n) − x̄²
Often faster. Mean of the squares minus the square of the mean.
Standard deviation
σ = √variance
In the same units as the data.
Variance for a frequency distribution
σ² = (Σfx² ÷ Σf) − x̄²
Here x̄ = Σfx ÷ Σf. Σf is the total frequency.
Coefficient of variation
CV = (standard deviation ÷ mean) × 100%
Compares relative spread. Lower CV means more consistent data.
Simple price index
Price index = (Current price ÷ Base price) × 100
Works the same for quantity or any other single series.
Laspeyres price index
Σ(P1 × Q0) ÷ Σ(P0 × Q0) × 100
P0 and P1 are base and current prices. Q0 is base-period quantity. Weights stay fixed.
Paasche price index
Σ(P1 × Q1) ÷ Σ(P0 × Q1) × 100
Q1 is current-period quantity. Weights change each period.
Laspeyres quantity index
Σ(Q1 × P0) ÷ Σ(Q0 × P0) × 100
Quantities change, prices held at base. Paasche quantity index uses current prices: Σ(Q1 × P1) ÷ Σ(Q0 × P1) × 100.
Deflating a value
Real value = Current value ÷ Price index × 100
Gives the value in base-period prices.
Rate of change between periods
% change = (Index later − Index earlier) ÷ Index earlier × 100
Do not simply subtract index points unless the earlier index is 100.
Changing the base
New index = Old index ÷ Old index of new base period × 100
The new base year then shows 100.
Additive model
Y = T + S + R
Seasonal variation is an absolute amount. Seasonal variations should sum to zero over a full cycle.
Multiplicative model
Y = T × S × R
Seasonal variation is a factor or percentage. Factors should average 1 (sum to the number of periods in the cycle).
Seasonal variation, additive
S = Actual - Trend
Calculate for each period, then average the figures for the same season across years.
Seasonal variation, multiplicative
S = Actual ÷ Trend
Average the ratios for the same season across years.
Moving average, odd number of periods
Sum of n values ÷ n
Placed against the middle period. Example: 7-day average for daily data.
Centred moving average, even number of periods
(Average 1 + Average 2) ÷ 2, or (½ first + middle values + ½ last) ÷ 4 for quarters
Places the trend against an actual period.
Adjusting seasonal variations (additive)
Adjustment = Total of averages ÷ number of seasons; subtract it from each average
Makes the variations sum to zero.
Forecast, additive
Forecast = Projected trend + Seasonal variation
Project the trend first, using the average change per period.
Forecast, multiplicative
Forecast = Projected trend × Seasonal factor
Use the factor for the correct season.
Seasonally adjusted figure
Additive: Actual - S. Multiplicative: Actual ÷ S
Removes the seasonal effect so you can see the underlying trend.
Correlation coefficient
r = (nΣxy − ΣxΣy) ÷ √[(nΣx² − (Σx)²) × (nΣy² − (Σy)²)]
n is the number of pairs. The answer must lie between −1 and +1. If it does not, you have made an arithmetic error.
Coefficient of determination
r² = r × r
Gives the proportion (or percentage) of variation in y explained by x. Never negative.
Regression line
y = a + bx
y is the dependent variable, x the independent variable. In cost estimation, a is fixed cost and b is variable cost per unit.
Slope (b)
b = (nΣxy − ΣxΣy) ÷ (nΣx² − (Σx)²)
The numerator is the same as in the r formula. Calculate it once and reuse it.
Intercept (a)
a = (Σy − bΣx) ÷ n, or a = ȳ − b x̄
Calculate b first. ȳ and x̄ are the mean values of y and x.
High-low variable cost
b = (cost at highest activity − cost at lowest activity) ÷ (highest activity − lowest activity)
Then fixed cost = total cost at either point − (b × activity at that point). Use the highest and lowest activity levels, not the highest and lowest costs.

Quick revision

  • Primary data is collected first-hand for a purpose; secondary data already exists and was collected by someone else.
  • Random sampling gives each item a known chance of selection; quota and convenience sampling are non-random and can bias results.
  • Mean = Σx ÷ n; the median is the middle value of ordered data; the mode is the most frequent value.
  • The mean uses every value but is affected by extreme values; the median is not.
  • Range = highest value - lowest value; standard deviation is the square root of variance.
  • A larger standard deviation means more spread around the mean.
  • Index = (current value ÷ base value) × 100.
  • Time series: actual = trend + seasonal variation + random variation in the additive model.
  • Forecasts become less reliable the further ahead you project the trend.
  • Correlation coefficient r lies from -1 to +1; it shows association, not cause and effect.
  • Regression line: y = a + bx, where b is the change in y for each one-unit change in x.
  • Linear programming finds the best use of scarce resources subject to constraints.

Common mistakes

  • Treating internal data as always primary and external data as always secondary. Fix: Use two separate tests. Primary or secondary depends on who collected the data and why. Internal or external depends on where the source sits. Last year's sales ledger used for a new forecast is internal but secondary.
  • Calling a measured quantity such as time or weight discrete. Fix: Ask how the value arises. If it is measured on a scale, it is continuous, even if reported to the nearest whole number.
  • Confusing stratified and cluster sampling. Fix: Stratified takes some items from every group. Cluster takes every item from some groups.
  • Calling quota sampling a random method. Fix: In quota sampling the interviewer picks people by judgement. Stratified picks randomly within each group.
  • Plotting raw frequency on a histogram with unequal class widths. Fix: Remember area equals frequency. Divide each frequency by its class width and plot frequency density.
  • Plotting an ogive against the class midpoint or lower boundary. Fix: Cumulative frequency is reached only at the end of a class, so plot it against the upper class boundary.
  • Taking a simple average when a weighted average is needed. Fix: If items have different quantities or importance, multiply by weights and divide by total weights.
  • Finding the median without ordering the data. Fix: Always sort from smallest to largest first.
  • Giving the variance when the question asks for standard deviation, or the other way round. Fix: Underline the measure asked for. If it is standard deviation, finish by taking the square root.
  • Forgetting to square the mean in the shortcut formula. Fix: Write the formula as (Σx² ÷ n) − (x̄)² and calculate x̄² as a separate line.

Exam tips

  • Read the question stem twice. Many marks are lost on one clue word such as 'published' or 'own survey'.
  • In multiple response questions, select exactly the number of answers stated. Extra choices usually score nothing.
  • Keep the two tests apart: primary or secondary is about who collected the data; internal or external is about where the source is.
  • For advantages and disadvantages, think cost, speed, relevance and reliability.
  • Do not spend long on these questions. They are quick marks if you know the definitions.
  • Learn one-line definitions for all six methods. Many objective questions are pure recognition.
  • In multiple response questions, read how many answers to select, and check each statement against the method.
  • For 'which method is most suitable' questions, match to the scenario constraint such as cost, spread or need for representation.