ACCA Applied Knowledge · Management Accounting
Summarising and analysing data: formula sheet
Key formulas
- Primary vs secondary data
- Primary = collected first-hand for this purpose; Secondary = already collected by someone else
- The test is who collected the data and why, not whether it is internal or external.
- Quantitative vs qualitative data
- Quantitative = numerical and measurable; Qualitative = descriptive, based on opinion or quality
- Ratings scored from 1 to 5 are still based on opinion, but the scores are numerical.
- Discrete vs continuous data
- Discrete = counted, separate values; Continuous = measured, any value in a range
- Ask: can the value sensibly be a fraction or decimal? If yes, it is continuous.
- Internal vs external source
- Internal = from within the organisation; External = from outside it
- Internal data can be primary or secondary. Check the wording of the question.
- Systematic sampling interval
- Interval (n) = population size ÷ sample size
- Pick a random start within the first n items, then take every nth item. Round sensibly if the answer is not whole.
- Proportionate stratified sample
- Sample from a stratum = (stratum size ÷ population size) × total sample size
- Use this to allocate the sample across strata. Round so the total still equals the sample size.
- Stratified vs cluster rule
- Stratified: sample from every group. Cluster: sample all items from some groups.
- Strata should differ from each other and be similar inside. Clusters should each resemble the whole population.
- Class width
- Class width = upper boundary − lower boundary
- For 10 – under 20 the width is 10. Use boundaries, not stated limits, when classes have gaps.
- Frequency density
- Frequency density = frequency ÷ class width
- Plot this on the vertical axis of a histogram when class widths are unequal.
- Standard-width adjustment
- Adjusted height = frequency ÷ (class width ÷ standard width)
- An alternative to frequency density. A class twice the standard width has its frequency halved.
- Pie chart angle
- Angle = (category value ÷ total) × 360°
- Angles must add to 360°.
- Cumulative frequency
- Cumulative frequency = running total of frequencies
- Plot against the upper class boundary for an ogive.
- Relative frequency
- Relative frequency = frequency ÷ total frequency
- Can be shown as a fraction, decimal or percentage.
- Median position on an ogive
- Read the value at cumulative frequency = n ÷ 2
- Quartiles are read at n ÷ 4 and 3n ÷ 4.
- Arithmetic mean (raw data)
- x̄ = Σx ÷ n
- Add all values, divide by the number of values.
- Mean (frequency or grouped data)
- x̄ = Σfx ÷ Σf
- For grouped data, x is the class mid-point. The result is an estimate.
- Class mid-point
- Mid-point = (lower limit + upper limit) ÷ 2
- Check class boundaries are continuous before using this.
- Median position (raw data, ordered)
- Position = (n + 1) ÷ 2
- If the position is a half, take the average of the two middle values.
- Median (grouped data)
- Median = L + [(n ÷ 2 − cf) ÷ f] × w
- L = lower boundary of median class, cf = cumulative frequency before it, f = its frequency, w = class width.
- Mode (grouped data)
- Mode = L + [d1 ÷ (d1 + d2)] × w
- L = lower boundary of modal class, d1 = modal frequency minus previous frequency, d2 = modal frequency minus next frequency. Equal class widths assumed.
- Weighted average
- Weighted mean = Σwx ÷ Σw
- w is the weight, such as units, hours or proportion.
- Range
- Range = highest value − lowest value
- Uses only two values, so it is affected by extreme figures.
- Interquartile range
- IQR = Q3 − Q1
- Measures the spread of the middle 50% of the data. Q1 and Q3 are the lower and upper quartiles.
- Variance (population, ungrouped)
- σ² = Σ(x − x̄)² ÷ n
- Use n when the data is treated as the whole population, as is usual in MA questions. Check whether the question says sample.
- Variance (shortcut form)
- σ² = (Σx² ÷ n) − x̄²
- Often faster. Mean of the squares minus the square of the mean.
- Standard deviation
- σ = √variance
- In the same units as the data.
- Variance for a frequency distribution
- σ² = (Σfx² ÷ Σf) − x̄²
- Here x̄ = Σfx ÷ Σf. Σf is the total frequency.
- Coefficient of variation
- CV = (standard deviation ÷ mean) × 100%
- Compares relative spread. Lower CV means more consistent data.
- Simple price index
- Price index = (Current price ÷ Base price) × 100
- Works the same for quantity or any other single series.
- Laspeyres price index
- Σ(P1 × Q0) ÷ Σ(P0 × Q0) × 100
- P0 and P1 are base and current prices. Q0 is base-period quantity. Weights stay fixed.
- Paasche price index
- Σ(P1 × Q1) ÷ Σ(P0 × Q1) × 100
- Q1 is current-period quantity. Weights change each period.
- Laspeyres quantity index
- Σ(Q1 × P0) ÷ Σ(Q0 × P0) × 100
- Quantities change, prices held at base. Paasche quantity index uses current prices: Σ(Q1 × P1) ÷ Σ(Q0 × P1) × 100.
- Deflating a value
- Real value = Current value ÷ Price index × 100
- Gives the value in base-period prices.
- Rate of change between periods
- % change = (Index later − Index earlier) ÷ Index earlier × 100
- Do not simply subtract index points unless the earlier index is 100.
- Changing the base
- New index = Old index ÷ Old index of new base period × 100
- The new base year then shows 100.
- Additive model
- Y = T + S + R
- Seasonal variation is an absolute amount. Seasonal variations should sum to zero over a full cycle.
- Multiplicative model
- Y = T × S × R
- Seasonal variation is a factor or percentage. Factors should average 1 (sum to the number of periods in the cycle).
- Seasonal variation, additive
- S = Actual - Trend
- Calculate for each period, then average the figures for the same season across years.
- Seasonal variation, multiplicative
- S = Actual ÷ Trend
- Average the ratios for the same season across years.
- Moving average, odd number of periods
- Sum of n values ÷ n
- Placed against the middle period. Example: 7-day average for daily data.
- Centred moving average, even number of periods
- (Average 1 + Average 2) ÷ 2, or (½ first + middle values + ½ last) ÷ 4 for quarters
- Places the trend against an actual period.
- Adjusting seasonal variations (additive)
- Adjustment = Total of averages ÷ number of seasons; subtract it from each average
- Makes the variations sum to zero.
- Forecast, additive
- Forecast = Projected trend + Seasonal variation
- Project the trend first, using the average change per period.
- Forecast, multiplicative
- Forecast = Projected trend × Seasonal factor
- Use the factor for the correct season.
- Seasonally adjusted figure
- Additive: Actual - S. Multiplicative: Actual ÷ S
- Removes the seasonal effect so you can see the underlying trend.
- Correlation coefficient
- r = (nΣxy − ΣxΣy) ÷ √[(nΣx² − (Σx)²) × (nΣy² − (Σy)²)]
- n is the number of pairs. The answer must lie between −1 and +1. If it does not, you have made an arithmetic error.
- Coefficient of determination
- r² = r × r
- Gives the proportion (or percentage) of variation in y explained by x. Never negative.
- Regression line
- y = a + bx
- y is the dependent variable, x the independent variable. In cost estimation, a is fixed cost and b is variable cost per unit.
- Slope (b)
- b = (nΣxy − ΣxΣy) ÷ (nΣx² − (Σx)²)
- The numerator is the same as in the r formula. Calculate it once and reuse it.
- Intercept (a)
- a = (Σy − bΣx) ÷ n, or a = ȳ − b x̄
- Calculate b first. ȳ and x̄ are the mean values of y and x.
- High-low variable cost
- b = (cost at highest activity − cost at lowest activity) ÷ (highest activity − lowest activity)
- Then fixed cost = total cost at either point − (b × activity at that point). Use the highest and lowest activity levels, not the highest and lowest costs.
Quick revision
- Primary data is collected first-hand for a purpose; secondary data already exists and was collected by someone else.
- Random sampling gives each item a known chance of selection; quota and convenience sampling are non-random and can bias results.
- Mean = Σx ÷ n; the median is the middle value of ordered data; the mode is the most frequent value.
- The mean uses every value but is affected by extreme values; the median is not.
- Range = highest value - lowest value; standard deviation is the square root of variance.
- A larger standard deviation means more spread around the mean.
- Index = (current value ÷ base value) × 100.
- Time series: actual = trend + seasonal variation + random variation in the additive model.
- Forecasts become less reliable the further ahead you project the trend.
- Correlation coefficient r lies from -1 to +1; it shows association, not cause and effect.
- Regression line: y = a + bx, where b is the change in y for each one-unit change in x.
- Linear programming finds the best use of scarce resources subject to constraints.
Common mistakes
- Treating internal data as always primary and external data as always secondary. Fix: Use two separate tests. Primary or secondary depends on who collected the data and why. Internal or external depends on where the source sits. Last year's sales ledger used for a new forecast is internal but secondary.
- Calling a measured quantity such as time or weight discrete. Fix: Ask how the value arises. If it is measured on a scale, it is continuous, even if reported to the nearest whole number.
- Confusing stratified and cluster sampling. Fix: Stratified takes some items from every group. Cluster takes every item from some groups.
- Calling quota sampling a random method. Fix: In quota sampling the interviewer picks people by judgement. Stratified picks randomly within each group.
- Plotting raw frequency on a histogram with unequal class widths. Fix: Remember area equals frequency. Divide each frequency by its class width and plot frequency density.
- Plotting an ogive against the class midpoint or lower boundary. Fix: Cumulative frequency is reached only at the end of a class, so plot it against the upper class boundary.
- Taking a simple average when a weighted average is needed. Fix: If items have different quantities or importance, multiply by weights and divide by total weights.
- Finding the median without ordering the data. Fix: Always sort from smallest to largest first.
- Giving the variance when the question asks for standard deviation, or the other way round. Fix: Underline the measure asked for. If it is standard deviation, finish by taking the square root.
- Forgetting to square the mean in the shortcut formula. Fix: Write the formula as (Σx² ÷ n) − (x̄)² and calculate x̄² as a separate line.
Exam tips
- Read the question stem twice. Many marks are lost on one clue word such as 'published' or 'own survey'.
- In multiple response questions, select exactly the number of answers stated. Extra choices usually score nothing.
- Keep the two tests apart: primary or secondary is about who collected the data; internal or external is about where the source is.
- For advantages and disadvantages, think cost, speed, relevance and reliability.
- Do not spend long on these questions. They are quick marks if you know the definitions.
- Learn one-line definitions for all six methods. Many objective questions are pure recognition.
- In multiple response questions, read how many answers to select, and check each statement against the method.
- For 'which method is most suitable' questions, match to the scenario constraint such as cost, spread or need for representation.