CA Foundation · Quantitative Aptitude
Measures of Central Tendency and Dispersion: formula sheet
Key formulas
- Simple AM (ungrouped)
- x̄ = Σx ÷ n
- Equivalent to Σx = n × x̄. Use this to move between mean and total.
- AM of a frequency distribution
- x̄ = Σfx ÷ Σf
- For class intervals, x is the mid-point of the class.
- Weighted AM
- x̄w = Σwx ÷ Σw
- w is the weight of each value x.
- Combined AM of two groups
- x̄ = (n₁x̄₁ + n₂x̄₂) ÷ (n₁ + n₂)
- Extends to more groups by adding terms in numerator and denominator.
- Step-deviation method
- d = (x − A) ÷ h; x̄ = A + h × (Σfd ÷ Σf)
- A is an assumed mean and h is a common factor of the class width or differences.
- Corrected mean
- Correct Σx = Wrong Σx − wrong value + correct value
- Wrong Σx = n × wrong mean. Divide the correct total by n (or the new n if items are added or removed).
- Properties of AM
- Σ(x − x̄) = 0; if y = a + bx then ȳ = a + b·x̄
- Adding, subtracting, multiplying or dividing every item changes the mean in the same way. The mean is affected by extreme values.
- Median, ungrouped data
- Median = value of the (n + 1) ÷ 2 th item in sorted data
- If n is even, the position ends in .5. Take the average of the two middle items.
- Quartiles, ungrouped data
- Qi = value of the i(n + 1) ÷ 4 th item, i = 1, 2, 3
- If the position is fractional, take the lower item plus that fraction of the gap to the next item.
- Deciles and percentiles, ungrouped data
- Di = value of the i(n + 1) ÷ 10 th item; Pi = value of the i(n + 1) ÷ 100 th item
- Sort the data first. Use the same interpolation for fractional positions.
- Median, grouped data
- Median = L + [(N ÷ 2 − cf) ÷ f] × h
- L = lower boundary of median class, N = total frequency, cf = cumulative frequency of the class before it, f = frequency of the median class, h = class width. Classes must be continuous.
- Quartile, decile, percentile, grouped data
- Qi = L + [(iN ÷ 4 − cf) ÷ f] × h; Di = L + [(iN ÷ 10 − cf) ÷ f] × h; Pi = L + [(iN ÷ 100 − cf) ÷ f] × h
- Same formula as the median. Only the target position changes. L, cf, f and h refer to the class that contains that position.
- Equivalences
- Q2 = D5 = P50 = Median; Q1 = P25; Q3 = P75; Di = P(10i)
- Use these to convert one question into another type.
- Mode for ungrouped data
- Mode = value with the highest frequency
- There can be no mode, one mode or several modes.
- Mode for grouped data
- Mode = L + (f1 − f0) ÷ (2f1 − f0 − f2) × h
- L = lower limit of modal class, f1 = frequency of modal class, f0 = frequency of class before it, f2 = frequency of class after it, h = class width. Classes must be continuous (exclusive) with equal width.
- Empirical relation
- Mode = 3 Median − 2 Mean
- Approximate, for moderately skewed distributions.
- Empirical relation (alternate form)
- Mean − Mode = 3 (Mean − Median)
- Same relation rearranged. Median = (Mode + 2 Mean) ÷ 3 and Mean = (3 Median − Mode) ÷ 2.
- Inclusive to exclusive classes
- Adjustment = (next lower limit − previous upper limit) ÷ 2
- Subtract it from every lower limit and add it to every upper limit. For 10-19, 20-29 the adjustment is 0.5, so the classes become 9.5-19.5, 19.5-29.5.
- Variance (ungrouped)
- σ² = Σ(x − x̄)² ÷ N
- N is the number of observations. CA Foundation questions normally divide by N.
- Shortcut form of variance
- σ² = Σx² ÷ N − (x̄)²
- Use this when the mean is not a whole number. It avoids many subtractions.
- Standard deviation
- σ = √variance
- Always the positive root.
- Grouped data (frequency distribution)
- σ² = Σf(x − x̄)² ÷ N = Σfx² ÷ N − (Σfx ÷ N)²
- N = Σf. For classes, x is the class mark (mid-value).
- Step deviation method
- d = (x − A) ÷ h; σ = h × √[Σfd² ÷ N − (Σfd ÷ N)²]
- A is the assumed mean and h is the common class width. Multiply by h at the end for SD, and by h² for variance.
- Change of origin and scale
- If y = a + bx, then σy = |b| × σx and variance of y = b² × variance of x
- a (the origin shift) has no effect. A negative b does not make SD negative.
- Combined mean
- x̄₁₂ = (n₁x̄₁ + n₂x̄₂) ÷ (n₁ + n₂)
- Find this first for the combined SD.
- Combined standard deviation
- σ₁₂ = √[(n₁σ₁² + n₂σ₂² + n₁d₁² + n₂d₂²) ÷ (n₁ + n₂)], where d₁ = x̄₁ − x̄₁₂ and d₂ = x̄₂ − x̄₁₂
- Works for two groups. If the two means are equal, d₁ = d₂ = 0.
- First n natural numbers
- Variance of 1, 2, 3, …, n = (n² − 1) ÷ 12
- The same holds for any n equally spaced values with common difference 1.
- Coefficient of variation
- CV = (σ ÷ x̄) × 100
- Used to compare the spread of two series. A lower CV means more consistency.
- Coefficient of variation
- CV = (σ ÷ x̄) × 100
- σ is the standard deviation, x̄ is the arithmetic mean. The result is a percentage. Use the mean as a positive value.
- Variance to SD
- σ = √Variance
- If a question gives variance, take the square root before computing CV.
- Consistency rule
- Smaller CV ⇒ more consistent; larger CV ⇒ more variable
- Compare CVs of the two series only. Do not compare SDs when means differ.
- Finding SD from CV
- σ = CV × x̄ ÷ 100
- Use this when CV and mean are given and SD is asked.
- Effect of change of origin and scale
- For y = a + bx: ȳ = a + b·x̄ and σy = |b|·σx
- Adding a constant does not change SD, but it changes the mean, so CV changes.
- Karl Pearson's coefficient (with mode)
- Sk = (Mean − Mode) ÷ SD
- Use when the mode is given or well defined. SD is the standard deviation, not the variance.
- Karl Pearson's coefficient (with median)
- Sk = 3 × (Mean − Median) ÷ SD
- Use when the mode is ill-defined or not given. It follows from Mode = 3 Median − 2 Mean.
- Empirical relation among averages
- Mode = 3 Median − 2 Mean
- Holds approximately for moderately skewed distributions. It is not an exact law.
- Bowley's coefficient
- Sk = (Q3 + Q1 − 2 × Median) ÷ (Q3 − Q1)
- Always lies between −1 and +1. It depends only on the middle 50% of the data.
- Order of averages
- Positive skew: Mean > Median > Mode. Negative skew: Mean < Median < Mode. Symmetric: Mean = Median = Mode
- Use this to check the sign of your answer.
- Reading the sign
- Sk > 0: positive skew. Sk < 0: negative skew. Sk = 0: symmetric
- The sign is decided by the numerator in both coefficients.
Quick revision
- Mean of grouped data: Σfx ÷ Σf; with step deviation, mean = A + h × (Σfd ÷ Σf).
- Median lies at the (N ÷ 2)th item in a continuous distribution; use cumulative frequency to find the class.
- Mode (grouped) = L + [(f1 − f0) ÷ (2f1 − f0 − f2)] × h, with f1 the modal class frequency.
- Empirical relation for moderately skewed data: Mode ≈ 3 Median − 2 Mean.
- For positive values that are not all equal: AM > GM > HM; they are equal only when all values are equal.
- For two positive values: GM² = AM × HM.
- Use GM for growth rates and ratios; use HM for averaging rates like speed over equal distances.
- Range = Largest − Smallest; Quartile Deviation = (Q3 − Q1) ÷ 2.
- Variance = (SD)²; SD = √[Σfx² ÷ Σf − (mean)²].
- Changing origin does not change SD or range; multiplying by k multiplies SD and range by |k|.
- CV = (SD ÷ Mean) × 100; the lower the CV, the more consistent the data.
- Karl Pearson's skewness = (Mean − Mode) ÷ SD; positive means a longer right tail.
Common mistakes
- Averaging the group means directly in a combined mean problem. Fix: Always weight by group size: (n₁x̄₁ + n₂x̄₂) ÷ (n₁ + n₂).
- Using class limits or class widths instead of mid-points in grouped data. Fix: Write a separate x column with mid-point = (lower + upper) ÷ 2 before doing anything else.
- Not sorting ungrouped data before finding the median or a quartile. Fix: Always rewrite the data in ascending order first. Do it even if the data looks nearly sorted.
- Using (n + 1) in the grouped formula, or using N ÷ 2 as a position in ungrouped data. Fix: Ungrouped data uses i(n + 1) ÷ parts. Grouped data uses iN ÷ parts inside the interpolation formula.
- Using the highest frequency number as the mode of a grouped table Fix: The highest frequency only identifies the modal class. Then apply the formula to get the mode.
- Not converting inclusive classes to exclusive classes Fix: Subtract 0.5 from lower limits and add 0.5 to upper limits first. Then L is 29.5, not 30, for the class 30-39.
- Giving the variance when SD is asked, or the reverse. Fix: Circle the word variance or SD in the question. Take the square root at the end only when SD is asked.
- Multiplying by h instead of h² for variance in the step deviation method. Fix: SD scales by h. Variance scales by h². Decide which you have before multiplying.
- Choosing the series with the larger CV as more consistent. Fix: Remember that CV measures variation. Less variation means more consistency, so pick the smaller CV.
- Comparing standard deviations instead of CVs when the means differ. Fix: Whenever the means are different, compute CV. Only use SD alone if the means and units are the same.
Exam tips
- Convert every mean into a total first. Most questions then become one line of arithmetic.
- In combined mean questions, check that your answer lies between the group means and nearer the larger group before marking it.
- Expect corrected-mean questions with two errors, or with an item added or removed. Update both the total and the count.
- For grouped data, a smart choice of A and h can cut your calculation time. Always recheck the sign of d for classes below A.
- Avoid guessing when you cannot eliminate at least two options, because each wrong answer costs 0.25 marks.
- Look at the options before calculating. In grouped data, finding the class that holds the target position usually removes two or three options.
- Read the question for the word used: quartile, decile or percentile. Write the target position (such as 3N ÷ 4 or 8N ÷ 10) at the top before touching the table.
- Check whether the classes are inclusive or continuous. Questions sometimes use inclusive classes to test boundary correction.