Actuarial Statistics · Exploratory data analysis
Numerical Summary Measures: Location, Spread, Skewness and Kurtosis
Updated 11 October 2026 · Fact-checked
Numerical summary measures describe a data set with a few numbers. Location (mean, median, mode) gives the centre. Spread (variance, standard deviation, IQR) gives the variability. Skewness gives asymmetry and kurtosis gives tail heaviness. To solve a question, identify the data type, apply the right formula, and state the convention you use.
Understand Numerical Summary Measures
A summary measure replaces a whole data set with one number that describes one feature of it. In CS1 you meet four features: where the data sit, how spread out they are, how lopsided they are, and how heavy the tails are.
Location tells you the centre. The mean is the sum divided by the count. It uses every value, so it is pulled by outliers. The median is the middle value once the data are ordered. It ignores how extreme the extremes are, so it is robust. The mode is the most frequent value, or the most frequent class for grouped data.
Spread tells you how far values sit from the centre. The variance is the average squared distance from the mean. The standard deviation is its square root, in the same units as the data. Quartiles split the ordered data into four parts. The interquartile range (IQR) is Q3 − Q1 and covers the middle half of the data. Percentiles are the same idea with 100 parts.
Skewness measures asymmetry. Positive skew means a long right tail, and usually mean > median. Negative skew means a long left tail. Claim sizes are typically positively skewed. Kurtosis measures how heavy the tails are compared with the peak. A normal distribution has kurtosis 3. Higher values mean more extreme observations. Skewness is about direction. Kurtosis is about tail weight.
Examiners want you to choose a sensible measure and say why. For skewed data with outliers, the median and IQR describe the data better than the mean and standard deviation.
Key rules to remember
- Sample mean
- x̄ = Σx ÷ n
- For grouped data use x̄ = Σfx ÷ Σf, with x as the class midpoint. The result is an estimate.
- Sample variance
- s² = Σ(x − x̄)² ÷ (n − 1) = (Σx² − n x̄²) ÷ (n − 1)
- Divide by n instead of n − 1 only if the question says population variance or divisor n. Standard deviation s = √s².
- Sum of squares
- Sxx = Σx² − (Σx)² ÷ n
- Then s² = Sxx ÷ (n − 1). The shortcut avoids calculating each deviation.
- Median and quartile positions (ungrouped)
- Median at (n + 1)/2; Q1 at (n + 1)/4; Q3 at 3(n + 1)/4; pth percentile at p(n + 1)/100
- If the position is not a whole number, interpolate between the two neighbouring ordered values. Say which convention you use.
- Interquartile range
- IQR = Q3 − Q1
- Measures the spread of the middle 50% of the data.
- Grouped percentile by interpolation
- Value = L + ((target − F) ÷ f) × w
- L = lower class boundary, F = cumulative frequency before the class, f = class frequency, w = class width. Target is n/4, n/2 or 3n/4 for quartiles and median.
- Coefficient of variation
- CV = s ÷ x̄
- Compares spread relative to the mean. Use only when the mean is positive and meaningful.
- Coefficient of skewness (moment form)
- Skewness = [Σ(x − x̄)³ ÷ n] ÷ σ³, where σ² = Σ(x − x̄)² ÷ n
- For a distribution: E[(X − μ)³] ÷ σ³. Positive means right-skewed. Divisors vary between texts, so state yours.
- Pearson's skewness
- 3 × (mean − median) ÷ standard deviation
- A quick approximate measure. It is not the moment-based coefficient.
- Kurtosis
- Kurtosis = [Σ(x − x̄)⁴ ÷ n] ÷ σ⁴; excess kurtosis = kurtosis − 3
- The normal distribution has kurtosis 3 and excess kurtosis 0.
How to solve Numerical Summary Measures questions
Use this routine for any question on numerical summaries. It works for raw data, frequency tables and grouped data.
- 1Read what is asked and note the convention given: sample or population variance, and the method for quartiles.
- 2Identify the data form: raw list, frequency table or grouped classes. For grouped data, find midpoints and cumulative frequencies.
- 3Order raw data before finding the median, quartiles or percentiles. Write the position formula, then locate the value and interpolate if needed.
- 4Compute the mean first. Then compute Σx² (or Σfx²) and use s² = (Σx² − n x̄²) ÷ (n − 1).
- 5For skewness or kurtosis, tabulate (x − x̄), then its cube or fourth power. Divide by n, then by σ³ or σ⁴.
- 6Check the answer for sense: Q1 ≤ median ≤ Q3, variance is not negative, and the sign of skewness matches mean versus median.
- 7Interpret in words and give units. Say what the result implies about the shape or risk of the data.
Quickest way: Shortcut with sums and cumulative frequencies
When to use it: Use this when you have a frequency table or many data points and little time. It also suits a calculator in statistics mode.
- Put the data into the calculator's statistics mode if the exam allows it. Read n, x̄ and the sample standard deviation directly. Check you pick the sample version (n − 1).
- If working by hand, build one table with columns f, x, fx and fx². Total each column once and use s² = (Σfx² − n x̄²) ÷ (n − 1).
- For quartiles from grouped data, add one cumulative frequency column. Find the class where the target n/4, n/2 or 3n/4 first falls inside, then interpolate once.
- For skewness direction, compare mean and median. Mean above median points to positive skew. Use this as a sense check, not as the answer to a calculation.
Common mistakes in Numerical Summary Measures
Dividing by n instead of n − 1 for sample variance.
The calculator or an earlier course used the population formula, or the student forgets the question says sample.
Fix: Default to n − 1 for a sample variance. Use n only when the question says population or gives that definition. Write the divisor in your working.
Finding the median or quartiles from unordered data.
The data list is long and the student rushes to the formula.
Fix: Always sort first. Then apply the position formula, such as (n + 1)/4 for Q1, and interpolate if the position is fractional.
Mixing up the two quartile conventions, (n + 1)/4 for ungrouped data and n/4 for grouped interpolation.
Both are called 'the quartile' and look similar.
Fix: For raw ordered data use p(n + 1)/100 positions. For grouped data use the interpolation formula with target n/4, n/2 or 3n/4. State the method in your answer.
Using class boundaries or frequencies instead of midpoints for the grouped mean and variance.
The student forgets that each class is represented by its midpoint.
Fix: Add a midpoint column first. Remember grouped results are estimates and say so.
Confusing skewness with kurtosis.
Both use standardised higher moments, the third and the fourth.
Fix: Skewness uses the cube, so its sign gives direction. Kurtosis uses the fourth power, so it is always positive and measures tail weight against 3 for the normal.
Saying a positive skewness always means mean > median.
It is a common rule of thumb, and students treat it as a law.
Fix: Say 'usually' or 'typically'. The rule can fail for some distributions, so base the conclusion on the calculated coefficient.
Worked examples
Example 1
A sample of 8 claim amounts (in ₹ thousand) is: 3, 5, 5, 7, 8, 10, 12, 14. Calculate the mean, median, sample variance, lower quartile, upper quartile, IQR and the moment coefficient of skewness (using divisor n).
Show the solution
- Mean: Σx = 3 + 5 + 5 + 7 + 8 + 10 + 12 + 14 = 64, so x̄ = 64 ÷ 8 = 8.
- Median: position (8 + 1)/2 = 4.5, so median = (7 + 8) ÷ 2 = 7.5.
- Deviations from the mean: −5, −3, −3, −1, 0, 2, 4, 6. Squares: 25, 9, 9, 1, 0, 4, 16, 36, summing to 100.
- Sample variance: s² = 100 ÷ 7 = 14.29 (to 2 d.p.).
- Q1: position 9/4 = 2.25. The 2nd and 3rd values are both 5, so Q1 = 5.
- Q3: position 27/4 = 6.75. The 6th value is 10 and the 7th is 12, so Q3 = 10 + 0.75 × 2 = 11.5. IQR = 11.5 − 5 = 6.5.
- Skewness: cubes of deviations are −125, −27, −27, −1, 0, 8, 64, 216, summing to 108. Mean cube = 108 ÷ 8 = 13.5.
- σ² = 100 ÷ 8 = 12.5, so σ³ = 12.5 × √12.5 = 12.5 × 3.5355 = 44.19. Skewness = 13.5 ÷ 44.19 = 0.31.
Answer: Mean = 8, median = 7.5, s² ≈ 14.29, Q1 = 5, Q3 = 11.5, IQR = 6.5, skewness ≈ 0.31 (mild positive skew, consistent with mean > median).
Example 2
Policy sizes for 40 policies are grouped as: 0–10: 4; 10–20: 10; 20–30: 16; 30–40: 8; 40–50: 2 (units of ₹ thousand). Estimate the mean, median, lower and upper quartiles, the IQR and the sample variance.
Show the solution
- Midpoints are 5, 15, 25, 35, 45. Cumulative frequencies are 4, 14, 30, 38, 40.
- Σfx = 4×5 + 10×15 + 16×25 + 8×35 + 2×45 = 20 + 150 + 400 + 280 + 90 = 940. Mean = 940 ÷ 40 = 23.5.
- Median: target n/2 = 20 falls in the 20–30 class (F = 14, f = 16). Median = 20 + (20 − 14) ÷ 16 × 10 = 20 + 3.75 = 23.75.
- Q1: target n/4 = 10 falls in the 10–20 class (F = 4, f = 10). Q1 = 10 + (10 − 4) ÷ 10 × 10 = 16.
- Q3: target 3n/4 = 30 falls in the 20–30 class (F = 14, f = 16). Q3 = 20 + (30 − 14) ÷ 16 × 10 = 30.
- IQR = 30 − 16 = 14.
- Σfx² = 4×25 + 10×225 + 16×625 + 8×1225 + 2×2025 = 100 + 2250 + 10000 + 9800 + 4050 = 26,200.
- n x̄² = 40 × 23.5² = 40 × 552.25 = 22,090. So Σfx² − n x̄² = 4,110.
- Sample variance = 4,110 ÷ 39 = 105.38 (to 2 d.p.).
Answer: Estimated mean ₹23.5 thousand, median ₹23.75 thousand, Q1 ₹16 thousand, Q3 ₹30 thousand, IQR ₹14 thousand, sample variance ≈ 105.38 (₹ thousand)². All values are estimates because the data are grouped.
Exam tips
- Write the convention you use for quartiles and variance in one line. Markers then follow your method even if a different convention gives a slightly different number.
- In MCQs, check the sign first. Mean above median suggests positive skew, and a variance can never be negative. This removes wrong options quickly.
- For 'comment on' questions, name the measure, give the number, then interpret it. For example: skewness 0.31 shows a mild right tail, so a few large claims lift the mean.
- Use Σx² − n x̄² to save time, but keep extra decimal places in x̄ to avoid rounding errors in the variance.
- In the computer-based paper, state the function you use and whether it uses divisor n or n − 1. Many software defaults for variance and standard deviation use n − 1, and quantile functions offer several methods.
Practice questions from Exploratory data analysis
- A scatterplot of monthly premium income (x, in ₹ lakh) against claims (y, in ₹ lakh) for a Mumbai insurer gives r = 0.80. Every x value is t…
- For 20 policies, the sum of claim amounts is Rs 400 thousand and the sum of squares of claims is 8,500 (Rs thousand squared). What is the sa…
- For a sample of 11 policy sizes, the ordered values give a lower quartile of ₹4 lakh, a median of ₹7 lakh and an upper quartile of ₹12 lakh.…
- An insurer's motor claims dataset contains the following fields: (i) vehicle registration state, (ii) number of claims in the last year, (ii…
- For paired data, Σx = 30, Σy = 60, Σx² = 200, Σy² = 460, Σxy = 380, n = 5. What is the sample correlation coefficient?
Numerical Summary Measures: frequently asked questions
How do I calculate quartiles and the IQR from grouped data?
Build cumulative frequencies. Find the class containing n/4 for Q1 and 3n/4 for Q3. Then use L + ((target − F) ÷ f) × w. The IQR is Q3 − Q1. The answer is an estimate because the exact values are lost in grouping.
What is the difference between skewness and kurtosis?
Skewness measures asymmetry. Its sign shows whether the longer tail is on the right or left. Kurtosis measures how heavy the tails are compared with the centre, with 3 as the normal benchmark. A data set can be symmetric yet have high kurtosis.
What is the coefficient of skewness formula in CS1?
The moment form is the average of (x − x̄)³ divided by σ³, where σ is the standard deviation. For a distribution it is E[(X − μ)³] ÷ σ³. Pearson's approximate form is 3 × (mean − median) ÷ standard deviation. Check which one the question asks for.
When should I use the median and IQR instead of the mean and standard deviation?
Use them when the data are skewed or contain outliers, such as insurance claim amounts. The median and IQR are not much affected by a few extreme values. The mean and standard deviation are pulled towards them.
Do I divide by n or n − 1 for variance?
For a sample variance used to estimate the population variance, divide by n − 1. Use n only when the question defines the population variance or the moment coefficients with divisor n. Always state which you used.