Skip to content

Financial Management and Business Data Analytics · Data Analysis and Modelling

Descriptive Statistics and Data Summarisation for CMA Inter

Updated 10 October 2026 · Fact-checked

Descriptive statistics summarise a dataset with a few numbers and tables instead of listing every value. You use a measure of central tendency (mean, median, mode) for the typical value, a measure of dispersion (range, standard deviation, coefficient of variation) for spread, and skewness for shape. Compute them step by step, then interpret.

Understand Descriptive Statistics and Data Summarisation

A raw dataset of hundreds of figures tells you nothing at a glance. Descriptive statistics reduce it to a few numbers and a table so that you can describe what the data looks like. They describe the data you have. They do not predict or test anything.

There are three ideas to learn. Central tendency answers 'what is a typical value?' The mean is the arithmetic average. The median is the middle value after sorting. The mode is the most frequent value. Dispersion answers 'how spread out are the values?' The range, variance, standard deviation and coefficient of variation measure this. Shape answers 'is the data symmetric or lopsided?' Skewness measures this.

Why do you need all three? Two branches can both average ₹20 lakh monthly sales. One may sell ₹19-21 lakh every month, while the other swings between ₹5 lakh and ₹35 lakh. The averages match, but the risk does not. The mean is also pulled by extreme values, so the median is a better 'typical' figure when the data has a few very large or very small values.

For large data, you first build a frequency distribution: classes, frequencies, cumulative frequencies and sometimes relative frequencies. You then compute the statistics from class mid-points. The results are estimates because individual values inside a class are not known.

The usual pattern in an exam answer is: organise the data, compute the measures, then comment. For a right-skewed (positively skewed) distribution, the mean is above the median, and the median is above the mode. For a left-skewed (negatively skewed) distribution, the order reverses. In a symmetric distribution, the three are close or equal.

Key rules to remember

Arithmetic mean (ungrouped)
x̄ = Σx ÷ n
Add all values and divide by the number of values.
Arithmetic mean (grouped)
x̄ = Σfx ÷ Σf
x is the class mid-point, f is the frequency. Mid-point = (lower limit + upper limit) ÷ 2.
Weighted mean
x̄w = Σwx ÷ Σw
Use when values carry different importance, such as prices with quantities.
Median (ungrouped)
Sort the data. Odd n: middle value. Even n: average of the two middle values.
Position of the middle value for odd n is (n + 1) ÷ 2.
Median (grouped)
Median = L + [(N ÷ 2 − cf) ÷ f] × h
L = lower limit of median class, N = Σf, cf = cumulative frequency before the median class, f = its frequency, h = class width. The median class is the first class whose cumulative frequency reaches N ÷ 2.
Mode (grouped)
Mode = L + [(f1 − f0) ÷ (2f1 − f0 − f2)] × h
f1 = frequency of modal class, f0 = frequency of class before it, f2 = frequency of class after it. Classes must have equal width.
Empirical relation
Mode ≈ 3 × Median − 2 × Mean
An approximation for moderately skewed data, not an exact law. Use it only when the question asks or the mode is missing.
Range and coefficient of range
Range = Largest − Smallest; Coefficient of range = (Largest − Smallest) ÷ (Largest + Smallest)
Quick but depends only on the two extreme values.
Quartile deviation
QD = (Q3 − Q1) ÷ 2
Based on the middle 50% of the data, so it is not affected by extreme values.
Variance and standard deviation (population)
σ² = Σ(x − x̄)² ÷ n; σ = √σ²; grouped: σ² = Σf(x − x̄)² ÷ Σf
Use this when the data covers the whole group. For a sample, divide by (n − 1) instead of n.
Coefficient of variation
CV = (σ ÷ x̄) × 100
Relative spread in percentage. The series with the lower CV is more consistent.
Karl Pearson's coefficient of skewness
Sk = (Mean − Mode) ÷ σ
Positive means right-skewed, negative means left-skewed, near zero means roughly symmetric.
Bowley's coefficient of skewness
Sk = (Q3 + Q1 − 2 × Median) ÷ (Q3 − Q1)
Quartile-based, useful when extreme values distort the mean and standard deviation.
Relative frequency
Relative frequency = f ÷ Σf
Multiply by 100 for a percentage. All relative frequencies add up to 1 (or 100%).

How to solve Descriptive Statistics and Data Summarisation questions

Use this order for any descriptive statistics question, whether the data is a short list or a frequency table.

  1. 1Read what is asked. Note which measures are required and whether the data is a sample or the whole group.
  2. 2Organise the data. Sort a raw list in ascending order. For a frequency table, add columns for mid-point (x), fx, and cumulative frequency as needed.
  3. 3Compute the central tendency measures asked for. Show the formula and the substitution.
  4. 4Compute dispersion. Find the deviations from the mean, square them, then divide and take the square root. Keep a column for f(x − x̄)² in grouped data.
  5. 5Compute relative measures (CV or skewness) if the question compares two series or asks about shape.
  6. 6Check reasonableness. The mean must lie between the smallest and largest value. The median class must contain N ÷ 2. The standard deviation cannot be negative.
  7. 7Interpret in one or two sentences: typical value, consistency, and the direction of skew. Link it to the business context.
  8. 8Write the final answer with units (₹, units, days) and a clear label for each measure.

Quickest way: Table-first method for exam speed

When to use it: Use this for grouped data or any question that asks for several measures at once, where marks go to a clean table.

  1. Draw one table with columns: class, f, x, fx, cumulative f, and (x − x̄)² × f if SD is needed.
  2. Fill the mid-point and fx columns and total them. Get the mean first, because every later column depends on it.
  3. Read the median class from the cumulative column and the modal class from the highest frequency. Substitute into the formulas right away.
  4. For SD, compute the deviations only after the mean is final. Then do Σf(x − x̄)² ÷ Σf and take the root.
  5. For a comparison of two series, compute only the mean, SD and CV, and state the lower CV as more consistent.
  6. Write one line of interpretation at the end. Do not skip it.

Common mistakes in Descriptive Statistics and Data Summarisation

  • Taking the mean of grouped data as Σx ÷ number of classes.

    Students forget that each class stands for many observations.

    Fix: Always weight by frequency: Σfx ÷ Σf. Check that Σf equals the total observations.

  • Finding the median without sorting the data first.

    The raw list looks ready to use, so the middle position is read from the unsorted order.

    Fix: Sort ascending before picking the middle value. For an even count, average the two middle values.

  • Dividing by n − 1 when the question treats the data as the whole group, or by n when it says sample.

    Students memorise one version of the variance formula.

    Fix: Read the wording. Whole group or complete record: divide by n. Sample drawn from a larger group: divide by n − 1. If unspecified, state your assumption.

  • Declaring the series with the higher standard deviation as more variable when the means differ.

    Standard deviation is in the units of the data, so it depends on the size of the numbers.

    Fix: Compare using the coefficient of variation. The lower CV means more consistent.

  • Using the modal class formula with the wrong neighbouring frequencies, or with unequal class widths.

    f0 and f2 get swapped or the highest frequency is read from the wrong class.

    Fix: Mark f1 (highest), f0 (just above in the table) and f2 (just below) before substituting. If class widths are unequal, adjust the frequencies first or state the limitation.

  • Stopping at the numbers with no interpretation.

    Students treat the calculation as the whole answer.

    Fix: Add a line such as 'Mean exceeds median, so the data is right-skewed and a few high values pull the average up.'

Worked examples

Example 1

A company recorded monthly sales (₹ lakh) for eight months: 12, 15, 15, 18, 20, 22, 28, 30. Treating these as the complete record, find the mean, median, mode, range, standard deviation, coefficient of variation and Karl Pearson's coefficient of skewness. Comment on the result.

Show the solution
  1. Data is already in ascending order. n = 8.
  2. Mean = Σx ÷ n = (12 + 15 + 15 + 18 + 20 + 22 + 28 + 30) ÷ 8 = 160 ÷ 8 = ₹20 lakh.
  3. Median: the two middle values are the 4th and 5th, which are 18 and 20. Median = (18 + 20) ÷ 2 = ₹19 lakh.
  4. Mode: 15 occurs twice and every other value once, so mode = ₹15 lakh.
  5. Range = 30 − 12 = ₹18 lakh.
  6. Deviations from 20: −8, −5, −5, −2, 0, 2, 8, 10. Squares: 64, 25, 25, 4, 0, 4, 64, 100. Sum = 286.
  7. Variance = 286 ÷ 8 = 35.75. Standard deviation = √35.75 = 5.98 (approx.), so ₹5.98 lakh.
  8. CV = (5.98 ÷ 20) × 100 = 29.9% (approx.).
  9. Karl Pearson's skewness = (Mean − Mode) ÷ σ = (20 − 15) ÷ 5.98 = 0.84 (approx.).
  10. Comment: mean (20) > median (19) > mode (15) and the skewness is positive. The distribution is moderately right-skewed because a few high-sales months pull the average up. A CV near 30% shows noticeable month-to-month variation.

Answer: Mean ₹20 lakh; median ₹19 lakh; mode ₹15 lakh; range ₹18 lakh; SD ≈ ₹5.98 lakh; CV ≈ 29.9%; skewness ≈ +0.84 (positively skewed).

Example 2

The daily wages (₹ hundred) of 50 workers are: 0-10: 5 workers; 10-20: 10; 20-30: 15; 30-40: 12; 40-50: 8. Find the mean, median and mode, and comment on the shape.

Show the solution
  1. Mid-points x: 5, 15, 25, 35, 45. Frequencies f: 5, 10, 15, 12, 8. Σf = 50.
  2. fx: 25, 150, 375, 420, 360. Σfx = 1,330.
  3. Mean = Σfx ÷ Σf = 1,330 ÷ 50 = 26.6.
  4. Cumulative frequencies: 5, 15, 30, 42, 50. N ÷ 2 = 25, which first falls in the 20-30 class, so the median class is 20-30.
  5. Median = L + [(N ÷ 2 − cf) ÷ f] × h = 20 + [(25 − 15) ÷ 15] × 10 = 20 + 6.67 = 26.67 (approx.).
  6. Modal class is the one with the highest frequency (15): 20-30. So f1 = 15, f0 = 10, f2 = 12, L = 20, h = 10.
  7. Mode = 20 + [(15 − 10) ÷ (2 × 15 − 10 − 12)] × 10 = 20 + (5 ÷ 8) × 10 = 20 + 6.25 = 26.25.
  8. Comment: mean 26.6, median 26.67 and mode 26.25 are very close, so the distribution is almost symmetric.

Answer: Mean = 26.6, median ≈ 26.67, mode = 26.25 (all in ₹ hundred). The distribution is nearly symmetric.

Exam tips

  • In the MCQ section, many questions are one-step: identify the measure, sort the data, or read the effect of an outlier. Remember that an extreme value moves the mean most, the median little, and the mode usually not at all. There is no negative marking, so always attempt every MCQ.
  • For written answers, draw the working table first. Step marks usually go for the correct table, formula substitution and final interpretation, even if a small arithmetic slip occurs.
  • When two series are compared, calculate CV and name the more consistent one. Do not stop at the standard deviation.
  • State your assumption (population or sample) in one line when the question is silent, then stay consistent.
  • Link the shape to the business meaning, such as 'a few very large orders raise the average above the typical order'. Examiners reward a short, relevant comment.

Practice questions from Data Analysis and Modelling

Descriptive Statistics and Data Summarisation: frequently asked questions

What is the difference between mean, median and mode?

The mean is the arithmetic average of all values. The median is the middle value after sorting, and the mode is the most frequent value. Use the median when extreme values distort the average, and the mode for the most common category or size.

Which measure of dispersion should I use?

Use the standard deviation for most numerical work, because it uses every value. Use the range for a quick check and the quartile deviation when extreme values are present. Use the coefficient of variation to compare series with different means or units.

How do I know if data is skewed?

Compare the mean, median and mode. If the mean is above the median, which is above the mode, the data is right-skewed. If the order is reversed, it is left-skewed. You can also compute Karl Pearson's coefficient: (Mean − Mode) ÷ standard deviation.

Is the median of grouped data exact?

No. The grouped-data formula assumes the values are spread evenly within the median class, so the answer is an estimate. The same applies to the grouped mean and mode, which use mid-points and class frequencies.