Skip to content

IAI Actuarial Core Principles · Actuarial Statistics

Exploratory data analysis: formula sheet

Full chapter guide

Key formulas

Classification of data
Data → Categorical (nominal, ordinal) or Numerical (discrete, continuous)
Ask: is it a label, a count, or a measurement? Ordinal data has order but gaps between categories are not meaningful.
Outlier rule of thumb (IQR rule)
Flag x if x < Q1 − 1.5 × IQR or x > Q3 + 1.5 × IQR, where IQR = Q3 − Q1
A convention for flagging only, not proof of an error. Always investigate before removing.
Missing data types
MCAR, MAR, MNAR
Missing completely at random, missing at random (depends on observed data), missing not at random (depends on the missing value itself). MNAR is the most likely to bias results.
Proportion missing
Proportion missing = (number of missing values) ÷ (number of records)
Compute it per variable. A high proportion limits what any treatment can fix.
Sample mean
x̄ = Σx ÷ n
For grouped data use x̄ = Σfx ÷ Σf, with x as the class midpoint. The result is an estimate.
Sample variance
s² = Σ(x − x̄)² ÷ (n − 1) = (Σx² − n x̄²) ÷ (n − 1)
Divide by n instead of n − 1 only if the question says population variance or divisor n. Standard deviation s = √s².
Sum of squares
Sxx = Σx² − (Σx)² ÷ n
Then s² = Sxx ÷ (n − 1). The shortcut avoids calculating each deviation.
Median and quartile positions (ungrouped)
Median at (n + 1)/2; Q1 at (n + 1)/4; Q3 at 3(n + 1)/4; pth percentile at p(n + 1)/100
If the position is not a whole number, interpolate between the two neighbouring ordered values. Say which convention you use.
Interquartile range
IQR = Q3 − Q1
Measures the spread of the middle 50% of the data.
Grouped percentile by interpolation
Value = L + ((target − F) ÷ f) × w
L = lower class boundary, F = cumulative frequency before the class, f = class frequency, w = class width. Target is n/4, n/2 or 3n/4 for quartiles and median.
Coefficient of variation
CV = s ÷ x̄
Compares spread relative to the mean. Use only when the mean is positive and meaningful.
Coefficient of skewness (moment form)
Skewness = [Σ(x − x̄)³ ÷ n] ÷ σ³, where σ² = Σ(x − x̄)² ÷ n
For a distribution: E[(X − μ)³] ÷ σ³. Positive means right-skewed. Divisors vary between texts, so state yours.
Pearson's skewness
3 × (mean − median) ÷ standard deviation
A quick approximate measure. It is not the moment-based coefficient.
Kurtosis
Kurtosis = [Σ(x − x̄)⁴ ÷ n] ÷ σ⁴; excess kurtosis = kurtosis − 3
The normal distribution has kurtosis 3 and excess kurtosis 0.
Frequency density
frequency density = class frequency ÷ class width
Use this as bar height in a histogram with unequal class widths. Area of bar = class frequency.
Histogram area rule
bar area = frequency density × class width = frequency
Total area = total frequency. Relative frequency density divides by n as well, so total area = 1.
Quartile positions
median at (n + 1) ÷ 2; Q1 at (n + 1) ÷ 4; Q3 at 3(n + 1) ÷ 4
Interpolate between ordered values if the position is not a whole number. R's default quantile() uses a different rule (position 1 + (n − 1)p), so answers can differ slightly. State your method.
Interquartile range
IQR = Q3 − Q1
A measure of spread that is not affected by extreme values.
Outlier fences (common convention)
lower fence = Q1 − 1.5 × IQR; upper fence = Q3 + 1.5 × IQR
Points outside the fences are plotted individually as outliers. This is a convention, not a law. State the rule you use.
Cumulative frequency plotting
plot cumulative frequency against the upper class boundary
Join the points with a smooth curve or straight lines. Start at the lower boundary of the first class with cumulative frequency 0.
Sums of squares and products
sxx = Σx² − (Σx)² ÷ n; syy = Σy² − (Σy)² ÷ n; sxy = Σxy − (Σx)(Σy) ÷ n
Equivalent to Σ(x − x̄)², Σ(y − ȳ)² and Σ(x − x̄)(y − ȳ). Use these for fast calculator work.
Sample covariance
sample covariance = sxy ÷ (n − 1)
Some texts divide by n. State which one you use. The correlation does not change, because the divisor cancels.
Pearson correlation coefficient
r = sxy ÷ √(sxx × syy)
Always between −1 and 1. Unchanged if you add a constant to the data or multiply by a positive constant. Measures linear association only.
Coefficient of determination (simple linear regression)
R² = r²
Proportion of variation in y explained by the fitted straight line on x.
Spearman rank correlation (general)
ρ = Pearson correlation of the ranks = sRxRy ÷ √(sRxRx × sRyRy)
Valid with or without ties. Tied values get the average of the ranks they occupy.
Spearman shortcut (no ties)
ρ = 1 − 6 Σd² ÷ (n(n² − 1)), where d = rank of x − rank of y
Exact only when there are no ties. With a few ties it is only approximate.
Principal component
Z₁ = φ₁₁X₁ + φ₂₁X₂ + … + φₚ₁Xₚ
Variables are usually centred (mean subtracted) first. The score of an observation is this sum using its centred values. Standardise too if you use the correlation matrix.
Loading normalisation
Σⱼ φⱼᵢ² = 1 for each component i
Loading vectors of different components are also orthogonal: Σⱼ φⱼᵢ φⱼₖ = 0 for i ≠ k.
Variance of a component
Var(Zᵢ) = λᵢ
λᵢ is the i-th largest eigenvalue of the covariance (or correlation) matrix, with λ₁ ≥ λ₂ ≥ … ≥ λₚ ≥ 0.
Total variance
Σ λᵢ = trace of the matrix
For a covariance matrix this is the sum of the variances. For a correlation matrix it equals p, the number of variables.
Proportion of variance explained
PVEᵢ = λᵢ ÷ Σ λⱼ
Cumulative PVE for the first k components is (λ₁ + … + λₖ) ÷ Σ λⱼ.
Eigenvalue equation
Σφ = λφ, i.e. (Σ − λI)φ = 0
Solve det(Σ − λI) = 0 for λ, then find φ and scale it to length 1. Σ here is the covariance or correlation matrix.
Correlation of component with variable
Corr(Zᵢ, Xⱼ) = φⱼᵢ √λᵢ ÷ sⱼ
sⱼ is the standard deviation of Xⱼ. For a correlation matrix sⱼ = 1, so it is simply φⱼᵢ √λᵢ.

Quick revision

  • Data can be numerical (discrete or continuous) or categorical (nominal or ordinal); the type decides the summary and graph.
  • Check quality first: missing values, outliers, errors and inconsistent recording.
  • Sample mean = Σx ÷ n; sample variance = Σ(x − x̄)² ÷ (n − 1).
  • The median is less affected by outliers than the mean.
  • Interquartile range = upper quartile − lower quartile; it measures spread and ignores extremes.
  • Right-skewed data usually has mean above median; left-skewed usually the reverse.
  • A histogram shows shape; a boxplot shows median, quartiles and outliers; a Q-Q plot checks fit to a distribution.
  • Sxx = Σx² − (Σx)² ÷ n; Sxy = Σxy − (Σx)(Σy) ÷ n; r = Sxy ÷ √(Sxx × Syy).
  • Correlation r lies between −1 and 1 and measures linear association only; it does not prove causation.
  • Always plot the data: a scatterplot can reveal curves or outliers that r hides.
  • PCA finds uncorrelated linear combinations of the variables that capture maximum variance, in decreasing order.
  • Proportion of variance explained by a component = its eigenvalue ÷ sum of all eigenvalues.

Common mistakes

  • Treating numerically coded categories as numerical data, such as policy type 1, 2, 3. Fix: Ask whether the differences and averages of the codes have meaning. If not, the data is categorical.
  • Calling all numbers with decimals continuous and all whole numbers discrete. Fix: Judge by the underlying quantity. Age in completed years is a rounded measure of a continuous quantity. Number of claims is truly discrete.
  • Dividing by n instead of n − 1 for sample variance. Fix: Default to n − 1 for a sample variance. Use n only when the question says population or gives that definition. Write the divisor in your working.
  • Finding the median or quartiles from unordered data. Fix: Always sort first. Then apply the position formula, such as (n + 1)/4 for Q1, and interpolate if the position is fractional.
  • Plotting frequency instead of frequency density when class widths are unequal. Fix: Check the widths first. If any differ, divide each frequency by its width. Test your graph: bar area should equal frequency.
  • Leaving gaps between histogram bars or treating a histogram like a bar chart. Fix: Continuous data means a continuous scale, so bars touch. Use separate bars only for categories or discrete values.
  • Using (Σx)² as Σx² (or the reverse) in sxx. Fix: Write Σx² (square each value, then add) and (Σx)² (add, then square) as separate lines in your table.
  • Getting r outside the range −1 to 1, or a negative sxx. Fix: Check sxx ≥ 0 and syy ≥ 0 before dividing. Recompute the sums if either fails.
  • Running PCA on the covariance matrix when variables have very different units or scales. Fix: Standardise the variables, or use the correlation matrix, unless all variables share the same units and scale. State your choice.
  • Dividing by the wrong total when finding the proportion of variance explained. Fix: Always divide by the sum of all eigenvalues. That sum is the trace, and it equals p only for a correlation matrix.

Exam tips

  • Give a reason with every classification. One short phrase earns the mark, a bare label may not.
  • For discuss questions, cover both sides: advantages and limits of primary and secondary data, in the context given.
  • State the quartile method if you use the IQR rule, as methods differ slightly. Your limits must follow from the method you state.
  • When treating missing data or outliers, always justify the choice and mention the possible bias.
  • In computer-based papers, check data types and missing values before fitting any model, and note this in your commentary.
  • Write the convention you use for quartiles and variance in one line. Markers then follow your method even if a different convention gives a slightly different number.
  • In MCQs, check the sign first. Mean above median suggests positive skew, and a variance can never be negative. This removes wrong options quickly.
  • For 'comment on' questions, name the measure, give the number, then interpret it. For example: skewness 0.31 shows a mild right tail, so a few large claims lift the mean.