IAI Actuarial Core Principles · Actuarial Statistics
Exploratory data analysis: formula sheet
Key formulas
- Classification of data
- Data → Categorical (nominal, ordinal) or Numerical (discrete, continuous)
- Ask: is it a label, a count, or a measurement? Ordinal data has order but gaps between categories are not meaningful.
- Outlier rule of thumb (IQR rule)
- Flag x if x < Q1 − 1.5 × IQR or x > Q3 + 1.5 × IQR, where IQR = Q3 − Q1
- A convention for flagging only, not proof of an error. Always investigate before removing.
- Missing data types
- MCAR, MAR, MNAR
- Missing completely at random, missing at random (depends on observed data), missing not at random (depends on the missing value itself). MNAR is the most likely to bias results.
- Proportion missing
- Proportion missing = (number of missing values) ÷ (number of records)
- Compute it per variable. A high proportion limits what any treatment can fix.
- Sample mean
- x̄ = Σx ÷ n
- For grouped data use x̄ = Σfx ÷ Σf, with x as the class midpoint. The result is an estimate.
- Sample variance
- s² = Σ(x − x̄)² ÷ (n − 1) = (Σx² − n x̄²) ÷ (n − 1)
- Divide by n instead of n − 1 only if the question says population variance or divisor n. Standard deviation s = √s².
- Sum of squares
- Sxx = Σx² − (Σx)² ÷ n
- Then s² = Sxx ÷ (n − 1). The shortcut avoids calculating each deviation.
- Median and quartile positions (ungrouped)
- Median at (n + 1)/2; Q1 at (n + 1)/4; Q3 at 3(n + 1)/4; pth percentile at p(n + 1)/100
- If the position is not a whole number, interpolate between the two neighbouring ordered values. Say which convention you use.
- Interquartile range
- IQR = Q3 − Q1
- Measures the spread of the middle 50% of the data.
- Grouped percentile by interpolation
- Value = L + ((target − F) ÷ f) × w
- L = lower class boundary, F = cumulative frequency before the class, f = class frequency, w = class width. Target is n/4, n/2 or 3n/4 for quartiles and median.
- Coefficient of variation
- CV = s ÷ x̄
- Compares spread relative to the mean. Use only when the mean is positive and meaningful.
- Coefficient of skewness (moment form)
- Skewness = [Σ(x − x̄)³ ÷ n] ÷ σ³, where σ² = Σ(x − x̄)² ÷ n
- For a distribution: E[(X − μ)³] ÷ σ³. Positive means right-skewed. Divisors vary between texts, so state yours.
- Pearson's skewness
- 3 × (mean − median) ÷ standard deviation
- A quick approximate measure. It is not the moment-based coefficient.
- Kurtosis
- Kurtosis = [Σ(x − x̄)⁴ ÷ n] ÷ σ⁴; excess kurtosis = kurtosis − 3
- The normal distribution has kurtosis 3 and excess kurtosis 0.
- Frequency density
- frequency density = class frequency ÷ class width
- Use this as bar height in a histogram with unequal class widths. Area of bar = class frequency.
- Histogram area rule
- bar area = frequency density × class width = frequency
- Total area = total frequency. Relative frequency density divides by n as well, so total area = 1.
- Quartile positions
- median at (n + 1) ÷ 2; Q1 at (n + 1) ÷ 4; Q3 at 3(n + 1) ÷ 4
- Interpolate between ordered values if the position is not a whole number. R's default quantile() uses a different rule (position 1 + (n − 1)p), so answers can differ slightly. State your method.
- Interquartile range
- IQR = Q3 − Q1
- A measure of spread that is not affected by extreme values.
- Outlier fences (common convention)
- lower fence = Q1 − 1.5 × IQR; upper fence = Q3 + 1.5 × IQR
- Points outside the fences are plotted individually as outliers. This is a convention, not a law. State the rule you use.
- Cumulative frequency plotting
- plot cumulative frequency against the upper class boundary
- Join the points with a smooth curve or straight lines. Start at the lower boundary of the first class with cumulative frequency 0.
- Sums of squares and products
- sxx = Σx² − (Σx)² ÷ n; syy = Σy² − (Σy)² ÷ n; sxy = Σxy − (Σx)(Σy) ÷ n
- Equivalent to Σ(x − x̄)², Σ(y − ȳ)² and Σ(x − x̄)(y − ȳ). Use these for fast calculator work.
- Sample covariance
- sample covariance = sxy ÷ (n − 1)
- Some texts divide by n. State which one you use. The correlation does not change, because the divisor cancels.
- Pearson correlation coefficient
- r = sxy ÷ √(sxx × syy)
- Always between −1 and 1. Unchanged if you add a constant to the data or multiply by a positive constant. Measures linear association only.
- Coefficient of determination (simple linear regression)
- R² = r²
- Proportion of variation in y explained by the fitted straight line on x.
- Spearman rank correlation (general)
- ρ = Pearson correlation of the ranks = sRxRy ÷ √(sRxRx × sRyRy)
- Valid with or without ties. Tied values get the average of the ranks they occupy.
- Spearman shortcut (no ties)
- ρ = 1 − 6 Σd² ÷ (n(n² − 1)), where d = rank of x − rank of y
- Exact only when there are no ties. With a few ties it is only approximate.
- Principal component
- Z₁ = φ₁₁X₁ + φ₂₁X₂ + … + φₚ₁Xₚ
- Variables are usually centred (mean subtracted) first. The score of an observation is this sum using its centred values. Standardise too if you use the correlation matrix.
- Loading normalisation
- Σⱼ φⱼᵢ² = 1 for each component i
- Loading vectors of different components are also orthogonal: Σⱼ φⱼᵢ φⱼₖ = 0 for i ≠ k.
- Variance of a component
- Var(Zᵢ) = λᵢ
- λᵢ is the i-th largest eigenvalue of the covariance (or correlation) matrix, with λ₁ ≥ λ₂ ≥ … ≥ λₚ ≥ 0.
- Total variance
- Σ λᵢ = trace of the matrix
- For a covariance matrix this is the sum of the variances. For a correlation matrix it equals p, the number of variables.
- Proportion of variance explained
- PVEᵢ = λᵢ ÷ Σ λⱼ
- Cumulative PVE for the first k components is (λ₁ + … + λₖ) ÷ Σ λⱼ.
- Eigenvalue equation
- Σφ = λφ, i.e. (Σ − λI)φ = 0
- Solve det(Σ − λI) = 0 for λ, then find φ and scale it to length 1. Σ here is the covariance or correlation matrix.
- Correlation of component with variable
- Corr(Zᵢ, Xⱼ) = φⱼᵢ √λᵢ ÷ sⱼ
- sⱼ is the standard deviation of Xⱼ. For a correlation matrix sⱼ = 1, so it is simply φⱼᵢ √λᵢ.
Quick revision
- Data can be numerical (discrete or continuous) or categorical (nominal or ordinal); the type decides the summary and graph.
- Check quality first: missing values, outliers, errors and inconsistent recording.
- Sample mean = Σx ÷ n; sample variance = Σ(x − x̄)² ÷ (n − 1).
- The median is less affected by outliers than the mean.
- Interquartile range = upper quartile − lower quartile; it measures spread and ignores extremes.
- Right-skewed data usually has mean above median; left-skewed usually the reverse.
- A histogram shows shape; a boxplot shows median, quartiles and outliers; a Q-Q plot checks fit to a distribution.
- Sxx = Σx² − (Σx)² ÷ n; Sxy = Σxy − (Σx)(Σy) ÷ n; r = Sxy ÷ √(Sxx × Syy).
- Correlation r lies between −1 and 1 and measures linear association only; it does not prove causation.
- Always plot the data: a scatterplot can reveal curves or outliers that r hides.
- PCA finds uncorrelated linear combinations of the variables that capture maximum variance, in decreasing order.
- Proportion of variance explained by a component = its eigenvalue ÷ sum of all eigenvalues.
Common mistakes
- Treating numerically coded categories as numerical data, such as policy type 1, 2, 3. Fix: Ask whether the differences and averages of the codes have meaning. If not, the data is categorical.
- Calling all numbers with decimals continuous and all whole numbers discrete. Fix: Judge by the underlying quantity. Age in completed years is a rounded measure of a continuous quantity. Number of claims is truly discrete.
- Dividing by n instead of n − 1 for sample variance. Fix: Default to n − 1 for a sample variance. Use n only when the question says population or gives that definition. Write the divisor in your working.
- Finding the median or quartiles from unordered data. Fix: Always sort first. Then apply the position formula, such as (n + 1)/4 for Q1, and interpolate if the position is fractional.
- Plotting frequency instead of frequency density when class widths are unequal. Fix: Check the widths first. If any differ, divide each frequency by its width. Test your graph: bar area should equal frequency.
- Leaving gaps between histogram bars or treating a histogram like a bar chart. Fix: Continuous data means a continuous scale, so bars touch. Use separate bars only for categories or discrete values.
- Using (Σx)² as Σx² (or the reverse) in sxx. Fix: Write Σx² (square each value, then add) and (Σx)² (add, then square) as separate lines in your table.
- Getting r outside the range −1 to 1, or a negative sxx. Fix: Check sxx ≥ 0 and syy ≥ 0 before dividing. Recompute the sums if either fails.
- Running PCA on the covariance matrix when variables have very different units or scales. Fix: Standardise the variables, or use the correlation matrix, unless all variables share the same units and scale. State your choice.
- Dividing by the wrong total when finding the proportion of variance explained. Fix: Always divide by the sum of all eigenvalues. That sum is the trace, and it equals p only for a correlation matrix.
Exam tips
- Give a reason with every classification. One short phrase earns the mark, a bare label may not.
- For discuss questions, cover both sides: advantages and limits of primary and secondary data, in the context given.
- State the quartile method if you use the IQR rule, as methods differ slightly. Your limits must follow from the method you state.
- When treating missing data or outliers, always justify the choice and mention the possible bias.
- In computer-based papers, check data types and missing values before fitting any model, and note this in your commentary.
- Write the convention you use for quartiles and variance in one line. Markers then follow your method even if a different convention gives a slightly different number.
- In MCQs, check the sign first. Mean above median suggests positive skew, and a variance can never be negative. This removes wrong options quickly.
- For 'comment on' questions, name the measure, give the number, then interpret it. For example: skewness 0.31 shows a mild right tail, so a few large claims lift the mean.