Skip to content

IAI Actuarial Core Principles · Actuarial Statistics

Exploratory Data Analysis for IAI Actuarial Statistics

Exploratory data analysis (EDA) is the first look at a data set before you fit any model. You classify the data, check its quality, compute summary measures, draw graphs, measure correlation and, with many variables, reduce dimension using principal components. In the exam, you must calculate and interpret.

What this chapter covers

Exploratory data analysis is about understanding data before modelling it. You learn what type of data you have, whether it is reliable, how it is centred and spread, what shape it has, and how variables move together. The last topic, principal components analysis, handles data with many correlated variables by finding a few new variables that capture most of the variation.

The chapter is mostly practical. Typical tasks are: compute a mean, variance, quantiles or the sample correlation coefficient from a small data set; read a boxplot, histogram or Q-Q plot; and say what a scatterplot suggests. For PCA you work from a covariance or correlation matrix, and you must be able to interpret the proportion of variance explained and the loadings.

This chapter is the base for the rest of Actuarial Statistics. The summary measures return in random variables and inference. Correlation and scatterplots lead directly into regression. Graphical checks such as Q-Q plots are used to test distributional assumptions. PCA also links to the machine learning ideas that appear in CS2. The Paper B computer-based exam also expects you to produce these summaries and plots in R, so the same ideas are tested by software as well as by hand.

The Core Principles Actuarial Statistics subject is split into topics by syllabus weighting, and data analysis is one of the smaller ones. But the marks here are among the easiest to secure, because the questions are mostly computation and interpretation with clear steps. The skills also feed into regression and inference, which carry much larger weight. If you are weak on summary measures or correlation, you will lose marks later in topics that look unrelated. It is cheap to learn well and costly to neglect.

Exploratory data analysis: topics in the order to study them

  1. 1Types of Data and Data QualityIt sets the vocabulary and tells you which summaries and graphs suit which data, so everything after builds on it.
  2. 2Numerical Summary MeasuresMeasures of location, spread, skewness and quantiles are the core calculations and are used in every later topic.
  3. 3Graphical Data PresentationOnce you can compute summaries, you learn to see the same information in histograms, boxplots and Q-Q plots.
  4. 4Correlation and ScatterplotsThis moves from one variable to two and prepares you directly for regression.
  5. 5Principal Components AnalysisIt is the most advanced idea and needs variance, covariance and correlation to be secure first.

How to prepare Exploratory data analysis

Treat this chapter as a mix of short calculations and interpretation. Build speed on the formulas, then practise explaining results in a sentence.

  1. Write a one-page list of data types (numerical, categorical, discrete, continuous, cross-sectional, time series) and common quality problems such as missing values and outliers.
  2. Learn the formulas for sample mean, sample variance (divisor n − 1), standard deviation, median, quartiles, and skewness. Practise each on a small data set until you can do it without notes.
  3. Practise reading histograms, boxplots and Q-Q plots. For each one, state what it shows about shape, spread, outliers and fit to a distribution.
  4. Compute the sample covariance and correlation coefficient by hand using Sxx, Syy and Sxy. Then say in words what the value means and what it does not mean.
  5. For PCA, practise working from a given covariance or correlation matrix: find the proportion of total variance explained by each component and decide how many to keep.
  6. Repeat the same summaries, plots and correlation in R for Paper B, and learn the main commands so you can produce and interpret output quickly.
  7. Finish with past-paper style questions that mix calculation and comment, and time yourself.

Common mistakes in Exploratory data analysis

  • Dividing by n instead of n − 1 when computing sample variance.

    Fix: Check whether the question gives a sample. If it does, use n − 1 unless told otherwise, and write the formula before substituting.

  • Claiming that a high correlation shows one variable causes the other.

    Fix: Say that correlation shows linear association only. Mention possible common causes or chance when asked to comment.

  • Using the mean as the best summary for skewed data or data with outliers.

    Fix: Look at the shape first. For skewed data or outliers, prefer the median and interquartile range and say why.

  • Reading a Q-Q plot as a scatterplot of the data against time.

    Fix: Remember that it compares sample quantiles with theoretical quantiles. Points near a straight line suggest a good fit; systematic curves show skewness or heavy tails.

  • Doing PCA on variables with very different units without standardising.

    Fix: When variables have different units or scales, use the correlation matrix, which standardises them, and state this choice.

  • Giving a number without interpretation in written questions.

    Fix: Add one sentence saying what the result means for the data, and tie it to the context in the question.

Last-day revision: Exploratory data analysis

  • Data can be numerical (discrete or continuous) or categorical (nominal or ordinal); the type decides the summary and graph.
  • Check quality first: missing values, outliers, errors and inconsistent recording.
  • Sample mean = Σx ÷ n; sample variance = Σ(x − x̄)² ÷ (n − 1).
  • The median is less affected by outliers than the mean.
  • Interquartile range = upper quartile − lower quartile; it measures spread and ignores extremes.
  • Right-skewed data usually has mean above median; left-skewed usually the reverse.
  • A histogram shows shape; a boxplot shows median, quartiles and outliers; a Q-Q plot checks fit to a distribution.
  • Sxx = Σx² − (Σx)² ÷ n; Sxy = Σxy − (Σx)(Σy) ÷ n; r = Sxy ÷ √(Sxx × Syy).
  • Correlation r lies between −1 and 1 and measures linear association only; it does not prove causation.
  • Always plot the data: a scatterplot can reveal curves or outliers that r hides.
  • PCA finds uncorrelated linear combinations of the variables that capture maximum variance, in decreasing order.
  • Proportion of variance explained by a component = its eigenvalue ÷ sum of all eigenvalues.

Exploratory data analysis practice questions

Exploratory data analysis in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Exploratory data analysis: frequently asked questions

How much time should I give to exploratory data analysis?

Give it less time than regression or inference, but do not skip it. The ideas are quick to learn and are used throughout the paper. A few focused sessions of calculation practice are usually enough before you move on.

Do I need to learn PCA calculations by hand?

You should be able to work from a given covariance or correlation matrix, find variance proportions and interpret the loadings. Finding eigenvalues of a small matrix by hand is good practice. For larger data, expect to use R output.

Is this chapter tested in the computer-based Paper B?

Yes, the ideas can be tested there through R, such as computing summaries, drawing plots and finding correlation or components. Practise the commands as well as the theory so that you can interpret the output.

What is the difference between covariance and correlation?

Covariance measures how two variables move together but depends on their units. Correlation divides covariance by the product of the standard deviations, so it is unit-free and lies between −1 and 1.