IAI Actuarial Core Principles · Actuarial Statistics
Exploratory Data Analysis for IAI Actuarial Statistics
Exploratory data analysis (EDA) is the first look at a data set before you fit any model. You classify the data, check its quality, compute summary measures, draw graphs, measure correlation and, with many variables, reduce dimension using principal components. In the exam, you must calculate and interpret.
What this chapter covers
Exploratory data analysis is about understanding data before modelling it. You learn what type of data you have, whether it is reliable, how it is centred and spread, what shape it has, and how variables move together. The last topic, principal components analysis, handles data with many correlated variables by finding a few new variables that capture most of the variation.
The chapter is mostly practical. Typical tasks are: compute a mean, variance, quantiles or the sample correlation coefficient from a small data set; read a boxplot, histogram or Q-Q plot; and say what a scatterplot suggests. For PCA you work from a covariance or correlation matrix, and you must be able to interpret the proportion of variance explained and the loadings.
This chapter is the base for the rest of Actuarial Statistics. The summary measures return in random variables and inference. Correlation and scatterplots lead directly into regression. Graphical checks such as Q-Q plots are used to test distributional assumptions. PCA also links to the machine learning ideas that appear in CS2. The Paper B computer-based exam also expects you to produce these summaries and plots in R, so the same ideas are tested by software as well as by hand.
The Core Principles Actuarial Statistics subject is split into topics by syllabus weighting, and data analysis is one of the smaller ones. But the marks here are among the easiest to secure, because the questions are mostly computation and interpretation with clear steps. The skills also feed into regression and inference, which carry much larger weight. If you are weak on summary measures or correlation, you will lose marks later in topics that look unrelated. It is cheap to learn well and costly to neglect.
Exploratory data analysis: topics in the order to study them
- 1Types of Data and Data QualityIt sets the vocabulary and tells you which summaries and graphs suit which data, so everything after builds on it.
- 2Numerical Summary MeasuresMeasures of location, spread, skewness and quantiles are the core calculations and are used in every later topic.
- 3Graphical Data PresentationOnce you can compute summaries, you learn to see the same information in histograms, boxplots and Q-Q plots.
- 4Correlation and ScatterplotsThis moves from one variable to two and prepares you directly for regression.
- 5Principal Components AnalysisIt is the most advanced idea and needs variance, covariance and correlation to be secure first.
How to prepare Exploratory data analysis
Treat this chapter as a mix of short calculations and interpretation. Build speed on the formulas, then practise explaining results in a sentence.
- Write a one-page list of data types (numerical, categorical, discrete, continuous, cross-sectional, time series) and common quality problems such as missing values and outliers.
- Learn the formulas for sample mean, sample variance (divisor n − 1), standard deviation, median, quartiles, and skewness. Practise each on a small data set until you can do it without notes.
- Practise reading histograms, boxplots and Q-Q plots. For each one, state what it shows about shape, spread, outliers and fit to a distribution.
- Compute the sample covariance and correlation coefficient by hand using Sxx, Syy and Sxy. Then say in words what the value means and what it does not mean.
- For PCA, practise working from a given covariance or correlation matrix: find the proportion of total variance explained by each component and decide how many to keep.
- Repeat the same summaries, plots and correlation in R for Paper B, and learn the main commands so you can produce and interpret output quickly.
- Finish with past-paper style questions that mix calculation and comment, and time yourself.
Common mistakes in Exploratory data analysis
Dividing by n instead of n − 1 when computing sample variance.
Fix: Check whether the question gives a sample. If it does, use n − 1 unless told otherwise, and write the formula before substituting.
Claiming that a high correlation shows one variable causes the other.
Fix: Say that correlation shows linear association only. Mention possible common causes or chance when asked to comment.
Using the mean as the best summary for skewed data or data with outliers.
Fix: Look at the shape first. For skewed data or outliers, prefer the median and interquartile range and say why.
Reading a Q-Q plot as a scatterplot of the data against time.
Fix: Remember that it compares sample quantiles with theoretical quantiles. Points near a straight line suggest a good fit; systematic curves show skewness or heavy tails.
Doing PCA on variables with very different units without standardising.
Fix: When variables have different units or scales, use the correlation matrix, which standardises them, and state this choice.
Giving a number without interpretation in written questions.
Fix: Add one sentence saying what the result means for the data, and tie it to the context in the question.
Last-day revision: Exploratory data analysis
- Data can be numerical (discrete or continuous) or categorical (nominal or ordinal); the type decides the summary and graph.
- Check quality first: missing values, outliers, errors and inconsistent recording.
- Sample mean = Σx ÷ n; sample variance = Σ(x − x̄)² ÷ (n − 1).
- The median is less affected by outliers than the mean.
- Interquartile range = upper quartile − lower quartile; it measures spread and ignores extremes.
- Right-skewed data usually has mean above median; left-skewed usually the reverse.
- A histogram shows shape; a boxplot shows median, quartiles and outliers; a Q-Q plot checks fit to a distribution.
- Sxx = Σx² − (Σx)² ÷ n; Sxy = Σxy − (Σx)(Σy) ÷ n; r = Sxy ÷ √(Sxx × Syy).
- Correlation r lies between −1 and 1 and measures linear association only; it does not prove causation.
- Always plot the data: a scatterplot can reveal curves or outliers that r hides.
- PCA finds uncorrelated linear combinations of the variables that capture maximum variance, in decreasing order.
- Proportion of variance explained by a component = its eigenvalue ÷ sum of all eigenvalues.
Exploratory data analysis practice questions
- For a sample, the variances are sx² = 16 and sy² = 25, and the sample covariance is 15. Each x is then transformed to u = 3 - 2x, while y is…
- Which feature is a genuine advantage of a stem-and-leaf plot over a histogram for a small data set of 30 observations?
- Daily premium collections (₹ lakh) over 5 days are 2, 4, 6, 8 and 30. Which statement about the outlier rule 'flag values beyond 1.5 × IQR f…
- A scatterplot of claim size (y) against sum insured (x) shows a clear increasing relationship, but the vertical spread of points grows as x …
- For 20 policies, the sum of claim amounts is Rs 400 thousand and the sum of squares of claims is 9,800 (Rs thousand squared). What is the sa…
- A histogram of monthly claim amounts (in ₹ thousand) uses unequal class widths. The class 10–20 has frequency 30 and the class 20–40 has fre…
- Three variables are standardised before PCA, so PCA is applied to their correlation matrix. The first two eigenvalues are 1.8 and 0.9. Which…
- A scatterplot of monthly premium income (x, in ₹ lakh) against claims (y, in ₹ lakh) for a Mumbai insurer gives r = 0.80. Every x value is t…
Exploratory data analysis in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Exploratory data analysis: frequently asked questions
How much time should I give to exploratory data analysis?
Give it less time than regression or inference, but do not skip it. The ideas are quick to learn and are used throughout the paper. A few focused sessions of calculation practice are usually enough before you move on.
Do I need to learn PCA calculations by hand?
You should be able to work from a given covariance or correlation matrix, find variance proportions and interpret the loadings. Finding eigenvalues of a small matrix by hand is good practice. For larger data, expect to use R output.
Is this chapter tested in the computer-based Paper B?
Yes, the ideas can be tested there through R, such as computing summaries, drawing plots and finding correlation or components. Practise the commands as well as the theory so that you can interpret the output.
What is the difference between covariance and correlation?
Covariance measures how two variables move together but depends on their units. Correlation divides covariance by the product of the standard deviations, so it is unit-free and lies between −1 and 1.