Actuarial Statistics · Purpose and function of data analysis
Types of Data and Data Sources for Actuarial Exams
Updated 11 October 2026 · Fact-checked
Data can be classified by what it measures (categorical or numerical), by how it is collected over time (cross-sectional, time series or longitudinal) and by who collected it (primary or secondary). To answer an exam question, state the type, give a reason tied to the scenario, and name a likely source.
Understand Types of Data and Data Sources
Every data analysis starts with one question: what kind of data do I have? The answer decides which summaries, graphs and models are sensible. Averaging policy type codes makes no sense. Treating a trend over years as independent observations can mislead.
First, classify by what is measured. Categorical (qualitative) data puts items into groups, such as policy type, gender or state. If the groups have a natural order, such as claim severity low, medium or high, the data is ordinal. If not, it is nominal. Numerical (quantitative) data is a number on which arithmetic is meaningful. It is discrete if it takes separate values, such as number of claims. It is continuous if it can take any value in a range, such as claim amount or age at death.
Second, classify by how it is collected over time. Cross-sectional data records many units at a single point in time, such as the sum assured of all policies in force on 31 March. Time series data records one unit or quantity repeatedly over time, such as monthly claim totals for one portfolio. Longitudinal (panel) data follows the same units over time, such as the same policyholders tracked for several years. Longitudinal data is a mix of the two.
Third, classify by source. Primary data is collected by you or your organisation for the specific purpose at hand, for example a survey of customers or an experiment. It fits your needs but costs time and money. Secondary data was collected by someone else, or for another purpose, and you reuse it. It is cheaper and faster, but may not match your definitions, may be out of date, and its quality is harder to check. Examples are insurer policy and claim records used later for pricing, regulator publications, census data and published mortality tables.
Actuaries also meet big data. It is usually described by the Vs: volume (very large size), velocity (arrives fast, often in real time), variety (structured and unstructured forms such as text, images and telematics) and veracity (uncertain quality). It is often secondary data, so check its quality before use.
Key rules to remember
- Classification by measurement
- Data → Categorical (nominal, ordinal) or Numerical (discrete, continuous)
- Ask: is arithmetic meaningful? If not, it is categorical. Ordered categories are ordinal.
- Classification by time structure
- Cross-sectional = many units, one time; Time series = one unit, many times; Longitudinal = many units, many times
- Say which dimension varies. Longitudinal is also called panel data.
- Classification by source
- Primary = collected for this purpose; Secondary = collected by others or for another purpose
- Internal company records reused for a new purpose are secondary in the sense of being collected for another purpose.
- Big data characteristics
- Volume, Velocity, Variety, Veracity
- Some texts add further Vs. Learn these four and be ready to explain each.
How to solve Types of Data and Data Sources questions
Use this method for any question that asks you to classify data, compare data types or discuss sources.
- 1Read the scenario and identify the variable or variables being recorded.
- 2Decide if each variable is categorical or numerical. If numerical, decide discrete or continuous. If categorical, decide nominal or ordinal.
- 3Look at the time structure: one point in time, one series over time, or the same units followed over time.
- 4Decide whether the data is primary or secondary by asking who collected it and for what purpose.
- 5Name the likely source, such as internal records, survey, regulator or published tables.
- 6Give the consequence: which summaries or models suit this type, and which quality issues to check.
- 7Write each point with a short reason tied to the scenario, not just the label.
Quickest way: Three-label shortcut
When to use it: Use in multiple-choice questions and short classification parts where time is tight.
- Label 1, measurement: can you average it meaningfully? Yes means numerical. No means categorical.
- Label 2, time: how many units and how many time points? This gives cross-sectional, time series or longitudinal.
- Label 3, source: did the user collect it for this purpose? Yes means primary. No means secondary.
- Check for traps such as numeric codes for categories (for example 1 = male, 2 = female), which are still categorical.
Common mistakes in Types of Data and Data Sources
Calling numeric codes numerical data
The values look like numbers, such as policy type 1, 2 or 3.
Fix: Ask whether arithmetic means anything. If the average of the codes is meaningless, the data is categorical.
Confusing time series with longitudinal data
Both involve observations over time.
Fix: Time series tracks one quantity or unit. Longitudinal tracks many units over time. Check how many units there are.
Treating all claim counts as continuous
Students link all numbers with measurement.
Fix: Counts are discrete. Amounts, times and ages are continuous, even if recorded rounded.
Saying an insurer's own data is always primary
It belongs to the company, so it feels first-hand.
Fix: Judge by purpose. Records collected for administration and later reused for pricing were not collected for the analysis, so treat them as secondary. State your reasoning.
Listing advantages of secondary data without drawbacks
Cheap and quick is the memorable part.
Fix: Always give both sides: definitions may differ, data may be old, and you cannot control quality.
Giving only labels in written answers
Students think naming the type is enough.
Fix: Add a reason from the scenario and a consequence for analysis. Marks are for application.
Worked examples
Example 1
An insurer holds, for each motor policy in force on 31 March: vehicle type (car, two-wheeler, commercial), number of claims in the last year, and the largest claim amount in rupees. (a) Classify each variable. (b) State the time structure of the data set.
Show the solution
- Vehicle type puts policies into unordered groups, so it is categorical and nominal.
- Number of claims takes values 0, 1, 2, and so on, so it is numerical and discrete.
- Claim amount in rupees can take any value in a range, so it is numerical and continuous.
- All policies are observed at one date, 31 March, with many units and one time point.
Answer: Vehicle type: categorical (nominal). Number of claims: numerical (discrete). Claim amount: numerical (continuous). The data set is cross-sectional.
Example 2
An actuary needs to estimate how often health policyholders are hospitalised. She can (i) run a new survey of policyholders or (ii) use the insurer's existing claim records. Compare the two options in terms of primary and secondary data.
Show the solution
- Option (i) is primary data, collected for this specific question.
- Primary data advantages: questions and definitions match the need, and you control collection. Disadvantages: slow, costly, and survey responses may be biased or incomplete.
- Option (ii) is secondary data, as the claim records were collected for administration and payment.
- Secondary data advantages: quick, cheap and large. Disadvantages: definitions may not match, for example only insured claims are recorded, and fields may be missing or outdated.
- Conclude with a recommendation: start with the claim records, check their quality, and add a targeted survey only for gaps.
Answer: Option (i) is primary data: well suited but slow and costly. Option (ii) is secondary data: fast and cheap but may not fit the purpose, so check quality first. A sensible approach uses the records first and supplements with a survey if needed.
Exam tips
- Always tie each label to the scenario with a short reason. Bare labels earn few marks.
- Learn the four big data Vs with one insurance example each, such as telematics for velocity and variety.
- In source questions, give advantages and disadvantages for both primary and secondary data.
- Watch for coded categories and ordered categories. These are the usual MCQ traps.
- Link data type to analysis choice: categorical suggests counts and bar charts, numerical suggests means and histograms.
Practice questions from Purpose and function of data analysis
- An insurer has 2,000 motor policy records. During the cleaning step, 80 records have a missing vehicle age. The analyst finds that missingne…
- An actuary fits a simple linear regression of claim cost (Rs thousand) on driver age using past data: fitted line = 40 − 0.5 × age. She uses…
- A health insurer tracks 1,000 policyholders every month for 24 months, recording each person's monthly claim cost. Which description fits th…
- A pension scheme dataset has 500 members, and salary is missing for 100 of them. Among the 400 members with salary recorded, the mean salary…
- A motor insurer in India reviews 5,000 claims paid last year and reports the mean claim size, the median, and a histogram of claim amounts. …
Types of Data and Data Sources in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Types of Data and Data Sources: frequently asked questions
What is the difference between cross-sectional and longitudinal data?
Cross-sectional data observes many units at one point in time. Longitudinal data follows the same units over several time points. A single series for one unit, such as monthly total claims, is a time series.
What is the difference between primary and secondary data?
Primary data is collected by you for your specific purpose. Secondary data was collected by someone else or for another purpose and is reused. Secondary data is cheaper but may not fit your needs.
What are the characteristics of big data in CS1?
The usual description is volume, velocity, variety and veracity. These mean large size, fast arrival, mixed forms and uncertain quality. Be ready to explain what each means for an insurer.
Is ordinal data categorical or numerical?
Ordinal data is categorical, because the categories have a natural order but the gaps between them are not meaningful numbers. Examples are low, medium and high risk ratings.