Skip to content

Actuarial Statistics · Purpose and function of data analysis

Types of Data and Data Sources for Actuarial Exams

Updated 11 October 2026 · Fact-checked

Data can be classified by what it measures (categorical or numerical), by how it is collected over time (cross-sectional, time series or longitudinal) and by who collected it (primary or secondary). To answer an exam question, state the type, give a reason tied to the scenario, and name a likely source.

Understand Types of Data and Data Sources

Every data analysis starts with one question: what kind of data do I have? The answer decides which summaries, graphs and models are sensible. Averaging policy type codes makes no sense. Treating a trend over years as independent observations can mislead.

First, classify by what is measured. Categorical (qualitative) data puts items into groups, such as policy type, gender or state. If the groups have a natural order, such as claim severity low, medium or high, the data is ordinal. If not, it is nominal. Numerical (quantitative) data is a number on which arithmetic is meaningful. It is discrete if it takes separate values, such as number of claims. It is continuous if it can take any value in a range, such as claim amount or age at death.

Second, classify by how it is collected over time. Cross-sectional data records many units at a single point in time, such as the sum assured of all policies in force on 31 March. Time series data records one unit or quantity repeatedly over time, such as monthly claim totals for one portfolio. Longitudinal (panel) data follows the same units over time, such as the same policyholders tracked for several years. Longitudinal data is a mix of the two.

Third, classify by source. Primary data is collected by you or your organisation for the specific purpose at hand, for example a survey of customers or an experiment. It fits your needs but costs time and money. Secondary data was collected by someone else, or for another purpose, and you reuse it. It is cheaper and faster, but may not match your definitions, may be out of date, and its quality is harder to check. Examples are insurer policy and claim records used later for pricing, regulator publications, census data and published mortality tables.

Actuaries also meet big data. It is usually described by the Vs: volume (very large size), velocity (arrives fast, often in real time), variety (structured and unstructured forms such as text, images and telematics) and veracity (uncertain quality). It is often secondary data, so check its quality before use.

Key rules to remember

Classification by measurement
Data → Categorical (nominal, ordinal) or Numerical (discrete, continuous)
Ask: is arithmetic meaningful? If not, it is categorical. Ordered categories are ordinal.
Classification by time structure
Cross-sectional = many units, one time; Time series = one unit, many times; Longitudinal = many units, many times
Say which dimension varies. Longitudinal is also called panel data.
Classification by source
Primary = collected for this purpose; Secondary = collected by others or for another purpose
Internal company records reused for a new purpose are secondary in the sense of being collected for another purpose.
Big data characteristics
Volume, Velocity, Variety, Veracity
Some texts add further Vs. Learn these four and be ready to explain each.

How to solve Types of Data and Data Sources questions

Use this method for any question that asks you to classify data, compare data types or discuss sources.

  1. 1Read the scenario and identify the variable or variables being recorded.
  2. 2Decide if each variable is categorical or numerical. If numerical, decide discrete or continuous. If categorical, decide nominal or ordinal.
  3. 3Look at the time structure: one point in time, one series over time, or the same units followed over time.
  4. 4Decide whether the data is primary or secondary by asking who collected it and for what purpose.
  5. 5Name the likely source, such as internal records, survey, regulator or published tables.
  6. 6Give the consequence: which summaries or models suit this type, and which quality issues to check.
  7. 7Write each point with a short reason tied to the scenario, not just the label.

Quickest way: Three-label shortcut

When to use it: Use in multiple-choice questions and short classification parts where time is tight.

  1. Label 1, measurement: can you average it meaningfully? Yes means numerical. No means categorical.
  2. Label 2, time: how many units and how many time points? This gives cross-sectional, time series or longitudinal.
  3. Label 3, source: did the user collect it for this purpose? Yes means primary. No means secondary.
  4. Check for traps such as numeric codes for categories (for example 1 = male, 2 = female), which are still categorical.

Common mistakes in Types of Data and Data Sources

  • Calling numeric codes numerical data

    The values look like numbers, such as policy type 1, 2 or 3.

    Fix: Ask whether arithmetic means anything. If the average of the codes is meaningless, the data is categorical.

  • Confusing time series with longitudinal data

    Both involve observations over time.

    Fix: Time series tracks one quantity or unit. Longitudinal tracks many units over time. Check how many units there are.

  • Treating all claim counts as continuous

    Students link all numbers with measurement.

    Fix: Counts are discrete. Amounts, times and ages are continuous, even if recorded rounded.

  • Saying an insurer's own data is always primary

    It belongs to the company, so it feels first-hand.

    Fix: Judge by purpose. Records collected for administration and later reused for pricing were not collected for the analysis, so treat them as secondary. State your reasoning.

  • Listing advantages of secondary data without drawbacks

    Cheap and quick is the memorable part.

    Fix: Always give both sides: definitions may differ, data may be old, and you cannot control quality.

  • Giving only labels in written answers

    Students think naming the type is enough.

    Fix: Add a reason from the scenario and a consequence for analysis. Marks are for application.

Worked examples

Example 1

An insurer holds, for each motor policy in force on 31 March: vehicle type (car, two-wheeler, commercial), number of claims in the last year, and the largest claim amount in rupees. (a) Classify each variable. (b) State the time structure of the data set.

Show the solution
  1. Vehicle type puts policies into unordered groups, so it is categorical and nominal.
  2. Number of claims takes values 0, 1, 2, and so on, so it is numerical and discrete.
  3. Claim amount in rupees can take any value in a range, so it is numerical and continuous.
  4. All policies are observed at one date, 31 March, with many units and one time point.

Answer: Vehicle type: categorical (nominal). Number of claims: numerical (discrete). Claim amount: numerical (continuous). The data set is cross-sectional.

Example 2

An actuary needs to estimate how often health policyholders are hospitalised. She can (i) run a new survey of policyholders or (ii) use the insurer's existing claim records. Compare the two options in terms of primary and secondary data.

Show the solution
  1. Option (i) is primary data, collected for this specific question.
  2. Primary data advantages: questions and definitions match the need, and you control collection. Disadvantages: slow, costly, and survey responses may be biased or incomplete.
  3. Option (ii) is secondary data, as the claim records were collected for administration and payment.
  4. Secondary data advantages: quick, cheap and large. Disadvantages: definitions may not match, for example only insured claims are recorded, and fields may be missing or outdated.
  5. Conclude with a recommendation: start with the claim records, check their quality, and add a targeted survey only for gaps.

Answer: Option (i) is primary data: well suited but slow and costly. Option (ii) is secondary data: fast and cheap but may not fit the purpose, so check quality first. A sensible approach uses the records first and supplements with a survey if needed.

Exam tips

  • Always tie each label to the scenario with a short reason. Bare labels earn few marks.
  • Learn the four big data Vs with one insurance example each, such as telematics for velocity and variety.
  • In source questions, give advantages and disadvantages for both primary and secondary data.
  • Watch for coded categories and ordered categories. These are the usual MCQ traps.
  • Link data type to analysis choice: categorical suggests counts and bar charts, numerical suggests means and histograms.

Practice questions from Purpose and function of data analysis

Types of Data and Data Sources in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Types of Data and Data Sources: frequently asked questions

What is the difference between cross-sectional and longitudinal data?

Cross-sectional data observes many units at one point in time. Longitudinal data follows the same units over several time points. A single series for one unit, such as monthly total claims, is a time series.

What is the difference between primary and secondary data?

Primary data is collected by you for your specific purpose. Secondary data was collected by someone else or for another purpose and is reused. Secondary data is cheaper but may not fit your needs.

What are the characteristics of big data in CS1?

The usual description is volume, velocity, variety and veracity. These mean large size, fast arrival, mixed forms and uncertain quality. Be ready to explain what each means for an insurer.

Is ordinal data categorical or numerical?

Ordinal data is categorical, because the categories have a natural order but the gaps between them are not meaningful numbers. Examples are low, medium and high risk ratings.