Skip to content

Actuarial Statistics · Purpose and function of data analysis

Data Quality, Cleaning and Ethical Considerations in Actuarial Analysis

Updated 11 October 2026 · Fact-checked

Data quality means the data is accurate, complete, consistent and fit for the purpose. Cleaning is the process of finding and fixing errors, gaps and outliers, and recording what you did. Ethics means protecting personal data, being honest about limits, and following professional standards. In exams, state the issue, its effect and your action.

Understand Data Quality, Cleaning and Ethical Considerations

Every actuarial result depends on the data behind it. A perfect model fitted to poor data gives a poor answer. So the first job in any analysis is to ask whether the data is good enough for the question.

There are a few standard dimensions of quality. Accuracy: do the values reflect reality (is the age 35, or was it mistyped as 53)? Completeness: are records or fields missing? Consistency: do values agree across fields and sources (a date of death before the date of birth is an error; one source using years and another months is inconsistent)? Timeliness: is the data recent enough? Relevance: does it fit the purpose? Duplicates and different definitions across sources are further common problems.

Data cleaning (or validation) is how you deal with these. You run checks: range checks, cross-field checks, duplicate checks, and reconciliation to control totals such as total premiums or claim counts in the accounts. For missing data, you can remove the records, fill in values (for example the mean, or a value predicted from other fields), or treat missing as its own category. Each has a cost. Removing records can bias the result if the data is not missing at random. Filling with the mean reduces variability. For outliers, first check whether it is an error. If it is, correct it. If it is genuine, such as a very large claim, do not just delete it, because it may be the most important part of the risk. You can analyse with and without it, or model it separately.

Good practice is to keep the raw data, document every change, and make the process repeatable. State the limitations of the data and their likely effect on the results, and tell the user.

Ethics and professional duties sit alongside this. Personal data, especially health and financial data, must be handled lawfully, kept secure, and used only for the purpose it was collected for. Anonymise or aggregate where possible. In India, data protection law applies to personal data, so you should comply with the current legal requirements without quoting sections you are unsure of. Actuaries must also follow the professional standards and code of conduct of the IAI: act with integrity, be competent, communicate clearly, avoid bias and conflicts of interest, and not mislead users. Using data in ways that unfairly discriminate, or hiding data weaknesses, is a breach of these duties.

Key rules to remember

Missing data: mean imputation
x̂ᵢ = x̄ = (Σ xⱼ) ÷ n, summed over the n observed values
Keeps the mean unchanged but understates variance. Use only with a few missing values.
Outlier rule of thumb (IQR)
Flag x if x < Q₁ − 1.5 × IQR or x > Q₃ + 1.5 × IQR, where IQR = Q₃ − Q₁
This is a screening rule, not proof of an error. Investigate before removing.
Outlier screen (standard deviations)
z = (x − x̄) ÷ s; investigate if |z| is large, for example above 3
The mean and s are themselves affected by the outlier. The cut-off 3 is a convention.
Reconciliation check
Difference = total from dataset − control total from accounts
A non-zero difference signals missing, duplicated or wrong records.

How to solve Data Quality, Cleaning and Ethical Considerations questions

Use this order for any question on data quality, cleaning or ethics. It keeps your answer structured and earns marks for each part.

  1. 1Read the scenario and identify the purpose of the analysis and the data source.
  2. 2List the problems that apply: accuracy, completeness, consistency, duplicates, timeliness, relevance, definitions.
  3. 3For each problem, say how you would detect it (range check, cross-field check, reconciliation, summary statistics, plots).
  4. 4State the action for each: correct, remove, impute or model separately, and give the effect on results and bias.
  5. 5Document the changes, keep the raw data, and state the limitations to the user.
  6. 6Add the ethical and legal points: data protection, purpose limitation, security, anonymisation, professional standards.
  7. 7If numbers are given, calculate them (means, IQR, control total differences) and show the working.
  8. 8Finish with a short conclusion on whether the data is fit for the purpose.

Quickest way: Issue, check, action, effect

When to use it: Use for short written or MCQ questions where you must name and handle data problems quickly.

  1. For each problem write four short words: issue, check, action, effect.
  2. Never say only 'delete the outlier'. Say 'investigate, then correct or keep if genuine'.
  3. For missing data, always mention the bias risk and whether data is missing at random.
  4. Add one line each on documentation and data protection.

Common mistakes in Data Quality, Cleaning and Ethical Considerations

  • Deleting every outlier automatically.

    Outliers look like errors and removing them gives neat results.

    Fix: Investigate first. Genuine large claims are real risk and should be kept or modelled separately.

  • Saying mean imputation is harmless.

    The mean stays the same, so it seems safe.

    Fix: State that it understates variance and can distort relationships between variables.

  • Treating completeness and accuracy as the same thing.

    Both seem to mean 'good data'.

    Fix: Completeness is about missing values or records. Accuracy is whether recorded values are correct.

  • Ignoring ethics and data protection in a technical answer.

    Students think the topic is only about statistics.

    Fix: Add at least one point on lawful use, security, anonymisation or professional duty when personal data appears.

  • Not documenting the cleaning steps.

    Cleaning feels like preparation rather than part of the analysis.

    Fix: Say you will keep the raw data, log each change and tell the user the limitations.

  • Using the IQR rule as proof that a value is wrong.

    The rule gives a clear numerical answer.

    Fix: Call it a flag for investigation, not a verdict.

Worked examples

Example 1

Claim amounts (₹ thousands) for nine motor policies are: 12, 15, 14, 18, 16, 13, 17, 15, 95. Using the 1.5 × IQR rule with quartiles as the medians of the lower and upper halves (excluding the overall median), identify any outlier and comment on what to do.

Show the solution
  1. Order the data: 12, 13, 14, 15, 15, 16, 17, 18, 95.
  2. The median is the 5th value = 15. Lower half: 12, 13, 14, 15. Upper half: 16, 17, 18, 95.
  3. Q₁ = (13 + 14) ÷ 2 = 13.5. Q₃ = (17 + 18) ÷ 2 = 17.5.
  4. IQR = 17.5 − 13.5 = 4.
  5. Upper fence = 17.5 + 1.5 × 4 = 23.5. Lower fence = 13.5 − 6 = 7.5.
  6. 95 > 23.5, so it is flagged. No other value is outside the fences.
  7. Investigate 95: check for a keying error (such as 9.5 or 9,500) or a genuine large claim. Correct if an error. Keep it, or model it separately, if genuine.

Answer: The value 95 (₹95,000) is flagged as an outlier. Investigate before deciding: correct it if it is an error, keep or model it separately if it is genuine.

Example 2

An insurer's policy dataset has 1,000 records. Ages are missing for 40 records. The mean of the 960 known ages is 42.0 years. (a) Find the mean age if the 40 missing ages are filled with 42.0. (b) Give two risks of this approach and one ethical point on using the data.

Show the solution
  1. (a) Sum of known ages = 960 × 42.0 = 40,320.
  2. Imputed sum = 40 × 42.0 = 1,680.
  3. Total = 40,320 + 1,680 = 42,000. Mean = 42,000 ÷ 1,000 = 42.0.
  4. (b) Risk 1: the variance of age is understated, because 40 values are set equal to the mean.
  5. Risk 2: if ages are not missing at random (for example, older customers leave age blank), the imputed values are biased and relationships with claims may be distorted.
  6. Ethical point: the data is personal, so use it only for the purpose it was collected for, keep it secure, and anonymise where possible. Tell the user about the imputation as a limitation.

Answer: (a) The mean stays 42.0 years. (b) Variance is understated and bias is possible if data is not missing at random. Use personal data lawfully and only for its stated purpose, and disclose the imputation.

Exam tips

  • Structure written answers by problem: issue, check, action, effect. Examiners reward each clearly named point.
  • Use words like 'investigate' and 'justify' for outliers and missing data. Blanket rules lose marks.
  • If the scenario has personal data, always add a data protection or professional conduct point.
  • For numerical questions, show quartiles, fences or control-total differences step by step.
  • In MCQs, watch for options that overstate: 'always remove', 'always harmless'.

Practice questions from Purpose and function of data analysis

Data Quality, Cleaning and Ethical Considerations in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Data Quality, Cleaning and Ethical Considerations: frequently asked questions

What is the difference between data cleaning and data validation?

Validation is checking data against rules to find errors, such as range and consistency checks. Cleaning is acting on what you find, by correcting, removing or filling values. In practice the two go together and should be documented.

How should I deal with missing data in an exam answer?

First say why the data may be missing and whether it is missing at random. Then give options: remove records, impute values, or treat missing as a category. State the bias or variance effect of your choice.

Should outliers always be removed?

No. Check whether the value is an error. Genuine extreme values, such as large insurance claims, carry real information and should be kept or modelled separately. Removing them can understate risk.

What ethical issues apply to actuaries using data in India?

You must handle personal data lawfully and securely, use it only for its proper purpose, and follow IAI professional standards. Be honest about data limitations and avoid unfair bias. Check the current data protection law rather than relying on memory of sections.