Skip to content

Financial Management and Business Data Analytics · Introduction to Data Science for Business Decision-making

Data Preparation and Exploratory Data Analysis Explained

Updated 10 October 2026 · Fact-checked

Data preparation turns raw data into a clean, consistent, analysis-ready dataset. You remove duplicates, fix errors, treat missing values and outliers, and transform variables. Exploratory data analysis (EDA) then summarises and charts the data to spot patterns, spread and anomalies before any modelling or decision is made.

Understand Data Preparation and Exploratory Data Analysis

Raw business data is rarely ready to use. It comes from invoices, ERP systems, surveys and web forms. It has typing errors, blank cells, duplicate rows and mixed formats. If you analyse it as it is, your answers will be wrong. This is the idea of "garbage in, garbage out".

Data preparation (also called data wrangling or data munging) is the work of fixing this. It has four parts: cleaning (remove duplicates, correct errors, standardise formats), handling missing values, handling outliers, and transformation (changing data into a more useful form, such as converting dates, scaling values or creating new columns).

Missing values can be handled in two broad ways. You can delete the record or column (deletion), or you can fill the gap (imputation) using the mean, median, mode, the previous value, or a value predicted from other fields. Choose based on how much is missing and why.

Outliers are values far away from the rest. They may be errors (a sales figure of ₹50,00,000 typed instead of ₹5,000) or genuine rare events (one very large corporate order). You must investigate before deciding to correct, cap, keep or remove them.

Exploratory data analysis (EDA) comes after cleaning. You look at the data to understand it: summary statistics (mean, median, standard deviation, minimum, maximum), frequency tables, and charts such as histograms, box plots and scatter plots. EDA shows the shape of the data, relationships between variables and any remaining problems. It guides which technique to use next.

Key rules to remember

Mean
Mean = Σx ÷ n
Used to fill missing numeric values when data has no extreme outliers.
Median
Middle value of data arranged in order; for even n, average of the two middle values
Preferred for imputation when data is skewed or has outliers.
Mode
Most frequent value
Used to fill missing categorical values such as city or product type.
Interquartile range (IQR)
IQR = Q3 − Q1
Measures the spread of the middle 50% of the data.
IQR rule for outliers
Lower fence = Q1 − 1.5 × IQR; Upper fence = Q3 + 1.5 × IQR
A common rule of thumb: values outside the fences are flagged as possible outliers. Flagged does not mean wrong.
Z-score
z = (x − mean) ÷ standard deviation
Measures how many standard deviations a value is from the mean. An absolute value above 3 is often flagged as a possible outlier.
Min-max scaling
x' = (x − min) ÷ (max − min)
Rescales values to the range 0 to 1.
Missing percentage
Missing % = (Number of missing values ÷ Total records) × 100
Helps decide between deletion and imputation.

How to solve Data Preparation and Exploratory Data Analysis questions

For any question on data preparation or EDA, follow this order. It also gives you a clear written structure.

  1. 1Understand the data: note the source, the variables, their types (numeric or categorical) and the business purpose.
  2. 2Inspect for quality problems: duplicates, wrong formats, inconsistent spellings, impossible values, missing values.
  3. 3Clean: remove duplicates, correct errors, standardise formats and units.
  4. 4Treat missing values: calculate the missing percentage, then choose deletion or imputation (mean, median, mode) and justify it.
  5. 5Detect and treat outliers: use the IQR rule or z-score, then investigate whether it is an error or a genuine value before acting.
  6. 6Transform as needed: convert types, scale, group into categories, or create new fields.
  7. 7Explore (EDA): compute summary statistics and choose suitable charts, then state what the data shows.
  8. 8Conclude: state the decision or next step in business language.

Quickest way: Clean, fill, flag, summarise

When to use it: Use this when time is short, especially for MCQs and short-note questions.

  1. Duplicates and errors: remove or correct.
  2. Missing numeric data: median if skewed or outliers exist, mean if roughly symmetric; mode for categories.
  3. Outliers: compute Q1, Q3, IQR and the two fences; flag anything outside.
  4. Scaling: apply (x − min) ÷ (max − min).
  5. EDA: write mean, median, spread, then name the chart that fits (histogram for distribution, box plot for outliers, scatter plot for relationship).

Common mistakes in Data Preparation and Exploratory Data Analysis

  • Filling every missing value with the mean.

    The mean is the first average students learn.

    Fix: Use the median when data is skewed or has outliers, and the mode for categorical fields. State your reason.

  • Deleting every outlier automatically.

    Students treat outliers as errors.

    Fix: Investigate first. A large genuine order is real information. Remove or correct only when it is an error.

  • Treating the IQR rule as proof of an error.

    The rule gives a clear numeric cut-off, so it feels final.

    Fix: Say values outside the fences are flagged as possible outliers, then verify against the source.

  • Doing EDA before cleaning.

    Students jump to charts and averages.

    Fix: Clean first. Duplicates and wrong entries distort summary statistics and charts.

  • Mixing up cleaning and transformation.

    Both change the data.

    Fix: Cleaning fixes mistakes and gaps. Transformation reshapes correct data, for example scaling or creating a new column.

  • Calculating quartiles from unsorted data.

    Students rush through the arithmetic.

    Fix: Always arrange values in ascending order before finding the median, Q1 and Q3.

Worked examples

Example 1

Monthly sales (₹ in thousands) of a Pune retailer for 9 months are: 42, 45, 47, 48, 50, 52, 53, 55, 160. Using the IQR rule, identify any outlier. Use the median of the lower four and upper four values as Q1 and Q3.

Show the solution
  1. The data is already sorted. n = 9, so the median is the 5th value = 50.
  2. Lower half (excluding the median): 42, 45, 47, 48. Q1 = (45 + 47) ÷ 2 = 46.
  3. Upper half: 52, 53, 55, 160. Q3 = (53 + 55) ÷ 2 = 54.
  4. IQR = 54 − 46 = 8.
  5. Lower fence = 46 − 1.5 × 8 = 46 − 12 = 34.
  6. Upper fence = 54 + 1.5 × 8 = 54 + 12 = 66.
  7. The value 160 is above 66, so it is flagged. All others lie between 34 and 66.
  8. Next step: check the source. If 160 is a typing error (for example 60), correct it. If it is a genuine bulk order, keep it and report it separately.

Answer: The sales figure of ₹1,60,000 (160 thousand) is flagged as a possible outlier. Verify it against records before correcting or removing it.

Example 2

A customer dataset of 200 records has the Age field missing in 20 records. The known ages (in years) of 5 sample customers are 28, 30, 31, 33 and 78. Calculate the missing percentage, and decide whether to fill the gaps with the mean or the median of these five. Show the effect.

Show the solution
  1. Missing % = (20 ÷ 200) × 100 = 10%.
  2. 10% is modest, so imputation is better than deleting 20 records.
  3. Mean of sample = (28 + 30 + 31 + 33 + 78) ÷ 5 = 200 ÷ 5 = 40.
  4. Median of sample: sorted values are 28, 30, 31, 33, 78, so the median = 31.
  5. The value 78 pulls the mean up to 40, which is higher than four of the five customers. The median of 31 represents a typical customer better.
  6. Fill missing ages with the median, and check whether 78 is genuine.

Answer: Missing percentage is 10%. Use median imputation (31 years), because the mean of 40 is distorted by the high value 78.

Exam tips

  • For 'explain' or 'discuss' questions, follow the sequence: clean, missing values, outliers, transform, explore. Examiners reward a logical order.
  • Always justify your choice (mean, median, deletion). A choice without a reason loses marks.
  • In numerical questions, sort the data first and show Q1, Q3, IQR and both fences as separate lines for step marks.
  • In MCQs, watch the wording: 'most suitable for skewed data' points to the median, 'categorical field' points to the mode.
  • Link EDA to a business use, such as spotting a slow-selling product or an unusual expense, to show application.

Practice questions from Introduction to Data Science for Business Decision-making

Data Preparation and Exploratory Data Analysis in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Data Preparation and Exploratory Data Analysis: frequently asked questions

What is data wrangling with an example?

Data wrangling means converting raw data into a usable form. For example, a sales file may have dates as 05/03/26, 5-Mar-2026 and March 5, 2026. You convert all of them to one format, remove duplicate invoices and fill blank city names. After that, the file is ready for analysis.

How do I handle missing values in data?

First find how much is missing and why. If very little is missing, you can delete those records. Otherwise fill the gaps using the mean, median or mode, or a value predicted from other fields. Use the median for skewed numeric data and the mode for categories.

What is exploratory data analysis?

EDA is the first look at cleaned data to understand its shape, spread, relationships and problems. You use summary statistics and charts such as histograms, box plots and scatter plots. It helps you decide which analysis or model to use next.

Is an outlier always an error?

No. An outlier can be a data entry mistake or a genuine rare event. You should check the source before removing it. Removing genuine outliers can hide important business information.