Skip to content

CMA Intermediate · Financial Management and Business Data Analytics · Data Processing, Organisation, Cleaning and Validation

A finance team at an Indian retail company finds that the customer-master file lists 'Ravi Sharma, Mumbai' twice with identical PAN, phone number and address, each with a different customer ID. Which data-cleaning step directly addresses this problem?

Removing duplicate records after matching on key fields is correct. The same customer appears twice with identical PAN, phone and address, which is a duplication error. De-duplication keeps one master record. Imputation, scaling and winsorising address missing values, scale differences and outliers, not repeated records.

  1. AImputing the missing values with the column mean
  2. BRemoving duplicate records after matching on key fieldsCorrect
  3. CNormalising the values to a 0-1 scale
  4. DWinsorising the extreme values in the sales column

Explanation

Two records with identical identifying details represent the same customer, which is a duplicate-record problem. The cleaning step is de-duplication: match on key fields such as PAN and phone, then keep one master record. Imputation handles missing values, scaling changes the range of numbers, and winsorising caps outliers; none of them removes repeated records.

Did you get it right without looking?

One question tells you little. A timed set on Data Processing, Organisation, Cleaning and Validation shows your real accuracy, how long you take and where you lose marks.

More Data Processing, Organisation, Cleaning and Validation questions