Skip to content

CMA Intermediate · Financial Management and Business Data Analytics · Introduction to Data Science for Business Decision-making

In a dataset of 200 loan records, the 'Income' field is missing for 20 records. The analyst fills each blank with the average income of the 180 available records, which is ₹6,00,000. What is the main effect of this mean imputation on the Income variable?

Mean imputation keeps the average income at ₹6,00,000 but reduces the variance. The 20 filled values equal the mean, so they add no squared deviation while increasing the number of records, which understates the true spread of incomes.

  1. AIt increases the variance of income
  2. BIt leaves the mean unchanged at ₹6,00,000 but reduces the spread (variance)Correct
  3. CIt changes the mean to ₹5,40,000
  4. DIt removes all outliers in the column

Explanation

Adding 20 values equal to the existing mean keeps the overall mean at ₹6,00,000 (180×6,00,000 + 20×6,00,000 over 200). The imputed values deviate by zero from the mean, adding nothing to the sum of squared deviations while increasing the count, so variance shrinks. The mean-change option wrongly treats blanks as zeros.

Did you get it right without looking?

One question tells you little. A timed set on Introduction to Data Science for Business Decision-making shows your real accuracy, how long you take and where you lose marks.

More Introduction to Data Science for Business Decision-making questions