Skip to content

ACCA Strategic Professional · Strategic Business Leader · Big data and data analytics

Nordell Bank uses a machine learning model to approve loans. An internal review finds the model rejects applicants from certain postcodes at much higher rates, although postcode is not a protected characteristic and the model was never given ethnicity data. Which risk does this BEST illustrate?

This illustrates algorithmic bias through proxy variables. Postcode correlates with protected characteristics, so the model can reproduce historic bias in training data and produce discriminatory outcomes even though ethnicity was excluded, exposing the bank to legal, ethical and reputational risk.

  1. AAlgorithmic bias arising from proxy variables in the training data, creating discrimination and reputational riskCorrect
  2. BData latency caused by batch processing of loan applications
  3. CLoss of data integrity through hardware failure
  4. DOver-reliance on structured data rather than unstructured data

Explanation

Postcode can act as a proxy for protected characteristics, so the model reproduces historic bias in the training data. Excluding the protected attribute does not remove discriminatory outcomes. The other options describe unrelated technical issues.

Did you get it right without looking?

One question tells you little. A timed set on Big data and data analytics shows your real accuracy, how long you take and where you lose marks.

More Big data and data analytics questions