Skip to content

CS Professional · Artificial Intelligence, Data Analytics and Cyber Security - Laws and Practice · Computer Hardware and Software

A company wants to store large volumes of unstructured data such as emails, images and scanned contracts for later analytics, without first fixing a schema. Which storage approach best fits?

A data lake fits best because it stores raw structured and unstructured data in native form without requiring a predefined schema, applying structure only when the data is analysed. Fixed-schema relational tables, spreadsheets and CPU caches do not suit large, varied, schema-free data.

  1. AA data lakeCorrect
  2. BA normalised relational table with fixed columns
  3. CA spreadsheet with macros
  4. DA CPU cache

Explanation

A data lake holds raw data in native formats, structured or unstructured, with the schema applied only when read. A normalised relational table needs a predefined schema, spreadsheets do not scale, and a cache is temporary processor memory.

Did you get it right without looking?

One question tells you little. A timed set on Computer Hardware and Software shows your real accuracy, how long you take and where you lose marks.

More Computer Hardware and Software questions