Skip to content

CS Professional · Artificial Intelligence, Data Analytics and Cyber Security - Laws and Practice · Internet and Other Technologies

Pune-based Kaveri Retail uses a Hadoop-style framework to store huge log files across many low-cost servers and process them in parallel. Which feature of this design best explains how it handles very large datasets?

Such frameworks distribute both storage and processing across a cluster of many commodity nodes, so large datasets are split up and processed in parallel. This allows scaling by adding nodes, unlike a single high-end server, which limits capacity.

  1. AStoring all data on one high-end server with a single processor
  2. BDistributing data and processing across a cluster of nodesCorrect
  3. CConverting all unstructured data into paper records before analysis
  4. DEncrypting data so that it cannot be processed at all

Explanation

Hadoop-style systems split data across the nodes of a cluster (distributed file system) and run processing in parallel on those nodes, allowing scaling on commodity hardware. A single server is the traditional approach that Big Data frameworks aim to overcome. The other options are not features of such frameworks.

Did you get it right without looking?

One question tells you little. A timed set on Internet and Other Technologies shows your real accuracy, how long you take and where you lose marks.

More Internet and Other Technologies questions