Skip to content

Strategic Cost Management · Introduction to Tools for Data Analytics

Data Analytics Process and Data Sources Explained

Updated 11 October 2026 · Fact-checked

The data analytics process turns raw data into decisions through a sequence: define the question, collect data, clean it, process it, analyse it, interpret the results and communicate them. Data comes as structured, semi-structured or unstructured, and big data is described by the 5 Vs. In exams, name the steps in order and tie each to a cost decision.

Understand Data Analytics Process and Data Sources

Data analytics means examining data to find patterns that support a decision. In cost management, the decision may be which product to drop, where costs are leaking, or what a customer segment really earns you. The process matters because poor data, however clever the tool, gives a wrong answer.

The usual flow has these stages. First, define the objective: what question are you answering? Second, collect data from ERP records, cost sheets, invoices, sensors, websites and so on. Third, clean the data: remove duplicates, fix errors, handle missing values and standardise formats. Fourth, process or transform it: merge sources, group, code and structure it for use. Fifth, analyse it using tools such as descriptive statistics, regression or visualisation. Sixth, interpret the output in business terms. Last, report and act. Textbooks vary slightly in how they name and group these steps, so focus on the logical order.

Structured data sits in fixed rows and columns with a defined format, like a ledger, a bill of materials or a payroll table in a database. Unstructured data has no predefined format: emails, contracts, images, audio, social media posts. Semi-structured data has some tags or markers but no rigid table, such as XML or JSON files. Structured data is easy to query with SQL. Unstructured data needs extra processing before it can be analysed.

Big data means data so large, fast or varied that ordinary tools struggle to handle it. It is commonly described by the Vs: Volume (how much), Velocity (how fast it arrives), Variety (different forms), Veracity (how reliable it is) and Value (the usefulness you can extract). Some sources list only three or four Vs, so state the ones you use.

For a CMA, the link is practical. Better data means better cost allocation, more reliable variance analysis and sharper pricing and product decisions.

Key rules to remember

Analytics process sequence
Objective → Collection → Cleaning → Processing → Analysis → Interpretation → Reporting
No numerical formula applies. Write the steps in this order; cleaning always comes before analysis.
Big data 5 Vs
Volume, Velocity, Variety, Veracity, Value
Some texts list only 3 or 4 Vs. State which set you use and define each in one line.
Data types
Structured (tabular, fixed schema) | Semi-structured (tagged, flexible) | Unstructured (no schema)
Give one business example of each, such as a cost ledger, a JSON file and customer emails.
Common cleaning actions
Remove duplicates, correct errors, treat missing values, standardise formats, check outliers
Use these as a checklist when asked what cleaning involves.

How to solve Data Analytics Process and Data Sources questions

This topic is mostly descriptive. Use one method for any question, whether it is a short note, an MCQ or a case.

  1. 1Read the question and decide what is asked: a process, a data type, big data, or a case application.
  2. 2If it is a process question, list the stages in logical order and give one line for each.
  3. 3Link each stage to the case data given, for example sales records, machine sensor logs or supplier invoices.
  4. 4For data type questions, classify each item as structured, semi-structured or unstructured, and say why.
  5. 5For big data questions, define each V and attach a short business example.
  6. 6Point out risks, such as poor data quality, privacy or cost of tools, if the question asks for limitations.
  7. 7End with a clear conclusion or recommendation tied to the cost decision.

Quickest way: Order-and-classify shortcut

When to use it: Use it for MCQs and short notes when you have under two minutes.

  1. For sequence questions, remember: Define, Collect, Clean, Process, Analyse, Interpret, Report. Eliminate any option that analyses before cleaning.
  2. For data type, ask: does it fit neatly in rows and columns? Yes means structured. Partly tagged means semi-structured. Free text, image or audio means unstructured.
  3. For big data, say each V as a question: How much? How fast? How varied? How accurate? How useful?
  4. Check that your option matches the exact term asked, since the options often mix up Veracity and Value.

Common mistakes in Data Analytics Process and Data Sources

  • Placing analysis before data cleaning

    Students think analysis is the main step and rush to it.

    Fix: Remember that dirty data gives wrong results. Cleaning always comes before analysis.

  • Calling emails or PDFs structured because they are stored on a computer

    Storage in a digital file is confused with a fixed format.

    Fix: Judge by schema. If there are no fixed fields, it is unstructured.

  • Mixing up Veracity and Value

    Both start with V and both relate to quality.

    Fix: Veracity is trustworthiness of data. Value is the benefit gained from it.

  • Treating big data as only large volume

    The word 'big' suggests size alone.

    Fix: Mention velocity and variety as well, with an example of each.

  • Writing generic answers with no cost or business link

    Students recall textbook definitions without reading the case.

    Fix: Use the case's data, such as production logs or vendor invoices, in every stage you describe.

  • Saying cleaning means deleting all incomplete records

    Deletion looks like the simplest fix.

    Fix: Explain that missing values can be corrected, estimated or flagged, and deleting is one option, not the only one.

Worked examples

Example 1

A manufacturing company in Pune wants to find why scrap cost has risen. It has machine sensor readings, scrap entries in its ERP, and supervisor notes in free text. Classify each data source and list the steps to analyse the issue.

Show the solution
  1. Classify ERP scrap entries: structured, as they sit in fixed fields such as date, batch and quantity.
  2. Classify sensor readings: usually structured or semi-structured, depending on the log format, for example timestamped records or JSON.
  3. Classify supervisor notes: unstructured, as they are free text.
  4. Define the objective: identify causes of the rise in scrap cost.
  5. Collect the three sources for the same period.
  6. Clean: remove duplicate scrap entries, fix batch codes that do not match, and treat missing sensor readings.
  7. Process: merge the data by batch and machine, and convert text notes into coded causes where possible.
  8. Analyse: compare scrap rate by machine, shift and material, and test which factors move with scrap.
  9. Interpret and report: state the main causes and the cost impact, and recommend corrective action.

Answer: ERP data is structured, sensor data is structured or semi-structured, and supervisor notes are unstructured. The analysis follows objective, collection, cleaning, processing, analysis, interpretation and reporting.

Example 2

Which of the following describes the V called Veracity in big data? (A) The speed at which data is generated (B) The reliability and accuracy of data (C) The variety of formats in which data comes (D) The total amount of data stored

Show the solution
  1. Recall the 5 Vs: Volume is amount, Velocity is speed, Variety is forms, Veracity is reliability, Value is usefulness.
  2. Option A describes velocity.
  3. Option C describes variety.
  4. Option D describes volume.
  5. Option B describes reliability and accuracy, which is veracity.

Answer: (B) The reliability and accuracy of data

Exam tips

  • Write the process stages in order and add a one-line explanation for each. Marks are usually for sequence plus application.
  • In case questions, name the actual data in the scenario for each stage rather than giving generic text.
  • Prepare one business example each for structured, semi-structured and unstructured data and the 5 Vs.
  • Expect MCQs that test classification or the meaning of a V. Read options carefully, since distractors are similar terms.
  • State limitations such as data quality, privacy and tool cost when a question asks for an evaluation.

Practice questions from Introduction to Tools for Data Analytics

Data Analytics Process and Data Sources in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Data Analytics Process and Data Sources: frequently asked questions

What are the steps in the data analytics process?

The common sequence is defining the objective, collecting data, cleaning, processing, analysing, interpreting and reporting. Books may group or name them slightly differently, but cleaning comes before analysis in every version. Write the steps in logical order and link them to the case.

What is the difference between structured and unstructured data?

Structured data follows a fixed format of rows and columns, such as a cost ledger or payroll table. Unstructured data has no predefined format, such as emails, images or audio. Structured data is easier to query, while unstructured data needs extra processing.

What is data cleaning and why is it needed?

Data cleaning means fixing errors in the data: removing duplicates, correcting wrong entries, treating missing values and standardising formats. It is needed because analysis on faulty data gives misleading results and poor cost decisions.

What are the 5 Vs of big data?

They are Volume, Velocity, Variety, Veracity and Value. They describe how much data there is, how fast it arrives, how varied it is, how reliable it is and how useful it is. Some sources list only three or four Vs, so state the set you use.