Skip to content

Financial Management and Business Data Analytics · Introduction to Data Science for Business Decision-making

Types of Data and Data Sources for Business Decisions

Updated 10 October 2026 · Fact-checked

Business data is classified by format as structured (fixed rows and columns), semi-structured (tagged but flexible, like JSON or XML) and unstructured (no fixed model, like emails and images). Big data is described by the 5 Vs: volume, velocity, variety, veracity and value. Sources are internal (own records) or external (outside the firm). To answer, classify, justify, then give an example.

Understand Types of Data and Data Sources

Data is a set of raw facts and figures. It becomes information once it is processed and useful for a decision. A business collects data from many places and in many formats. Knowing the type and source tells you how to store it, how to analyse it and how far to trust it.

By format, data falls into three types. Structured data follows a fixed model of rows and columns with defined fields, such as a ledger, a sales register or a customer table in a database. It is easy to search with SQL and to total in a spreadsheet. Semi-structured data has no rigid table but carries tags or markers that give it some order. Examples are JSON files, XML files, HTML pages and email headers. Unstructured data has no predefined model. Examples are the body of emails, PDF contracts, social media posts, photos, audio and video. Most of the data firms hold today is unstructured, but it needs more effort to analyse.

Big data means data so large, fast or varied that ordinary tools struggle to handle it. It is commonly described by the 5 Vs. Volume is the sheer size. Velocity is the speed at which data is created and must be processed. Variety is the mix of formats: structured, semi-structured and unstructured. Veracity is how accurate and reliable the data is. Value is the usefulness of the data for decisions. Some books list only three Vs (volume, velocity, variety) or add more, so answer with the five unless the question says otherwise.

By source, data is internal or external. Internal data comes from inside the organisation: accounting records, sales and purchase data, inventory and payroll records, CRM data and production logs. It is usually cheap, quick and relevant. External data comes from outside: government publications, RBI and SEBI data, stock exchange data, industry reports, trade bodies, competitors' published results, surveys and social media. It adds market context but may need checking for reliability.

A related split is primary and secondary data. Primary data is collected first-hand for your own purpose, for example through a customer survey or interview. Secondary data was collected earlier by someone else, such as a published RBI report. Internal records can be primary or secondary depending on who collected them and why, so do not treat internal and primary as the same thing.

Key rules to remember

Three types by format
Structured (fixed schema) | Semi-structured (tags, flexible) | Unstructured (no model)
Give one business example for each. Semi-structured sits between the other two.
Big data 5 Vs
Volume, Velocity, Variety, Veracity, Value
Volume = size, Velocity = speed, Variety = formats, Veracity = quality and trust, Value = usefulness.
Source classification
Internal (inside the firm) | External (outside the firm)
Internal: ERP, sales, payroll. External: government data, market reports, social media.
Primary vs secondary
Primary = first-hand, new collection | Secondary = already collected by others
Based on who collected the data and for what purpose, not on location.

How to solve Types of Data and Data Sources questions

Use this routine for any question on data types, big data or sources. It works for MCQs and for short written answers.

  1. 1Read the question and spot the key word: type, format, source, characteristic, or difference.
  2. 2If it gives a data item, decide whether it has a fixed row-column model (structured), tags without a fixed table (semi-structured) or neither (unstructured).
  3. 3If it asks about big data, list the Vs asked for, name them in order, and add a one-line meaning of each.
  4. 4If it asks about source, ask whether the data comes from inside or outside the firm, then whether it was collected first-hand or earlier by others.
  5. 5For a difference question, write two or three points side by side in the same order: format, storage or tools, example.
  6. 6Add a business example from an Indian context, such as a retailer, bank or manufacturer.
  7. 7Close with one line on why the classification matters for analysis or decisions.

Quickest way: Three-question test for any data item

When to use it: Use in the MCQ section when you must classify a data item or source in under a minute.

  1. Can it sit in a table with fixed columns? If yes, it is structured.
  2. If not, does it carry tags or keys, like JSON, XML or HTML? If yes, it is semi-structured. Otherwise it is unstructured.
  3. Did the firm generate it internally, or does it come from outside? Then check whether it was collected first-hand (primary) or earlier by others (secondary).

Common mistakes in Types of Data and Data Sources

  • Calling emails unstructured without nuance, or calling all text unstructured.

    Students think text always means unstructured.

    Fix: Email headers (sender, date, subject) are semi-structured; the free-text body is unstructured. Read what the question asks about.

  • Treating JSON or XML as structured data.

    They look orderly, so they seem like tables.

    Fix: They use tags and flexible nesting without a fixed row-column schema, so classify them as semi-structured.

  • Mixing up Veracity and Value.

    Both words sound like quality.

    Fix: Veracity is about accuracy and trustworthiness. Value is about usefulness for decisions.

  • Assuming internal data is always primary and external data is always secondary.

    The two pairs of terms are learned together.

    Fix: Internal/external is about location of the source. Primary/secondary is about who collected the data and why.

  • Listing the Vs without meanings.

    Students memorise the names only.

    Fix: Add a short meaning and a business example for each V to earn full marks.

Worked examples

Example 1

Classify each of the following as structured, semi-structured or unstructured: (a) a company's GST sales register in a database table, (b) customer feedback as JSON files from a mobile app, (c) recorded customer-care calls. Give a reason for each.

Show the solution
  1. (a) The sales register has fixed columns such as invoice number, date, party and amount. It follows a fixed schema, so it is structured.
  2. (b) JSON files use keys and tags but no fixed table layout, and fields can vary between records. This is semi-structured.
  3. (c) Audio recordings have no predefined data model and cannot be placed in rows and columns directly. This is unstructured.

Answer: (a) Structured; (b) Semi-structured; (c) Unstructured.

Example 2

An Indian e-commerce company records millions of orders every hour, from app clicks, payment tables and product photos. Some records have wrong pincodes. Explain which of the 5 Vs each feature illustrates.

Show the solution
  1. Millions of orders in total points to Volume, the large size of data.
  2. Orders arriving every hour and needing quick processing points to Velocity, the speed of data creation.
  3. Clicks (semi-structured), payment tables (structured) and photos (unstructured) show Variety, the mix of formats.
  4. Wrong pincodes affect accuracy and reliability, which is Veracity.
  5. If the firm cleans the data and uses it to plan deliveries and offers, it gains Value, the usefulness for decisions.

Answer: Volume: millions of orders; Velocity: hourly flow; Variety: clicks, tables and photos; Veracity: wrong pincodes; Value: better delivery and offer decisions after analysis.

Exam tips

  • For MCQs, look for the format clue: fixed table, tags, or free form. This settles most classification questions.
  • Write the 5 Vs in the standard order and give each a one-line meaning; a bare list earns fewer marks.
  • For difference questions, use the same points for each type: format, tools, example.
  • Give Indian business examples such as GST records, bank transactions, UPI data or retailer sales.
  • When asked about sources, state internal or external first, then say whether it is primary or secondary only if the question asks.

Practice questions from Introduction to Data Science for Business Decision-making

Types of Data and Data Sources in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Types of Data and Data Sources: frequently asked questions

What is the difference between structured and unstructured data?

Structured data fits a fixed row-and-column schema, like an accounting ledger, and is easy to query. Unstructured data has no predefined model, like emails, images and videos, and needs special tools to analyse. Semi-structured data lies between them, using tags such as JSON or XML.

What are the 5 Vs of big data?

They are Volume, Velocity, Variety, Veracity and Value. Volume is size, velocity is speed, variety is the mix of formats, veracity is accuracy and reliability, and value is usefulness for decisions.

What are internal and external sources of business data?

Internal sources are inside the firm, such as sales, inventory, payroll and ERP records. External sources are outside, such as government statistics, regulator data, industry reports and social media. A good analysis often combines both.

Is primary data the same as internal data?

No. Primary data is collected first-hand for your own purpose, like a fresh customer survey. Secondary data was collected earlier by someone else. Internal data is simply data held within the firm and may be either, depending on how it was collected.