Artificial Intelligence, Data Analytics and Cyber Security - Laws and Practice · Data Analytics
Types of Data and Data Sources in Data Analytics
Updated 11 October 2026 · Fact-checked
Data is classed by format as structured (fixed rows and columns), semi-structured (tagged but flexible, like JSON or XML) or unstructured (no set model, like emails and video). Big data is described by its Vs: volume, velocity, variety, veracity and value. Sources are internal, external, primary or secondary.
Understand Types of Data and Data Sources
Data is raw facts and figures that have not yet been processed. Data analytics turns this raw data into useful findings. Before you analyse anything, you must know what form the data is in and where it came from. That is why this topic comes first in the chapter.
Data is classed by its format. Structured data follows a fixed model of rows and columns with defined fields. A company's ledger, a share register or a table in a relational database is structured. You can search it easily with SQL. Unstructured data has no predefined model. Emails, PDFs of board minutes, images, audio, video and social media posts are unstructured. It forms most of the data organisations hold, and it is harder to analyse. Semi-structured data sits between the two. It has no rigid table, but it carries tags or markers that separate fields and show hierarchy. XML, JSON, HTML pages and log files are common examples.
Big data means data sets so large, fast or varied that ordinary tools cannot handle them well. It is described by the Vs. The core three are volume (the size of data), velocity (the speed at which it is created and processed) and variety (the different formats). Many texts add veracity (how accurate and trustworthy the data is) and value (the usefulness you can get from it). Some texts add more Vs, such as variability. State clearly which list you are using.
Data sources are the places data comes from. Internal sources are inside the organisation, such as accounting records, ERP systems, HR files and CRM data. External sources are outside it, such as government portals, MCA filings, stock exchange data, market reports and social media. A second split is by how you obtain it. Primary data is collected first-hand for your own purpose, through surveys, interviews or sensors. Secondary data was collected earlier by someone else and is reused. Machine-generated sources include IoT sensors, server logs and transaction systems.
In a corporate setting, the type and source of data also decide the legal duties. Personal data in emails or customer records raises privacy and security issues, which the later topics on legal and ethical issues cover.
Key rules to remember
- Three-way classification by format
- Data = Structured + Semi-structured + Unstructured
- Classify by whether the data follows a fixed schema (structured), carries tags without a rigid schema (semi-structured) or has no model (unstructured).
- Core Vs of big data
- Volume, Velocity, Variety
- These three are the classic characteristics. Always give them first.
- Extended Vs of big data
- Veracity (quality and trust), Value (usefulness)
- Add these for a 5 Vs answer. Mention that some sources list more Vs.
- Classification of sources
- Internal vs External; Primary vs Secondary
- Two separate splits. Internal and external depend on location. Primary and secondary depend on who collected the data first.
How to solve Types of Data and Data Sources questions
Use this method for any question on types of data, big data characteristics or data sources.
- 1Read the command word: define, distinguish, explain, discuss or illustrate. It sets the depth and the layout.
- 2Start with a one-line definition of the key term, for example data, big data or structured data.
- 3List the categories or Vs in a clear order. Name each one in bold or underline it.
- 4Explain each point in one or two lines and add an Indian corporate example, such as a company's share register, board minutes, or stock exchange feeds.
- 5For a 'distinguish' question, draw a two-column comparison on points like format, storage, tools, ease of analysis and examples.
- 6Link to practice: say how the type of data affects storage, analysis or compliance, such as privacy duties for personal data.
- 7Close with a one-line conclusion that answers the question asked.
Quickest way: Format, Vs, Source in three lines
When to use it: Use it when you have under five minutes for a short-answer question on this topic.
- Write the definition in one line.
- Give the list in the right order: three formats, or the Vs, or internal and external sources.
- Add one company example for each item, then stop.
Common mistakes in Types of Data and Data Sources
Calling emails, JSON and XML all unstructured, or all semi-structured.
Students remember that they are not tables and stop there.
Fix: Emails as free text, images and video are unstructured. JSON, XML and HTML carry tags, so they are semi-structured.
Listing only three Vs when the question says 5 Vs, or the other way round.
Different books give three, four, five or more Vs.
Fix: Follow the number in the question. For five, give volume, velocity, variety, veracity and value. Name the extra ones if asked for more.
Mixing up velocity and volume.
Both sound like 'a lot of data'.
Fix: Volume is how much data. Velocity is how fast it arrives and must be processed. Think of stock tick data for velocity.
Confusing primary and secondary data with internal and external data.
Both pairs describe the origin of data.
Fix: Internal and external show where data sits relative to the organisation. Primary and secondary show whether you collected it yourself. An internal record can be secondary if it was gathered for another purpose.
Writing definitions with no example.
Students rush and treat it as pure theory.
Fix: Give at least one Indian corporate example for each type or source, since the paper tests applied understanding.
Worked examples
Example 1
Distinguish between structured and unstructured data with examples.
Show the solution
- Define both: structured data follows a fixed schema of rows and columns; unstructured data has no predefined model.
- Compare on format: structured data sits in tables with defined fields; unstructured data is free-form text, images, audio or video.
- Compare on storage and tools: structured data is held in relational databases and queried with SQL; unstructured data needs file stores, data lakes or specialised tools such as text mining.
- Compare on ease of analysis: structured data is easy to search and sort; unstructured data needs processing before it yields insight.
- Give examples: a company's general ledger and share transfer register are structured; board meeting audio, scanned agreements and customer emails are unstructured.
- Note that semi-structured data, such as JSON or XML files, lies between the two.
Answer: Structured data fits a fixed table format and is easy to query with SQL, as in a ledger. Unstructured data has no set model and needs special tools, as in emails, scanned contracts and videos. Semi-structured data, like JSON or XML, carries tags but no rigid schema.
Example 2
Explain the characteristics of big data. How could a listed company's compliance team meet them in practice?
Show the solution
- Define big data: very large, fast and varied data sets that ordinary tools cannot manage well.
- Explain volume: the size of data, for example years of filings, trades and customer records.
- Explain velocity: the speed of data generation and processing, for example live market feeds and online transactions.
- Explain variety: many formats at once, such as tables, XML filings, emails and social media posts.
- Explain veracity: the accuracy and reliability of data, since poor data leads to wrong decisions.
- Explain value: data matters only if analysis gives useful insight, such as spotting unusual transactions.
- Apply to the company: the team may combine internal records with external sources like exchange disclosures, and must check data quality and protect personal data.
Answer: Big data is described by volume, velocity, variety, veracity and value. A compliance team handles large amounts of mixed data arriving quickly, checks its accuracy, and uses analysis to gain useful insights while protecting personal data.
Exam tips
- For 'distinguish' questions, use a short two-column comparison with at least four points. It scores better than long paragraphs.
- Give the Vs in a fixed order and say how many you are listing. If the question does not give a number, give five and note that some texts add more.
- Attach an Indian corporate example to every category. It shows application, not rote learning.
- Spend a line on legal relevance, such as personal data in unstructured sources. This links the topic to the rest of the paper.
- In case-based questions, first identify the type and source of data in the facts, then give your analysis.
Practice questions from Data Analytics
- A company's analyst examines whether last year's rise in employee attrition was linked to delayed appraisals by comparing attrition across d…
- A retail chain in Pune studies its past sales records to find out why sales of winter wear fell sharply last year. Which type of analytics i…
- A Mumbai-based retailer collects customers' purchase histories and wants to run analytics on them. Under the data protection principle of pu…
- A bank's analytics team groups its customers into segments based on similar spending behaviour without using any pre-labelled categories. Wh…
- A private bank in India uses historical loan repayment records and customer attributes to estimate the probability that each new applicant w…
Types of Data and Data Sources in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Types of Data and Data Sources: frequently asked questions
What is the difference between structured and unstructured data?
Structured data follows a fixed table layout with defined fields and can be queried easily with SQL. Unstructured data has no predefined model, such as emails, images and videos. It needs special tools before it can be analysed.
What are the 5 Vs of big data?
They are volume, velocity, variety, veracity and value. Volume is size, velocity is speed, variety is format mix, veracity is trustworthiness and value is usefulness. Some books list more Vs, so follow the number in the question.
Is JSON structured or unstructured data?
JSON is usually treated as semi-structured data. It has tags and key-value pairs that give it some organisation, but it does not follow the rigid table model of a relational database.
What are the main sources of data?
Sources can be internal, such as company accounts, ERP and HR records, or external, such as government portals, exchange data and social media. They can also be primary, collected first-hand, or secondary, collected earlier by others.