CMA Intermediate · Financial Management and Business Data Analytics
Introduction to Data Science for Business Decision-making: formula sheet
Key formulas
- Three components of data science
- Data science = Statistics and Mathematics + Computing + Domain knowledge
- A memory frame for definition answers. Missing any one component weakens the result.
- Characteristics of big data (the 3 Vs)
- Volume + Velocity + Variety
- Some books add Veracity and Value. Name the version your study material uses and list all the Vs you give.
- Flow from data to decision
- Data → Information → Insight → Decision
- Useful for explaining importance and scope in one line.
- Life cycle sequence
- Business understanding → Data collection/understanding → Data preparation → Exploratory analysis → Modelling → Evaluation → Deployment → Monitoring
- Iterative: evaluation or monitoring can send you back to earlier stages.
- CRISP-DM phases
- Business understanding → Data understanding → Data preparation → Modelling → Evaluation → Deployment
- Six phases. Use the grouping your study material gives if it differs slightly.
- Three types by format
- Structured (fixed schema) | Semi-structured (tags, flexible) | Unstructured (no model)
- Give one business example for each. Semi-structured sits between the other two.
- Big data 5 Vs
- Volume, Velocity, Variety, Veracity, Value
- Volume = size, Velocity = speed, Variety = formats, Veracity = quality and trust, Value = usefulness.
- Source classification
- Internal (inside the firm) | External (outside the firm)
- Internal: ERP, sales, payroll. External: government data, market reports, social media.
- Primary vs secondary
- Primary = first-hand, new collection | Secondary = already collected by others
- Based on who collected the data and for what purpose, not on location.
- Mean
- Mean = Σx ÷ n
- Used to fill missing numeric values when data has no extreme outliers.
- Median
- Middle value of data arranged in order; for even n, average of the two middle values
- Preferred for imputation when data is skewed or has outliers.
- Mode
- Most frequent value
- Used to fill missing categorical values such as city or product type.
- Interquartile range (IQR)
- IQR = Q3 − Q1
- Measures the spread of the middle 50% of the data.
- IQR rule for outliers
- Lower fence = Q1 − 1.5 × IQR; Upper fence = Q3 + 1.5 × IQR
- A common rule of thumb: values outside the fences are flagged as possible outliers. Flagged does not mean wrong.
- Z-score
- z = (x − mean) ÷ standard deviation
- Measures how many standard deviations a value is from the mean. An absolute value above 3 is often flagged as a possible outlier.
- Min-max scaling
- x' = (x − min) ÷ (max − min)
- Rescales values to the range 0 to 1.
- Missing percentage
- Missing % = (Number of missing values ÷ Total records) × 100
- Helps decide between deletion and imputation.
- Descriptive analytics
- Question: What happened? Tools: mean, median, totals, ratios, charts, dashboards
- Looks backward. No forecast or recommendation.
- Predictive analytics
- Question: What is likely to happen? Tools: regression, time series, classification
- Uses historical data. Output is a probability or estimate, not a guarantee.
- Prescriptive analytics
- Question: What should we do? Tools: optimisation, simulation, decision rules
- Needs an objective and constraints. Builds on predictions.
- Simple linear regression
- Y = a + bX
- Predicts a numeric value Y from X. Regression shows association; it does not by itself prove cause.
- Supervised vs unsupervised learning
- Supervised: labelled outcomes. Unsupervised: no labels
- Regression and classification are supervised. Clustering is unsupervised.
Quick revision
- Data science combines statistics, computing and domain knowledge to get insight from data for decisions.
- Life cycle runs from problem definition to data collection, preparation, exploration, modelling, interpretation and action.
- Always begin with a clear business problem, since it decides what data you need.
- Structured data fits rows and columns; unstructured data such as text, images and audio does not.
- Primary data is collected first-hand; secondary data already exists from another source.
- Data preparation covers cleaning, handling missing values, removing duplicates and treating outliers.
- EDA summarises and visualises data to spot patterns, trends and anomalies before modelling.
- Descriptive analytics asks what happened; predictive asks what is likely to happen; prescriptive asks what should we do.
- Prescriptive analytics builds on prediction and recommends actions.
- Typical uses include customer segmentation, demand forecasting, fraud detection and pricing.
- Ethical issues include privacy, consent, bias, transparency and data security.
- Good decisions need good data: poor quality input leads to poor output.
Common mistakes
- Treating data science and statistics as the same thing Fix: Say statistics is a theoretical foundation, while data science adds computing, large and unstructured data, and domain knowledge.
- Describing evolution as a list of random terms Fix: Present it as stages in time order, with the driver of each stage such as computers, internet data or cheaper computing.
- Starting the life cycle with modelling or data collection. Fix: Always begin with business understanding: the problem and the decision to be supported.
- Treating the cycle as a one-way straight line. Fix: State that it is iterative and give an example of looping back, such as poor evaluation leading to more data preparation.
- Calling emails unstructured without nuance, or calling all text unstructured. Fix: Email headers (sender, date, subject) are semi-structured; the free-text body is unstructured. Read what the question asks about.
- Treating JSON or XML as structured data. Fix: They use tags and flexible nesting without a fixed row-column schema, so classify them as semi-structured.
- Filling every missing value with the mean. Fix: Use the median when data is skewed or has outliers, and the mode for categorical fields. State your reason.
- Deleting every outlier automatically. Fix: Investigate first. A large genuine order is real information. Remove or correct only when it is an error.
- Calling a sales forecast descriptive analytics Fix: Ask whether the output is about the future. If yes, it is predictive.
- Treating predictive output as a certain result Fix: Say it is an estimate with error, based on past patterns that may change.
Exam tips
- Expect MCQs on definitions, the three components, the Vs of big data and which stage came when. Learn these as short lists.
- For written answers, a small comparison table with 4 to 5 points earns clear step marks.
- Write the evolution in time order. Examiners reward sequence and cause, not buzzwords.
- Always attach one Indian business example. It shows application, not just recall.
- There is no negative marking in the MCQ section, so attempt every question.
- Learn the stages in order and be ready to write them as a numbered list with one line each.
- If a case is given, use its business context in every stage instead of generic text.
- Mention that the process is iterative and that models need monitoring; many answers miss this.