Skip to content

CMA Intermediate · Financial Management and Business Data Analytics

Data Analysis and Modelling: formula sheet

Full chapter guide

Key formulas

Mean imputation
Mean = Σx ÷ n (over available values only)
Divide by the count of values present, not the total number of records. Use when data has no strong outliers.
Median imputation
Median = middle value of the sorted data (average of the two middle values if n is even)
Preferred when data is skewed or has outliers.
Mode imputation
Mode = most frequent value
Used for categorical data such as city or product category.
IQR rule for outliers
IQR = Q3 − Q1; lower fence = Q1 − 1.5 × IQR; upper fence = Q3 + 1.5 × IQR
A value outside the fences is flagged as a possible outlier. It is a screening rule, not proof of an error.
Min-max scaling
x′ = (x − min) ÷ (max − min)
Rescales values to the range 0 to 1.
Missing-value percentage
Missing % = (Number of missing values ÷ Total records) × 100
Helps decide between deleting and imputing.
Arithmetic mean (ungrouped)
x̄ = Σx ÷ n
Add all values and divide by the number of values.
Arithmetic mean (grouped)
x̄ = Σfx ÷ Σf
x is the class mid-point, f is the frequency. Mid-point = (lower limit + upper limit) ÷ 2.
Weighted mean
x̄w = Σwx ÷ Σw
Use when values carry different importance, such as prices with quantities.
Median (ungrouped)
Sort the data. Odd n: middle value. Even n: average of the two middle values.
Position of the middle value for odd n is (n + 1) ÷ 2.
Median (grouped)
Median = L + [(N ÷ 2 − cf) ÷ f] × h
L = lower limit of median class, N = Σf, cf = cumulative frequency before the median class, f = its frequency, h = class width. The median class is the first class whose cumulative frequency reaches N ÷ 2.
Mode (grouped)
Mode = L + [(f1 − f0) ÷ (2f1 − f0 − f2)] × h
f1 = frequency of modal class, f0 = frequency of class before it, f2 = frequency of class after it. Classes must have equal width.
Empirical relation
Mode ≈ 3 × Median − 2 × Mean
An approximation for moderately skewed data, not an exact law. Use it only when the question asks or the mode is missing.
Range and coefficient of range
Range = Largest − Smallest; Coefficient of range = (Largest − Smallest) ÷ (Largest + Smallest)
Quick but depends only on the two extreme values.
Quartile deviation
QD = (Q3 − Q1) ÷ 2
Based on the middle 50% of the data, so it is not affected by extreme values.
Variance and standard deviation (population)
σ² = Σ(x − x̄)² ÷ n; σ = √σ²; grouped: σ² = Σf(x − x̄)² ÷ Σf
Use this when the data covers the whole group. For a sample, divide by (n − 1) instead of n.
Coefficient of variation
CV = (σ ÷ x̄) × 100
Relative spread in percentage. The series with the lower CV is more consistent.
Karl Pearson's coefficient of skewness
Sk = (Mean − Mode) ÷ σ
Positive means right-skewed, negative means left-skewed, near zero means roughly symmetric.
Bowley's coefficient of skewness
Sk = (Q3 + Q1 − 2 × Median) ÷ (Q3 − Q1)
Quartile-based, useful when extreme values distort the mean and standard deviation.
Relative frequency
Relative frequency = f ÷ Σf
Multiply by 100 for a percentage. All relative frequencies add up to 1 (or 100%).
Chart choice rule: comparison
Compare categories → bar or column chart
Categories are separate groups such as products, branches or departments.
Chart choice rule: trend
Change over time → line graph
Time goes on the horizontal axis. Use it for sales, cost or profit across periods.
Chart choice rule: composition
Parts of a whole → pie chart or stacked bar
Use pie only with few categories (about five or fewer) that add to 100%.
Chart choice rule: distribution
Spread of continuous data → histogram
Bars touch because class intervals are continuous.
Chart choice rule: relationship
Two numeric variables → scatter plot
Shows direction and strength of association. It does not prove cause.
Pie chart angle
Angle of a sector = (Value of item ÷ Total) × 360°
All angles must add to 360°.
Percentage share
Share (%) = (Value of item ÷ Total) × 100
Used for labels in pie charts and stacked charts.
Dashboard components
KPIs + visuals + filters / drill-down on one screen
A dashboard supports monitoring and quick decisions.
Descriptive analytics
Question: What happened? Input: past data. Output: summaries, reports, charts
Examples: sales dashboard, average cost per unit, branch-wise totals.
Diagnostic analytics
Question: Why did it happen? Input: past data split by factors. Output: causes and relationships
Examples: drill-down, variance analysis, correlation between two variables.
Predictive analytics
Question: What is likely to happen? Input: historical data and models. Output: forecasts and probabilities
Examples: demand forecast, credit default risk score, regression-based estimate.
Prescriptive analytics
Question: What should we do? Input: predictions plus constraints and objectives. Output: recommended action
Examples: optimal product mix, inventory plan, route optimisation.
Order of the four types
Descriptive → Diagnostic → Predictive → Prescriptive
Moves from hindsight to insight to foresight to action.
Karl Pearson's correlation coefficient
r = [nΣXY − ΣX·ΣY] ÷ √{[nΣX² − (ΣX)²] × [nΣY² − (ΣY)²]}
Use when you have raw paired data. r is always between −1 and +1.
Correlation using covariance
r = Cov(X, Y) ÷ (σx × σy)
Cov(X, Y) = ΣXY ÷ n − X̄·Ȳ. Use when the question gives covariance and standard deviations.
Regression line of Y on X
Y = a + bX, where b = [nΣXY − ΣX·ΣY] ÷ [nΣX² − (ΣX)²] and a = Ȳ − b·X̄
Use this to predict Y from X. The line always passes through (X̄, Ȳ).
Regression coefficients using r
b(yx) = r × σy ÷ σx ; b(xy) = r × σx ÷ σy
b(yx) is for Y on X. b(xy) is for X on Y. Both have the same sign as r.
Link between r and the two coefficients
r² = b(yx) × b(xy)
r = √(b(yx) × b(xy)), with the sign of the coefficients. Both coefficients must have the same sign.
Coefficient of determination
R² = 1 − SSE ÷ SST
SSE is the sum of squared errors. SST is the total sum of squares. In simple regression, R² = r².
Adjusted R²
Adjusted R² = 1 − (1 − R²) × (n − 1) ÷ (n − k − 1)
k is the number of independent variables. It penalises adding variables that add little.
Spearman's rank correlation (no tied ranks)
ρ = 1 − 6Σd² ÷ [n(n² − 1)]
d is the difference between the two ranks of each item. Use when data are ranks.
Additive model
Y = T + S + C + I
Use when seasonal swings stay about the same size.
Multiplicative model
Y = T × S × C × I
Use when seasonal swings grow with the level. Seasonal indices are then ratios or percentages.
Simple moving average (odd period, n = 3)
MA = (Y₁ + Y₂ + Y₃) ÷ 3
Place the result against the middle period, here the second one.
Centred moving average (even period, n = 4)
Centred MA = (sum of two consecutive 4-period averages) ÷ 2
A 4-period average falls between two periods, so average two consecutive ones to align with a real period.
Moving average forecast
Forecast for next period = average of the last n actual values
The simplest forecast. It lags behind if there is a strong trend.
Least squares trend line
Y = a + bX, where b = (nΣXY − ΣX·ΣY) ÷ (nΣX² − (ΣX)²) and a = (ΣY − bΣX) ÷ n
X is time, coded 0, 1, 2, ... or around zero.
Trend line with ΣX = 0
a = ΣY ÷ n and b = ΣXY ÷ ΣX²
Code time so the middle period is 0. Easy when n is odd (…, −2, −1, 0, 1, 2, …).
Seasonal index (average method)
Seasonal index = (average of that season ÷ grand average of all values) × 100
The indices over a full year average 100.
Seasonal index (ratio to moving average)
Ratio = (Actual Y ÷ centred moving average) × 100
Average the ratios for each season across years, then adjust so the indices average 100.
Data model vs analytical model
Data model = structure of data; Analytical model = logic applied to data to answer a question
The most common distinction asked in theory questions.
Relational link
Primary key (unique row identifier) ↔ Foreign key (reference in another table)
Use this to explain how tables are related in Power BI or a database.
Expected value for a decision model
EV = Σ (Probability × Outcome)
Choose the option with the highest EV for profit, or the lowest for cost, when the decision is based on expected value.
Simulation average
Average outcome = Σ (outcomes of all runs) ÷ number of runs
Simulation gives a range of results, not one certain answer.

Quick revision

  • Cleaning comes before analysis: fix missing values, duplicates, errors and outliers first.
  • Mean is affected by extreme values; median is not.
  • Standard deviation is the square root of variance.
  • Correlation lies between -1 and +1; the sign shows direction and the size shows strength.
  • Correlation shows association and does not prove that one variable causes the other.
  • In simple regression, Y = a + bX, where b is the slope and a is the intercept.
  • The regression coefficient b shows the change in Y for a one-unit change in X.
  • Descriptive analytics says what happened; diagnostic says why; predictive says what may happen; prescriptive says what to do.
  • Match the chart to the purpose: lines for trends over time, bars for comparing categories.
  • A moving average smooths short-term fluctuations to show the trend.
  • Forecasts are estimates and become less reliable the further ahead you go.
  • In MCQs, read all four options before choosing; there is no negative marking, so attempt every question.

Common mistakes

  • Deleting every record that has a missing value. Fix: Check the missing percentage first. If few records are affected, deletion is fine. If many are, impute, otherwise you lose too much data.
  • Using the mean to fill gaps when the data has extreme values. Fix: Use the median for skewed data. State the reason in your answer.
  • Taking the mean of grouped data as Σx ÷ number of classes. Fix: Always weight by frequency: Σfx ÷ Σf. Check that Σf equals the total observations.
  • Finding the median without sorting the data first. Fix: Sort ascending before picking the middle value. For an even count, average the two middle values.
  • Treating a bar chart and a histogram as the same thing. Fix: Bar chart: categories, gaps between bars. Histogram: continuous class intervals, bars touching, fixed order.
  • Using a pie chart for trend over time or for many categories. Fix: Use a line graph for time. Use a pie only for a few parts of one whole. Use a bar chart when categories are many.
  • Calling a forecast prescriptive because it helps decisions. Fix: Ask whether the output is an estimate or a recommended action. Estimate is predictive. A recommended action is prescriptive.
  • Treating diagnostic analytics as the same as descriptive. Fix: Descriptive reports the result. Diagnostic investigates the cause by drilling down or comparing factors.
  • Writing that a high correlation proves that X causes Y. Fix: Say that r shows association only. Two variables may move together because of a third factor or by chance.
  • Regressing in the wrong direction, using X on Y when Y on X is asked. Fix: Decide which variable is to be predicted. The two lines differ, with b(yx) = r·σy ÷ σx and b(xy) = r·σx ÷ σy.

Exam tips

  • In MCQs, match the technique to the data type: median for skewed numbers, mode for categories, deletion only when few records are affected.
  • Write written answers in the order: identify problem, treatment, reason, validation. Each part can carry a mark.
  • Always show the sorted list and the quartile working in outlier questions. Method marks are lost when only fences are written.
  • Use the words in the syllabus: cleaning, transformation, integration, validation, imputation. Define each in one line.
  • Remember that an outlier is flagged, not automatically wrong. Say that you would verify it.
  • In the MCQ section, many questions are one-step: identify the measure, sort the data, or read the effect of an outlier. Remember that an extreme value moves the mean most, the median little, and the mode usually not at all. There is no negative marking, so always attempt every MCQ.
  • For written answers, draw the working table first. Step marks usually go for the correct table, formula substitution and final interpretation, even if a small arithmetic slip occurs.
  • When two series are compared, calculate CV and name the more consistent one. Do not stop at the standard deviation.