Skip to content

FRM Part I · FRM Exam Part I

Multivariate Random Variables: formula sheet

Full chapter guide

Key formulas

Joint PMF (discrete)
f(x, y) = P(X = x, Y = y); f(x, y) ≥ 0; Σx Σy f(x, y) = 1
All cells of the table must add to 1.
Marginal PMF
fX(x) = Σy f(x, y); fY(y) = Σx f(x, y)
Sum across the other variable.
Joint PDF (continuous)
f(x, y) ≥ 0; ∫∫ f(x, y) dx dy = 1 over the whole range
Probabilities come from integrating over a region, not from the value of f.
Marginal PDF
fX(x) = ∫ f(x, y) dy; fY(y) = ∫ f(x, y) dx
Integrate over the full range of the other variable. Watch the limits.
Independence
f(x, y) = fX(x) × fY(y) for all x, y
It must hold for every pair, not just some.
Conditional distribution
f(x | y) = f(x, y) ÷ fY(y), for fY(y) > 0
Links the joint to the marginal.
Conditional probability (mass or density)
f(x | y) = f(x, y) ÷ f(y)
Valid only when f(y) > 0. Each conditional distribution sums (or integrates) to 1.
Marginal from joint
f(x) = Σ f(x, y) over y (discrete); f(x) = ∫ f(x, y) dy (continuous)
Sum across the row or column to get the marginal.
Independence
f(x, y) = f(x) × f(y) for all x, y
Must hold for every pair. One failure means dependent.
Independence, conditional form
f(x | y) = f(x) for all y with f(y) > 0
Equivalent to the product rule.
Conditional expectation (discrete)
E(X | Y = y) = Σ x × f(x | y)
Mean of the conditional distribution.
Law of iterated expectations
E(X) = E[E(X | Y)]
Weight each conditional mean by P(Y = y) and add.
Conditional variance
Var(X | Y = y) = E(X² | Y = y) − [E(X | Y = y)]²
Use conditional probabilities throughout.
Independence and moments
If independent: E(XY) = E(X)E(Y) and Cov(X, Y) = 0
The converse is false in general.
Covariance (definition)
Cov(X,Y) = E[(X − μX)(Y − μY)]
Average product of deviations from the means.
Covariance (shortcut)
Cov(X,Y) = E[XY] − E[X]E[Y]
Fastest form when you have a joint probability table.
Correlation
ρXY = Cov(X,Y) ÷ (σX × σY)
Always between −1 and +1. Needs both standard deviations to be non-zero.
Sample covariance
s_XY = Σ(Xi − X̄)(Yi − Ȳ) ÷ (n − 1)
Divide by n − 1 for the unbiased sample estimate. Divide by n for a population of n equally likely points.
Covariance with itself
Cov(X,X) = Var(X)
Variance is a special case of covariance.
Scaling and shifting
Cov(aX + b, cY + d) = ac × Cov(X,Y)
Constants b and d drop out. Correlation is unchanged if a and c have the same sign and flips sign if they differ.
Variance of a sum
Var(X + Y) = Var(X) + Var(Y) + 2Cov(X,Y)
For a difference, the covariance term is subtracted: Var(X − Y) = Var(X) + Var(Y) − 2Cov(X,Y).
Independence
X, Y independent ⇒ Cov(X,Y) = 0
The reverse is not true in general.
Mean of a linear combination
E(aX + bY + c) = aE(X) + bE(Y) + c
Always true. No independence needed.
Variance of a weighted sum
Var(aX + bY) = a²σX² + b²σY² + 2ab·Cov(X, Y)
A constant added to the sum does not change the variance.
Covariance and correlation link
Cov(X, Y) = ρ·σX·σY
Correlation is covariance scaled to lie between −1 and +1.
Two-asset portfolio variance
σp² = w1²σ1² + w2²σ2² + 2w1w2ρσ1σ2
Weights are the portfolio weights. Volatility is σp = √σp².
Variance of a difference
Var(X − Y) = σX² + σY² − 2Cov(X, Y)
The sign of the covariance term flips.
Independent or uncorrelated case
Var(aX + bY) = a²σX² + b²σY²
Valid when Cov(X, Y) = 0. Independence implies zero covariance, but not the reverse.
Matrix form
σp² = wᵀΣw
Σ is the covariance matrix. Diagonal entries are variances.
Skewness
S = E[(X − μ)³] ÷ σ³
Zero for symmetric distributions. Negative means a longer left tail.
Kurtosis
K = E[(X − μ)⁴] ÷ σ⁴
Normal = 3. Always positive.
Excess kurtosis
Excess K = K − 3
Positive means fatter tails than the normal.
Coskewness
S(X,X,Y) = E[(X − μX)²(Y − μY)] ÷ (σX² σY)
Swap the squared variable to get S(X,Y,Y). Powers sum to 3.
Cokurtosis
K(X,X,Y,Y) = E[(X − μX)²(Y − μY)²] ÷ (σX² σY²)
Powers sum to 4. Other versions use powers 3,1 or 1,3.
Sample skewness (equal weights)
Ŝ = [Σ(xi − x̄)³ ÷ n] ÷ σ̂³
Use the same σ̂ convention as the question states.
Covariance from correlation
Cov(X, Y) = ρ × σx × σy
Off-diagonal entry of the covariance matrix. ρ is between −1 and 1.
Conditional mean of Y given X = x
E(Y | X = x) = μy + ρ × (σy ÷ σx) × (x − μx)
Linear in x. The slope ρσy/σx is the regression slope of Y on X.
Conditional variance of Y given X = x
Var(Y | X = x) = σy² × (1 − ρ²)
Does not depend on x. Take the square root for the standard deviation.
Linear combination of two jointly normal variables
aX + bY ~ Normal(aμx + bμy, a²σx² + b²σy² + 2abρσxσy)
Used for portfolio return and risk. Watch the sign of ρ.
Standardisation
Z = (W − mean) ÷ standard deviation
Convert to a standard normal to find probabilities.
Multivariate normal parameters
X ~ N(μ, Σ), where μ is the mean vector and Σ is the covariance matrix
Σ is symmetric. Any linear combination wᵀX is normal with mean wᵀμ and variance wᵀΣw.
Independence rule
For jointly normal X and Y: ρ = 0 ⇔ X and Y are independent
Holds only under joint normality, not for marginally normal variables in general.

Quick revision

  • Marginal probability: sum the joint probabilities over the other variable.
  • Conditional probability: P(X | Y) = P(X, Y) ÷ P(Y), for P(Y) > 0.
  • Independent variables satisfy P(X, Y) = P(X)P(Y) for all values.
  • Cov(X, Y) = E(XY) − E(X)E(Y).
  • Correlation = Cov(X, Y) ÷ (σX σY), always between −1 and +1.
  • Independent variables (with finite variances) have zero covariance, but zero covariance does not imply independence.
  • E(aX + bY) = aE(X) + bE(Y), whether or not X and Y are independent.
  • Var(aX + bY) = a²Var(X) + b²Var(Y) + 2ab·Cov(X, Y).
  • Var(X + Y) = Var(X) + Var(Y) if the covariance is zero.
  • Covariance changes with units; correlation does not.
  • Skewness is the third standardised moment and kurtosis the fourth. A normal has skewness 0 and kurtosis 3.
  • For jointly normal variables, zero correlation implies independence.

Common mistakes

  • Treating a joint density value f(x, y) as a probability. Fix: For continuous variables, always integrate over a region to get probability.
  • Summing the wrong direction when finding a marginal. Fix: To get the marginal of X, you sum over Y, so each X value gets one total.
  • Dividing by the wrong marginal, or not dividing at all. Fix: Divide by the marginal of the variable you are conditioning on, the one after the bar.
  • Concluding independence from zero correlation. Fix: Zero correlation only rules out linear dependence. Test the product rule. Only for a bivariate normal does zero correlation imply independence.
  • Treating zero correlation as independence. Fix: Remember that correlation captures only linear dependence. Y = X² with symmetric X has zero correlation but full dependence.
  • Forgetting to subtract E[X]E[Y] when computing covariance. Fix: Always write Cov = E[XY] − E[X]E[Y] first and fill in all three terms.
  • Adding standard deviations to get portfolio risk. Fix: Add variances plus the covariance term. Adding σ's works only when ρ = 1 and weights are positive.
  • Leaving out the factor 2 on the covariance term. Fix: Always write 2ab·Cov(X, Y) or 2w1w2ρσ1σ2.
  • Saying kurtosis of the normal is 0. Fix: Normal kurtosis is 3. Excess kurtosis is 0. Read which one the question asks for.
  • Dividing by variance instead of σ³ or σ⁴. Fix: Divide by σ raised to the same power as the moment order.

Exam tips

  • Expect a small joint table (2×2 or 3×3) and a question on a marginal, a conditional probability or independence.
  • Always write the margins first. Most options are built from common slips such as a joint cell used as a marginal.
  • For density questions, check the limits of integration before computing. A triangular support is a common trap.
  • If a constant k is unknown, use total probability = 1 before anything else.
  • Independence is tested against the product of marginals; zero covariance alone does not prove it.
  • Questions often hand you a joint table. Compute marginals first, then the conditional you need.
  • For independence, find one cell that fails. It is faster than checking them all.
  • Expect a distractor that says uncorrelated variables are independent. It is true only for the bivariate normal.