Skip to content

Actuarial Statistics · Jointly distributed random variables

Conditional Distributions and Conditional Expectation Explained

Updated 11 October 2026 · Fact-checked

A conditional distribution describes X once you know Y = y. Find it by dividing the joint pdf or pmf by the marginal of Y. Conditional expectation E[X | Y = y] is the mean of that distribution. The tower law then gives E[X] = E[E[X | Y]], and variance splits into two parts.

Understand Conditional Distributions and Conditional Expectation

A joint distribution tells you how X and Y behave together. Often you learn the value of one and want to update the other. A conditional distribution does this. It is the distribution of X among only those outcomes where Y = y.

For discrete variables, take the joint pmf and divide by the marginal: P(X = x | Y = y) = P(X = x, Y = y) ÷ P(Y = y). For continuous variables the idea is the same: f(x | y) = f(x, y) ÷ f_Y(y), valid where f_Y(y) > 0. For each fixed y, this is a proper density in x. It integrates to 1 over x.

Once you have the conditional distribution, treat it like any other distribution. Conditional expectation E[X | Y = y] is its mean. Conditional variance Var(X | Y = y) is its variance. Both are functions of y.

Now let y be random. Then E[X | Y] is itself a random variable, because it is a function of Y. This gives the tower law: E[X] = E[E[X | Y]]. It lets you find a mean by conditioning on something easier, such as the number of claims.

The same idea gives the variance decomposition: Var(X) = E[Var(X | Y)] + Var(E[X | Y]). The first term is the average spread within each y. The second is the spread of the conditional means between different y. If X and Y are independent, f(x | y) = f_X(x), so conditioning changes nothing.

Key rules to remember

Conditional pmf
P(X = x | Y = y) = P(X = x, Y = y) ÷ P(Y = y)
Needs P(Y = y) > 0.
Conditional pdf
f(x | y) = f(x, y) ÷ f_Y(y)
Needs f_Y(y) > 0. Integrates to 1 over x for each fixed y.
Marginal from joint
f_Y(y) = ∫ f(x, y) dx (or Σ over x for a pmf)
Integrate over the full range of x for that y.
Conditional expectation
E[X | Y = y] = ∫ x f(x | y) dx (or Σ x P(X = x | Y = y))
A function of y. The same rule gives E[g(X) | Y = y] using g(x).
Conditional variance
Var(X | Y = y) = E[X² | Y = y] − (E[X | Y = y])²
Use the conditional second moment, not the unconditional one.
Tower law (law of total expectation)
E[X] = E[E[X | Y]]
The outer expectation is over Y.
Variance decomposition (law of total variance)
Var(X) = E[Var(X | Y)] + Var(E[X | Y])
Within-group variance plus between-group variance.
Independence
f(x | y) = f_X(x) for all y, so E[X | Y] = E[X]
Independence implies this. The converse for the mean alone does not hold.
Random sum
S = X₁ + … + X_N, with N independent of the Xᵢ (iid, mean μ, variance σ²): E[S] = μE[N], Var(S) = σ²E[N] + μ²Var(N)
Follows directly from the tower law and variance decomposition.

How to solve Conditional Distributions and Conditional Expectation questions

Use this method for any question on conditional distributions, conditional moments or the tower law.

  1. 1Write down the joint pdf or pmf and its exact support. Note whether the limits of one variable depend on the other.
  2. 2Find the marginal of the conditioning variable by integrating (or summing) the joint over the other variable, using the correct limits.
  3. 3Divide the joint by that marginal to get the conditional density or pmf. Check it integrates or sums to 1 over the variable of interest.
  4. 4Compute E[X | Y = y] from the conditional distribution. Compute E[X² | Y = y] too if you need the variance.
  5. 5Get Var(X | Y = y) = E[X² | Y = y] − (E[X | Y = y])².
  6. 6If the question asks for an unconditional mean or variance, apply E[X] = E[E[X | Y]] and Var(X) = E[Var(X | Y)] + Var(E[X | Y]), using the distribution of Y.
  7. 7Check the result: the answer must lie within the support, and the unconditional mean must lie between the smallest and largest conditional means.

Quickest way: Condition first, then use the tower law

When to use it: Use when the question gives a hierarchical model, such as 'given N = n, X is ...', or asks only for E[X] or Var(X) rather than the full conditional density.

  1. Identify the variable to condition on. It is usually the one whose value makes the rest simple.
  2. Write E[X | Y] and Var(X | Y) directly from the given conditional distribution. You do not need the joint pdf.
  3. Find E[E[X | Y]] using the mean of Y, or Var(E[X | Y]) using the variance of Y, when E[X | Y] is linear in Y.
  4. Add E[Var(X | Y)] and Var(E[X | Y]) for the total variance.
  5. For random sums use E[S] = μE[N] and Var(S) = σ²E[N] + μ²Var(N) directly.

Common mistakes in Conditional Distributions and Conditional Expectation

  • Dividing by the wrong marginal, or not dividing at all, so the conditional density does not integrate to 1.

    Students confuse f(x, y), f_X(x) and f_Y(y), or use the joint as if it were conditional.

    Fix: To condition on Y = y, divide by f_Y(y). Always check that the result integrates to 1 over x.

  • Using wrong limits when finding the marginal or the conditional mean.

    The support is not a rectangle, for example 0 < x < y < 1, and the dependence between limits is missed.

    Fix: Sketch the region. For f_Y(y), integrate x over its range for that fixed y. Use the same range for E[X | Y = y].

  • Writing Var(X) = E[Var(X | Y)] and stopping.

    Students forget the second term of the decomposition.

    Fix: Always include Var(E[X | Y]). It is zero only when E[X | Y] does not depend on Y.

  • Treating E[X | Y] as a number.

    Students forget it is a function of y, and a random variable once Y is random.

    Fix: Keep y in the answer. Take the outer expectation or variance over Y only in the last step.

  • Using E[X²] where E[X² | Y = y] is needed when computing conditional variance.

    Mixing unconditional and conditional moments.

    Fix: Compute both moments from the same conditional density, then subtract the square of the conditional mean.

  • Assuming E[X | Y] = E[X] implies independence.

    Students over-read the independence rule.

    Fix: Independence implies the conditional mean is constant. A constant conditional mean alone does not prove independence.

Worked examples

Example 1

X and Y have joint pdf f(x, y) = x + y for 0 < x < 1, 0 < y < 1, and 0 otherwise. (a) Find f(x | y). (b) Find E[X | Y = y]. (c) Use the tower law to find E[X].

Show the solution
  1. Marginal of Y: f_Y(y) = ∫₀¹ (x + y) dx = 1/2 + y, for 0 < y < 1.
  2. (a) f(x | y) = (x + y) ÷ (y + 1/2), for 0 < x < 1. Check at y = 0: f(x | 0) = 2x, which integrates to 1.
  3. (b) E[X | Y = y] = ∫₀¹ x(x + y) dx ÷ (y + 1/2) = (1/3 + y/2) ÷ (y + 1/2).
  4. Simplify: (1/3 + y/2) = (2 + 3y)/6 and (y + 1/2) = (2y + 1)/2, so E[X | Y = y] = (2 + 3y) ÷ (3(2y + 1)). Check at y = 0: 2/3, which matches the mean of 2x on (0, 1).
  5. (c) E[X] = ∫₀¹ E[X | Y = y] f_Y(y) dy = ∫₀¹ (1/3 + y/2) dy, since E[X | Y = y] × (y + 1/2) = 1/3 + y/2.
  6. This equals 1/3 + 1/4 = 7/12. Check directly: f_X(x) = x + 1/2, so E[X] = ∫₀¹ x(x + 1/2) dx = 1/3 + 1/4 = 7/12.

Answer: (a) f(x | y) = (x + y) ÷ (y + 1/2) for 0 < x < 1. (b) E[X | Y = y] = (2 + 3y) ÷ (3(2y + 1)). (c) E[X] = 7/12.

Example 2

The number of claims N in a year is Poisson with mean 20. Each claim amount is independent of N, with mean ₹5,000 and standard deviation ₹2,000. Let S be the total annual claims. Find E[S], Var(S) and the standard deviation of S.

Show the solution
  1. Condition on N. Given N = n, E[S | N = n] = 5,000n and Var(S | N = n) = n × 2,000² = 4,000,000n, as claims are independent.
  2. So E[S | N] = 5,000N and Var(S | N) = 4,000,000N.
  3. Tower law: E[S] = E[5,000N] = 5,000 × 20 = ₹1,00,000.
  4. E[Var(S | N)] = 4,000,000 × E[N] = 4,000,000 × 20 = 80,000,000.
  5. Var(E[S | N]) = 5,000² × Var(N) = 25,000,000 × 20 = 500,000,000, since a Poisson variance equals its mean.
  6. Var(S) = 80,000,000 + 500,000,000 = 580,000,000.
  7. Standard deviation = √580,000,000 ≈ 24,083.

Answer: E[S] = ₹1,00,000. Var(S) = 580,000,000 (in ₹²). Standard deviation ≈ ₹24,083.

Exam tips

  • In MCQs, start by asking whether the tower law or variance decomposition applies. It often avoids finding any joint density.
  • For written questions, show the marginal, the conditional density and a check that it integrates to 1. Each step usually earns marks.
  • Always state the support and limits. Many marks are lost on triangular regions such as 0 < x < y.
  • For total variance, write both terms separately and label them. Then a small arithmetic slip costs you less.
  • In computer-based paper work, you may simulate Y first, then X given Y, and compare sample means with E[E[X | Y]]. State the assumptions you use.

Practice questions from Jointly distributed random variables

Conditional Distributions and Conditional Expectation in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Conditional Distributions and Conditional Expectation: frequently asked questions

How do I find E[X | Y = y] from a joint pdf?

First find f_Y(y) by integrating the joint pdf over x. Then form f(x | y) = f(x, y) ÷ f_Y(y). Finally integrate x × f(x | y) over the range of x for that y. The answer is a function of y.

What is the difference between E[X | Y = y] and E[X | Y]?

E[X | Y = y] is a number or function of a fixed y. E[X | Y] is the same function with the random variable Y put in, so it is a random variable. The tower law takes the expectation of this random variable.

When can I use the law of total variance?

Use it whenever X has finite variance and you can describe X given Y. It is most useful for mixtures and random sums, where the conditional mean and variance are simple. Remember to include both terms.

Is the conditional pdf always a valid density?

Yes, for each y where f_Y(y) > 0, f(x | y) is non-negative and integrates to 1 over x. If your result does not integrate to 1, recheck the marginal or the limits.