CA Foundation · Quantitative Aptitude
Theoretical Distributions: formula sheet
Key formulas
- Conditions for a discrete distribution
- p(x) ≥ 0 for all x, and Σ p(x) = 1
- Use this to find an unknown constant in a given table.
- Conditions for a continuous distribution
- f(x) ≥ 0, and total area under f(x) = 1
- For a continuous variable, P(X = a) = 0, so P(a ≤ X ≤ b) is the same whether endpoints are included or not.
- Expectation (discrete)
- E(X) = Σ x·p(x)
- Multiply each value by its probability and add.
- Expectation of X²
- E(X²) = Σ x²·p(x)
- Square the value x, not the probability.
- Variance
- Var(X) = E(X²) − [E(X)]²
- Always non-negative. If you get a negative value, recheck your arithmetic.
- Standard deviation
- SD(X) = √Var(X)
- Same units as X.
- Linear change
- E(aX + b) = a·E(X) + b; Var(aX + b) = a²·Var(X)
- Adding a constant b does not change variance.
- Cumulative probability
- P(X ≤ x) = Σ p(t) for all t ≤ x
- For discrete X, P(X > x) = 1 − P(X ≤ x).
- Binomial probability
- P(X = r) = nCr × p^r × q^(n−r), r = 0, 1, 2, ..., n
- n = number of trials, r = number of successes, p = probability of success, q = 1 − p.
- Combination
- nCr = n! ÷ [r! × (n − r)!]
- Counts the ways to place r successes among n trials. nCr = nC(n−r).
- Probability of failure
- q = 1 − p
- p + q = 1 in every trial.
- Total probability
- Σ P(X = r) for r = 0 to n equals 1
- Use it to check answers or to find 'at least' by subtraction.
- At least one success
- P(X ≥ 1) = 1 − q^n
- Faster than adding P(1) + P(2) + ... + P(n).
- Mean and variance
- Mean = np; Variance = npq
- Standard deviation = √(npq). Since q < 1 for p > 0, variance npq < mean np. They are equal only in the trivial case p = 0.
- Recurrence relation
- P(r + 1) = [(n − r) ÷ (r + 1)] × (p ÷ q) × P(r)
- Useful when you need several consecutive terms starting from P(0) = q^n.
- Mean
- μ = np
- Average number of successes in n trials.
- Variance
- σ² = npq, where q = 1 − p
- For 0 < p < 1, q < 1, so the variance is less than the mean.
- Standard deviation
- σ = √(npq)
- Take the square root of the variance.
- Probability of r successes
- P(X = r) = nCr × p^r × q^(n − r), r = 0, 1, ..., n
- Used for fitting and for mode checks.
- Mode
- Compute (n + 1)p. If not an integer, mode = integer part. If an integer, two modes: (n + 1)p and (n + 1)p − 1
- Always check whether (n + 1)p is a whole number.
- Additive property
- X ~ B(n1, p), Y ~ B(n2, p), independent ⇒ X + Y ~ B(n1 + n2, p)
- Needs the same p and independence.
- Finding n and p
- q = variance ÷ mean; p = 1 − q; n = mean ÷ p
- Valid only when variance < mean, which holds for 0 < p < 1.
- Expected frequency in fitting
- Expected frequency = N × P(X = r)
- N is the total frequency (number of sets of trials).
- Poisson probability
- P(X = r) = e^(−m) × m^r ÷ r!, for r = 0, 1, 2, ...
- m is the average number of occurrences in the interval. Exam questions usually give the value of e^(−m).
- Mean and variance
- Mean = m; Variance = m; Standard deviation = √m
- Mean equals variance. This is the property most used to identify m.
- Recurrence relation
- P(r + 1) = m ÷ (r + 1) × P(r)
- Gives each probability from the previous one without computing factorials.
- Poisson as limit of binomial
- m = np
- Use when n is large and p is small. The binomial probabilities are then close to the Poisson ones.
- Probability of at least one
- P(X ≥ 1) = 1 − e^(−m)
- Use the complement for 'at least' questions.
- Sum of independent Poisson variables
- If X ~ Poisson(m₁) and Y ~ Poisson(m₂) are independent, X + Y ~ Poisson(m₁ + m₂)
- Useful when the interval is extended or two sources are combined.
- Mode
- Mode = integer part of m. If m is a whole number, there are two modes: m − 1 and m
- Check this when the question asks for the most likely value.
- Normal variable notation
- X ~ N(μ, σ²)
- μ is the mean, σ is the standard deviation, σ² is the variance.
- Probability density function
- f(x) = [1 ÷ (σ√(2π))] × e^(−(x − μ)² ÷ (2σ²)), for −∞ < x < +∞
- Rarely used for calculation. You may need to recognise it. The π here is 3.14159..., and e is about 2.718.
- Central equality
- Mean = Median = Mode = μ
- True because the curve is symmetric and has a single peak.
- Skewness and kurtosis
- Skewness (β₁ and γ₁) = 0; β₂ = 3
- β₂ = 3 means the curve is mesokurtic.
- Empirical rule
- P(μ − σ < X < μ + σ) ≈ 68.27%; P(μ − 2σ < X < μ + 2σ) ≈ 95.45%; P(μ − 3σ < X < μ + 3σ) ≈ 99.73%
- Exact values are 68.27%, 95.45% and 99.73%. The rounded values 68%, 95% and 99.7% are usually used in the options.
- Area on each side
- P(X < μ) = P(X > μ) = 0.5
- Total area under the curve = 1.
- Quartiles
- Q₁ = μ − 0.6745σ; Q₃ = μ + 0.6745σ
- Quartile deviation = 0.6745σ, which is about (2/3)σ. Mean deviation is about 0.7979σ, which is about (4/5)σ.
- Quartile deviation, mean deviation and SD (approximate)
- QD : MD : SD ≈ 10 : 12 : 15, from QD ≈ (2/3)σ and MD ≈ (4/5)σ
- This ratio is only an approximation. The exact ratio 0.6745 : 0.7979 : 1 is roughly 10.1 : 12 : 15. With σ = 15, QD ≈ 10 and MD ≈ 12.
- Points of inflection
- x = μ − σ and x = μ + σ
- The curve changes from bending downward to bending upward at these points.
- Standardisation
- Z = (X − μ) ÷ σ
- Z has mean 0 and SD 1. Table use is covered in the Z-table topic.
- Standard normal variate
- Z = (X − μ) ÷ σ
- μ is the mean and σ is the standard deviation of X. Z has mean 0 and SD 1.
- Symmetry
- P(Z < 0) = P(Z > 0) = 0.5 and P(Z > a) = P(Z < −a)
- Total area under the curve is 1. Use this to handle negative Z.
- Area from the mean (0-to-Z table)
- P(0 < Z < a) = P(−a < Z < 0) = table value at a
- The table value for a negative Z is read at |Z|.
- Right tail
- P(Z > a) = 0.5 − P(0 < Z < a), for a ≥ 0
- For a < 0, P(Z > a) = 0.5 + P(0 < Z < |a|).
- Interval on opposite sides of the mean
- P(−a < Z < b) = P(0 < Z < a) + P(0 < Z < b), for a, b > 0
- Add the two areas.
- Interval on the same side of the mean
- P(a < Z < b) = P(0 < Z < b) − P(0 < Z < a), for 0 ≤ a < b
- Subtract the smaller area from the larger.
- Number of items
- Expected number = N × probability
- N is the total number of items in the group.
- Standard reference areas
- P(−1 < Z < 1) ≈ 0.6826, P(−2 < Z < 2) ≈ 0.9545, P(−3 < Z < 3) ≈ 0.9973
- Also P(−1.96 < Z < 1.96) ≈ 0.95. Use these to check or skip table lookups.
- Binomial probability
- P(X = r) = nCr × p^r × q^(n−r), where q = 1 − p, r = 0, 1, ..., n
- Use when n is fixed and each trial is independent with the same p.
- Binomial mean and variance
- Mean = np; Variance = npq; SD = √(npq)
- Variance is always less than the mean because q < 1.
- Poisson probability
- P(X = r) = e^(−m) × m^r ÷ r!, where r = 0, 1, 2, ...
- m is the average number of occurrences. The values of r have no upper limit.
- Poisson mean and variance
- Mean = m; Variance = m; SD = √m
- Mean equals variance. This is the key clue for Poisson.
- Poisson as limit of binomial
- m = np (use when n is large and p is small)
- A common rule of thumb is np less than about 5. It is an approximation, not an exact rule.
- Normal distribution parameters
- X ~ N(μ, σ²); mean = median = mode = μ; skewness = 0
- The curve is a symmetric bell. The total area under it is 1.
- Standard normal variate
- Z = (X − μ) ÷ σ
- Z has mean 0 and SD 1. Use it with the Z-table.
- Area rule for normal curve
- μ ± 1σ ≈ 68.27%; μ ± 2σ ≈ 95.45%; μ ± 3σ ≈ 99.73%
- These are the standard approximate areas under the curve.
- Normal approximation to binomial
- X ≈ N(np, npq), so Z = (X − np) ÷ √(npq)
- Use when n is large and p is not near 0 or 1. A common check is that both np and nq are at least 5 or so. Apply the ±0.5 continuity correction if asked.
Quick revision
- A discrete variable takes countable values; a continuous variable takes any value in a range.
- The probabilities of all values of a random variable add up to 1.
- Binomial needs fixed n, two outcomes, constant p and independent trials.
- Binomial: P(X = r) = nCr × p^r × q^(n − r), with q = 1 − p.
- Binomial mean = np and variance = npq; variance is always less than the mean.
- Poisson: P(X = r) = e^(−m) × m^r ÷ r!, used for rare events.
- Poisson mean and variance are both equal to m.
- Normal curve is bell-shaped and symmetric about the mean; mean, median and mode are equal.
- Total area under the Normal curve is 1, and half lies on each side of the mean.
- Standard normal variate: Z = (X − μ) ÷ σ, with mean 0 and variance 1.
- For a Normal distribution, about 68%, 95% and 99.7% of values lie within 1, 2 and 3 standard deviations of the mean.
- Binomial variance is less than its mean, Poisson has them equal, and Normal has no link between them.
Common mistakes
- Forgetting to check that Σ p(x) = 1 before using a table. Fix: Add the probabilities first. If one is unknown, solve for it. This often finds the answer on its own.
- Computing E(X²) as [E(X)]². Fix: Square each x first, multiply by p(x), then add. Only then subtract [E(X)]² to get variance.
- Forgetting the nCr term and writing only p^r × q^(n−r). Fix: Always ask 'in how many positions can the successes occur?' and multiply by nCr.
- Using p for the wrong event, such as taking p as the defective rate when the question counts good items. Fix: Success is whatever the question counts. Write 'success = ...' in your rough work before choosing p.
- Using variance = np instead of npq. Fix: Remember that variance has the extra q. For 0 < p < 1, q < 1, so variance is smaller than mean.
- Treating the given standard deviation as the variance. Fix: Square the SD first. Then use npq = SD².
- Forgetting to change m when the interval changes. Fix: Scale m with the interval. A rate of 2 per page becomes m = 6 for 3 pages. Write 'm = ...' before using the formula.
- Using m = n × q or m = p instead of m = np. Fix: For a binomial approximation, always take m = np. Here p is the probability of the rare event.
- Saying the mean, median and mode of a normal distribution are different. Fix: For a normal curve, symmetry and a single peak make all three equal to μ.
- Treating 68%, 95% and 99.7% as areas from the mean to μ + σ, μ + 2σ, μ + 3σ. Fix: The rule is for μ ± kσ. The one-sided area from μ to μ + σ is about 34%.
Exam tips
- Questions often give a table with one unknown. Solve Σ p(x) = 1 first, because later parts depend on it.
- Learn Var(aX + b) = a²Var(X) cold. It gives a quick answer without any table work.
- Check whether the question asks for variance or standard deviation before you pick an option.
- Expect conceptual MCQs: which variable is discrete, why P(X = a) = 0 for continuous X, or what a theoretical distribution is. Revise the definitions.
- Wrong answers cost 0.25 marks. Skip a long table question if you are unsure and come back at the end.
- Questions are usually direct: given n and p, find P(exactly r) or an 'at least' probability. Practise these until the setup takes seconds.
- Check for the phrase 'independent' or 'with replacement'. It confirms the binomial model is meant.
- Expect the same data to be used in the mean-and-variance questions, so remember np and npq with this topic.