Risk Modelling and Survival Analysis · Maximum likelihood estimators for transition intensities
Properties and Variance of the MLE of a Transition Intensity
Updated 11 October 2026 · Fact-checked
For a constant transition intensity estimated from v observed transitions and total waiting time T, the MLE is μ̂ = v ÷ T. For large samples μ̂ is approximately normal with mean μ and variance μ² ÷ E[V], estimated by v ÷ T². This gives a confidence interval μ̂ ± z × √v ÷ T.
Understand Properties and Variance of the MLE
In a two-state model, you observe lives for a period and record two things. V is the number of transitions (for example deaths). T is the total time spent in the starting state, often called the central exposed to risk. If the intensity μ is constant, the likelihood is proportional to μ^v × e^(−μT). Maximising it gives μ̂ = V ÷ T.
The MLE is only an estimate. Different samples give different values, so you need its distribution. For large samples, maximum likelihood theory says the MLE is approximately normal. Its mean is the true value μ. Its variance is the Cramér-Rao lower bound, which is 1 ÷ I(μ), where I(μ) = −E[d²l/dμ²] is the Fisher information and l is the log-likelihood.
Here l = V ln μ − μT + constant. So d²l/dμ² = −V ÷ μ². Taking expectations gives I(μ) = E[V] ÷ μ². The variance is therefore μ² ÷ E[V]. This is an asymptotic result. It is a good approximation when the expected number of transitions is large.
In practice you do not know μ or E[V]. You replace E[V] with the observed v and μ with μ̂. This gives the estimated variance v ÷ T², which equals μ̂² ÷ v. Use it to build a confidence interval. Note that T is treated as fixed here. The randomness comes from V. In the model with T fixed, E[V] = μT, so the variance is also μ ÷ T.
The result also shows how precision works. The variance depends on the number of transitions, not just the number of lives. More deaths give a smaller relative error. The standard error relative to μ̂ is 1 ÷ √v.
Key rules to remember
- MLE of constant intensity
- μ̂ = v ÷ T
- v is observed number of transitions; T is total time spent in the starting state (central exposed to risk).
- Log-likelihood
- l(μ) = v ln μ − μT + constant
- Likelihood is proportional to μ^v × e^(−μT).
- Asymptotic distribution
- μ̂ ≈ N(μ, μ² ÷ E[V])
- Valid for large samples, with T treated as fixed.
- Cramér-Rao lower bound
- Var(μ̂) ≈ 1 ÷ I(μ), where I(μ) = −E[d²l/dμ²]
- Here −d²l/dμ² = V ÷ μ², so I(μ) = E[V] ÷ μ².
- Estimated variance
- Var(μ̂) ≈ v ÷ T² = μ̂² ÷ v
- Replace E[V] by v and μ by μ̂.
- Approximate confidence interval
- μ̂ ± z × √v ÷ T
- For 95%, z = 1.96. This is the normal-approximation interval.
How to solve Properties and Variance of the MLE questions
Use this method for any question asking for the MLE, its variance or a confidence interval for a constant transition intensity.
- 1Define the model: state the transition, the assumption that μ is constant, and what v and T are in the data.
- 2Write the likelihood L ∝ μ^v × e^(−μT) and the log-likelihood l = v ln μ − μT.
- 3Differentiate: dl/dμ = v ÷ μ − T. Set it to zero to get μ̂ = v ÷ T. Check d²l/dμ² = −v ÷ μ² < 0, so it is a maximum.
- 4Find the variance: I(μ) = −E[d²l/dμ²] = E[V] ÷ μ², so Var(μ̂) ≈ μ² ÷ E[V].
- 5Substitute estimates to get Var(μ̂) ≈ v ÷ T² and the standard error √v ÷ T.
- 6State the asymptotic distribution: μ̂ is approximately N(μ, v ÷ T²).
- 7Build the interval μ̂ ± z × standard error, using z = 1.96 for 95%. Give the interval and interpret it.
- 8State the assumptions: constant intensity, large sample, T fixed.
Quickest way: Shortcut: μ̂ and standard error from v and T
When to use it: When the question gives you the number of transitions and the total exposure and asks for an estimate, a standard error or a confidence interval.
- Compute μ̂ = v ÷ T.
- Compute the standard error as μ̂ ÷ √v, which equals √v ÷ T.
- Multiply by the z value (1.96 for 95%).
- Write the interval μ̂ ± z × se.
- If asked for the lower bound only, subtract and check it is positive.
Common mistakes in Properties and Variance of the MLE
Using the number of lives instead of the total time T as the denominator.
Students confuse the crude rate with a rate per life.
Fix: The denominator is total time in the state, summed across all lives, including part-years.
Writing the variance as μ ÷ v or v ÷ T instead of v ÷ T².
Students forget to square T or mix up the formulas.
Fix: Derive it: μ̂ = v ÷ T, so the variance is μ² ÷ E[V], which becomes v ÷ T² after substitution. Check units.
Treating the variance result as exact for small samples.
The formula is memorised without its condition.
Fix: Say it is asymptotic. It is a good approximation only when the expected number of transitions is large.
Forgetting to substitute μ̂ when μ is unknown.
Students leave the answer in terms of μ.
Fix: A numerical interval needs estimates, so use v for E[V] and μ̂ for μ.
Skipping the second derivative check.
Students assume the stationary point is a maximum.
Fix: Show d²l/dμ² = −v ÷ μ² < 0 whenever you derive the MLE.
Worked examples
Example 1
In a study of a two-state model with constant force of mortality μ, 40 deaths are observed over a total of 2,000 years of exposure. Estimate μ and its standard error, and give an approximate 95% confidence interval.
Show the solution
- v = 40 and T = 2,000.
- μ̂ = 40 ÷ 2,000 = 0.02.
- Estimated variance = v ÷ T² = 40 ÷ 4,000,000 = 0.00001.
- Standard error = √0.00001 = 0.003162.
- Margin = 1.96 × 0.003162 = 0.006198.
- Interval = 0.02 ± 0.006198.
Answer: μ̂ = 0.02, standard error ≈ 0.00316, 95% confidence interval ≈ (0.0138, 0.0262).
Example 2
For the model with likelihood L ∝ μ^v × e^(−μT), derive the MLE of μ and show that its asymptotic variance is μ² ÷ E[V].
Show the solution
- l = v ln μ − μT + constant.
- dl/dμ = v ÷ μ − T. Setting it to zero gives μ̂ = v ÷ T.
- d²l/dμ² = −v ÷ μ², which is negative, so this is a maximum.
- Replace v by the random variable V: −d²l/dμ² = V ÷ μ².
- I(μ) = E[V] ÷ μ².
- By the Cramér-Rao result, Var(μ̂) ≈ 1 ÷ I(μ) = μ² ÷ E[V].
Answer: μ̂ = v ÷ T, and asymptotically μ̂ ≈ N(μ, μ² ÷ E[V]), estimated by v ÷ T².
Exam tips
- Always show the derivation of μ̂ and the second derivative check if asked to derive. Marks are given for method.
- State that the normal result is asymptotic and that T is treated as fixed.
- Keep units clear. μ is per year, and T is in years.
- In computer-based questions, compute v, T, μ̂ and the standard error in code, then state the interval in words.
- If you are given E[V] ≈ μT, you can also write the variance as μ ÷ T. Show the link to v ÷ T².
Practice questions from Maximum likelihood estimators for transition intensities
- In a two-state alive-dead model with constant force of mortality mu, a study observes n lives, with total observed time exposed to risk v an…
- For the constant-force model with D deaths and total waiting time V, the Cramér–Rao lower bound for an unbiased estimator of μ is approximat…
- In the two-state alive-dead model with a constant force of mortality μ, a study observes n lives. Let D be the number of deaths and V the to…
- In a two-state model (Alive to Dead) with constant force of mortality mu, a study observes n lives. The total observed time exposed to risk …
- With a constant force of mortality, 40 deaths are observed over a total waiting time of 800 life-years. Using the asymptotic distribution of…
Properties and Variance of the MLE: frequently asked questions
Why is the MLE of μ equal to v ÷ T?
Maximising the log-likelihood v ln μ − μT gives v ÷ μ = T. So μ̂ = v ÷ T, the number of transitions divided by total time in the state.
What is the Cramér-Rao lower bound here?
It is the smallest variance an unbiased estimator can have, equal to 1 ÷ I(μ). For this model it is μ² ÷ E[V]. The MLE reaches this bound asymptotically.
When is the normal confidence interval reliable?
It is reliable when the number of transitions is large. With few transitions the distribution of μ̂ is skewed, and the lower limit can even turn negative.
Do I use μ or μ̂ in the variance?
The true μ is unknown, so use μ̂. The estimated variance is μ̂² ÷ v, which equals v ÷ T².