Skip to content

Risk Modelling and Survival Analysis · Maximum likelihood estimators for transition intensities

Properties and Variance of the MLE of a Transition Intensity

Updated 11 October 2026 · Fact-checked

For a constant transition intensity estimated from v observed transitions and total waiting time T, the MLE is μ̂ = v ÷ T. For large samples μ̂ is approximately normal with mean μ and variance μ² ÷ E[V], estimated by v ÷ T². This gives a confidence interval μ̂ ± z × √v ÷ T.

Understand Properties and Variance of the MLE

In a two-state model, you observe lives for a period and record two things. V is the number of transitions (for example deaths). T is the total time spent in the starting state, often called the central exposed to risk. If the intensity μ is constant, the likelihood is proportional to μ^v × e^(−μT). Maximising it gives μ̂ = V ÷ T.

The MLE is only an estimate. Different samples give different values, so you need its distribution. For large samples, maximum likelihood theory says the MLE is approximately normal. Its mean is the true value μ. Its variance is the Cramér-Rao lower bound, which is 1 ÷ I(μ), where I(μ) = −E[d²l/dμ²] is the Fisher information and l is the log-likelihood.

Here l = V ln μ − μT + constant. So d²l/dμ² = −V ÷ μ². Taking expectations gives I(μ) = E[V] ÷ μ². The variance is therefore μ² ÷ E[V]. This is an asymptotic result. It is a good approximation when the expected number of transitions is large.

In practice you do not know μ or E[V]. You replace E[V] with the observed v and μ with μ̂. This gives the estimated variance v ÷ T², which equals μ̂² ÷ v. Use it to build a confidence interval. Note that T is treated as fixed here. The randomness comes from V. In the model with T fixed, E[V] = μT, so the variance is also μ ÷ T.

The result also shows how precision works. The variance depends on the number of transitions, not just the number of lives. More deaths give a smaller relative error. The standard error relative to μ̂ is 1 ÷ √v.

Key rules to remember

MLE of constant intensity
μ̂ = v ÷ T
v is observed number of transitions; T is total time spent in the starting state (central exposed to risk).
Log-likelihood
l(μ) = v ln μ − μT + constant
Likelihood is proportional to μ^v × e^(−μT).
Asymptotic distribution
μ̂ ≈ N(μ, μ² ÷ E[V])
Valid for large samples, with T treated as fixed.
Cramér-Rao lower bound
Var(μ̂) ≈ 1 ÷ I(μ), where I(μ) = −E[d²l/dμ²]
Here −d²l/dμ² = V ÷ μ², so I(μ) = E[V] ÷ μ².
Estimated variance
Var(μ̂) ≈ v ÷ T² = μ̂² ÷ v
Replace E[V] by v and μ by μ̂.
Approximate confidence interval
μ̂ ± z × √v ÷ T
For 95%, z = 1.96. This is the normal-approximation interval.

How to solve Properties and Variance of the MLE questions

Use this method for any question asking for the MLE, its variance or a confidence interval for a constant transition intensity.

  1. 1Define the model: state the transition, the assumption that μ is constant, and what v and T are in the data.
  2. 2Write the likelihood L ∝ μ^v × e^(−μT) and the log-likelihood l = v ln μ − μT.
  3. 3Differentiate: dl/dμ = v ÷ μ − T. Set it to zero to get μ̂ = v ÷ T. Check d²l/dμ² = −v ÷ μ² < 0, so it is a maximum.
  4. 4Find the variance: I(μ) = −E[d²l/dμ²] = E[V] ÷ μ², so Var(μ̂) ≈ μ² ÷ E[V].
  5. 5Substitute estimates to get Var(μ̂) ≈ v ÷ T² and the standard error √v ÷ T.
  6. 6State the asymptotic distribution: μ̂ is approximately N(μ, v ÷ T²).
  7. 7Build the interval μ̂ ± z × standard error, using z = 1.96 for 95%. Give the interval and interpret it.
  8. 8State the assumptions: constant intensity, large sample, T fixed.

Quickest way: Shortcut: μ̂ and standard error from v and T

When to use it: When the question gives you the number of transitions and the total exposure and asks for an estimate, a standard error or a confidence interval.

  1. Compute μ̂ = v ÷ T.
  2. Compute the standard error as μ̂ ÷ √v, which equals √v ÷ T.
  3. Multiply by the z value (1.96 for 95%).
  4. Write the interval μ̂ ± z × se.
  5. If asked for the lower bound only, subtract and check it is positive.

Common mistakes in Properties and Variance of the MLE

  • Using the number of lives instead of the total time T as the denominator.

    Students confuse the crude rate with a rate per life.

    Fix: The denominator is total time in the state, summed across all lives, including part-years.

  • Writing the variance as μ ÷ v or v ÷ T instead of v ÷ T².

    Students forget to square T or mix up the formulas.

    Fix: Derive it: μ̂ = v ÷ T, so the variance is μ² ÷ E[V], which becomes v ÷ T² after substitution. Check units.

  • Treating the variance result as exact for small samples.

    The formula is memorised without its condition.

    Fix: Say it is asymptotic. It is a good approximation only when the expected number of transitions is large.

  • Forgetting to substitute μ̂ when μ is unknown.

    Students leave the answer in terms of μ.

    Fix: A numerical interval needs estimates, so use v for E[V] and μ̂ for μ.

  • Skipping the second derivative check.

    Students assume the stationary point is a maximum.

    Fix: Show d²l/dμ² = −v ÷ μ² < 0 whenever you derive the MLE.

Worked examples

Example 1

In a study of a two-state model with constant force of mortality μ, 40 deaths are observed over a total of 2,000 years of exposure. Estimate μ and its standard error, and give an approximate 95% confidence interval.

Show the solution
  1. v = 40 and T = 2,000.
  2. μ̂ = 40 ÷ 2,000 = 0.02.
  3. Estimated variance = v ÷ T² = 40 ÷ 4,000,000 = 0.00001.
  4. Standard error = √0.00001 = 0.003162.
  5. Margin = 1.96 × 0.003162 = 0.006198.
  6. Interval = 0.02 ± 0.006198.

Answer: μ̂ = 0.02, standard error ≈ 0.00316, 95% confidence interval ≈ (0.0138, 0.0262).

Example 2

For the model with likelihood L ∝ μ^v × e^(−μT), derive the MLE of μ and show that its asymptotic variance is μ² ÷ E[V].

Show the solution
  1. l = v ln μ − μT + constant.
  2. dl/dμ = v ÷ μ − T. Setting it to zero gives μ̂ = v ÷ T.
  3. d²l/dμ² = −v ÷ μ², which is negative, so this is a maximum.
  4. Replace v by the random variable V: −d²l/dμ² = V ÷ μ².
  5. I(μ) = E[V] ÷ μ².
  6. By the Cramér-Rao result, Var(μ̂) ≈ 1 ÷ I(μ) = μ² ÷ E[V].

Answer: μ̂ = v ÷ T, and asymptotically μ̂ ≈ N(μ, μ² ÷ E[V]), estimated by v ÷ T².

Exam tips

  • Always show the derivation of μ̂ and the second derivative check if asked to derive. Marks are given for method.
  • State that the normal result is asymptotic and that T is treated as fixed.
  • Keep units clear. μ is per year, and T is in years.
  • In computer-based questions, compute v, T, μ̂ and the standard error in code, then state the interval in words.
  • If you are given E[V] ≈ μT, you can also write the variance as μ ÷ T. Show the link to v ÷ T².

Practice questions from Maximum likelihood estimators for transition intensities

Properties and Variance of the MLE: frequently asked questions

Why is the MLE of μ equal to v ÷ T?

Maximising the log-likelihood v ln μ − μT gives v ÷ μ = T. So μ̂ = v ÷ T, the number of transitions divided by total time in the state.

What is the Cramér-Rao lower bound here?

It is the smallest variance an unbiased estimator can have, equal to 1 ÷ I(μ). For this model it is μ² ÷ E[V]. The MLE reaches this bound asymptotically.

When is the normal confidence interval reliable?

It is reliable when the number of transitions is large. With few transitions the distribution of μ̂ is skewed, and the lower limit can even turn negative.

Do I use μ or μ̂ in the variance?

The true μ is unknown, so use μ̂. The estimated variance is μ̂² ÷ v, which equals v ÷ T².