Skip to content

FRM Exam Part I · Stationary Time Series

Model Selection, Estimation and Forecasting in Time Series

Updated 11 October 2026 · Fact-checked

You pick a time series model by comparing information criteria (AIC, BIC; lower is better), then check that residuals look like white noise using the Box-Pierce or Ljung-Box Q test. You forecast by applying the fitted model recursively, so AR(1) forecasts decay toward the mean as the horizon grows.

Understand Model Selection, Estimation and Forecasting

A time series model is only useful if it captures the dependence in the data. After you fit an AR, MA or ARMA model, you ask two questions. Is this the best model among the candidates? And have the residuals been left with any pattern the model missed?

Model selection balances fit against complexity. Adding lags always lowers the sum of squared residuals, but it risks overfitting. AIC and BIC (also called SIC) reward fit and penalise the number of parameters. You compute each for every candidate model and choose the one with the lowest value. BIC penalises extra parameters more heavily for realistic sample sizes, so it tends to pick smaller models. AIC is the better guide when forecasting accuracy matters most; BIC is consistent, meaning it picks the true model in large samples if it is among the candidates.

Residual diagnostics test whether the model is adequate. If the model is right, residuals are white noise: zero mean, constant variance, no autocorrelation. The Box-Pierce and Ljung-Box statistics test the joint null that the first m residual autocorrelations are all zero. A large Q (above the chi-squared critical value) means leftover autocorrelation, so the model is misspecified.

Forecasting rests on Wold's representation theorem: any covariance stationary process can be written as a deterministic component plus an infinite moving average of white noise shocks. Lag operators let you write models compactly, for example AR(1) as (1 − φL)Yt = c + εt. For forecasting, you set future shocks to their expected value of zero and substitute earlier forecasts for unknown future values. Forecasts of a stationary model revert to the unconditional mean as the horizon grows.

Key formulas to remember

Box-Pierce Q statistic
Q_BP = T × Σ(k=1 to m) ρ̂k²
T = sample size, ρ̂k = sample autocorrelation at lag k. Under the null of no autocorrelation, Q is approximately chi-squared with m degrees of freedom (reduced by the number of estimated ARMA parameters when applied to residuals).
Ljung-Box Q statistic
Q_LB = T(T + 2) × Σ(k=1 to m) [ρ̂k² ÷ (T − k)]
Better small-sample behaviour than Box-Pierce. Same chi-squared reference distribution. Reject the null of white noise if Q exceeds the critical value.
AIC
AIC = ln(σ̂²) + 2k ÷ T
σ̂² = SSR ÷ T, k = number of estimated parameters. Choose the lowest value. Some texts write AIC = −2 ln L + 2k; the ranking is the same.
BIC (SIC)
BIC = ln(σ̂²) + k ln(T) ÷ T
Penalty is larger than AIC's when T ≥ 8 (ln T > 2). Prefers more parsimonious models.
Wold's representation
Yt = μ + Σ(i=0 to ∞) ψi εt−i, with ψ0 = 1 and Σψi² < ∞
Applies to covariance stationary processes. εt is white noise. Any AR or ARMA model that is stationary has such a form.
Lag operator
L Yt = Yt−1; Lⁿ Yt = Yt−n
AR(1): (1 − φL)Yt = c + εt. For |φ| < 1, Yt = (c ÷ (1 − φ)) + Σ φⁱ εt−i.
AR(1) multi-step forecast
E[Yt+h | Yt] = μ + φʰ (Yt − μ), where μ = c ÷ (1 − φ)
Valid for |φ| < 1. The forecast decays geometrically to the mean μ. One-step: Ŷt+1 = c + φYt.

How to solve Model Selection, Estimation and Forecasting questions

Use this order for any question on selection, diagnostics or forecasting.

  1. 1Identify what is asked: choosing a model, testing residuals, or forecasting.
  2. 2For selection, list each model's criterion value and pick the lowest. If you must compute, use the formula given and the model's k and T. Remember BIC and AIC can disagree.
  3. 3For a Q test, compute the statistic from the autocorrelations (Box-Pierce: T × sum of squares; Ljung-Box: weight each term by T(T+2) ÷ (T−k)).
  4. 4Set degrees of freedom to m (or m minus the number of fitted ARMA parameters if the question says residuals of a fitted model) and compare Q with the chi-squared critical value.
  5. 5Conclude: Q above the critical value means reject white noise, so residual autocorrelation remains and the model is inadequate.
  6. 6For forecasts, compute the mean μ = c ÷ (1 − φ) first, then apply the one-step rule recursively or use μ + φʰ(Yt − μ).
  7. 7Set future shocks to zero. Check that your long-horizon forecast moves toward μ.
  8. 8Sanity check the sign and size of the answer against the last observation and the mean.

Quickest way: Shortcut for AR(1) forecasts and Q tests

When to use it: Use when the question gives c and φ for an AR(1), or a short list of autocorrelations.

  1. AR(1) forecast: find μ = c ÷ (1 − φ). Then forecast = μ + φʰ × (last value − μ). This avoids h rounds of recursion.
  2. For Q tests with a few lags, square each autocorrelation, add, then multiply by T. For Ljung-Box, divide each squared term by (T − k) first, then multiply by T(T+2).
  3. Remember Ljung-Box is always slightly larger than Box-Pierce for the same data, since T(T+2) ÷ (T−k) > T.
  4. For AIC vs BIC with options, eliminate answers that pick a model with a higher criterion value, then check the question for which criterion it names.

Common mistakes in Model Selection, Estimation and Forecasting

  • Choosing the model with the highest AIC or BIC.

    Students link a larger number with a better fit, as with R².

    Fix: Lower is better for both criteria. Say it to yourself before reading the options.

  • Saying BIC always picks a smaller model than AIC regardless of sample size.

    Students overstate the rule of thumb.

    Fix: BIC's penalty per parameter is ln(T) against AIC's 2, so it is heavier when T ≥ 8. State it as a condition on sample size.

  • Reading a significant Q statistic as proof the model is good.

    Confusion over the null hypothesis.

    Fix: The null is no autocorrelation (white noise). A large Q rejects it, which means the model is inadequate.

  • Forecasting an AR(1) by adding φ each step without subtracting the mean.

    Students apply Yt+1 = φYt when the model has an intercept.

    Fix: Use Ŷt+1 = c + φYt, or work in deviations from μ with μ + φʰ(Yt − μ).

  • Including future shocks in the forecast.

    Students keep εt+1 in the equation.

    Fix: Future shocks have expectation zero, so drop them. Only the model's known parameters and past values matter.

  • Confusing Wold's theorem with a claim that all series are stationary.

    The theorem sounds universal.

    Fix: It applies to covariance stationary series only. Nonstationary series need differencing or detrending first.

Worked examples

Example 1

An AR(1) model is estimated as Yt = 2 + 0.6 Yt−1 + εt. The last observation is Yt = 8. Forecast Y two periods ahead.

Show the solution
  1. Mean: μ = c ÷ (1 − φ) = 2 ÷ 0.4 = 5.
  2. One-step forecast: Ŷt+1 = 2 + 0.6 × 8 = 6.8.
  3. Two-step forecast: Ŷt+2 = 2 + 0.6 × 6.8 = 6.08.
  4. Check with the shortcut: 5 + 0.6² × (8 − 5) = 5 + 0.36 × 3 = 6.08.

Answer: 6.08

Example 2

For T = 100 residuals from a fitted model, the first three sample autocorrelations are 0.20, −0.10 and 0.10. Compute the Box-Pierce Q statistic for m = 3 and decide at the 5% level, given the chi-squared critical value with 3 degrees of freedom is 7.81.

Show the solution
  1. Square the autocorrelations: 0.04, 0.01, 0.01. Sum = 0.06.
  2. Q_BP = T × sum = 100 × 0.06 = 6.0.
  3. Compare: 6.0 < 7.81, so we do not reject the null.
  4. Interpretation: no evidence of remaining autocorrelation in the first three lags.

Answer: Q = 6.0, below 7.81, so you fail to reject white noise residuals.

Exam tips

  • When a question lists AIC and BIC for several models, read which criterion it names. The two can point to different models.
  • Expect a conceptual question on Wold's theorem: it requires covariance stationarity and expresses the process as a moving average of white noise.
  • For multi-step AR(1) forecasts, compute the mean first. It gives you the long-run limit and a quick check on your answer.
  • A Q test null is always no autocorrelation. Memorise this so you do not reverse the conclusion under time pressure.
  • Know that Ljung-Box is preferred over Box-Pierce in small samples.

Practice questions from Stationary Time Series

Model Selection, Estimation and Forecasting: frequently asked questions

What is the difference between the Box-Pierce and Ljung-Box tests?

Both test whether the first m autocorrelations are jointly zero. Ljung-Box scales each squared autocorrelation by T(T+2) ÷ (T−k), which improves accuracy in small samples. For large T the two give nearly the same result.

Should I use AIC or BIC to select a model?

Both choose the lowest value. AIC suits forecasting because it tolerates a bit more complexity. BIC penalises parameters more when T ≥ 8 and is consistent, picking the true model in large samples if it is among those considered.

What does Wold's representation theorem say?

Any covariance stationary process can be written as a deterministic part plus an infinite weighted sum of current and past white noise shocks. It justifies using ARMA models to approximate stationary series.

How does an AR(1) forecast behave as the horizon grows?

If |φ| < 1, the forecast moves geometrically toward the unconditional mean μ = c ÷ (1 − φ). The gap from the mean shrinks by a factor of φ each step.