Risk Modelling and Survival Analysis · Applications of time series models
Forecasting with AR, MA and ARIMA Models
Updated 11 October 2026 · Fact-checked
Forecasting with a time series model means predicting X(n+k) from data up to time n. Replace future noise terms by zero and future values by their own forecasts, working forward one step at a time. The forecast error variance is σ² times the sum of the first k squared ψ weights. Exponential smoothing uses a weighted average of past values.
Understand Forecasting with Time Series Models
A forecast answers one question: given everything observed up to time n, what is the best estimate of X(n+k)? In CS2 "best" means the minimum mean square error forecast. This is the conditional expectation of X(n+k) given the history up to time n.
To compute it, use the model equation and apply two rules. Future noise terms e(n+1), e(n+2), ... have mean zero and are unrelated to the past, so you replace them by 0. Future values of X are replaced by their own forecasts. Past noise terms (up to time n) are treated as known. You work forward: first 1-step, then 2-step, and so on.
For a stationary AR(1), the forecast decays geometrically towards the mean μ as k grows. For an MA(q), the forecast is just μ once k is larger than q, because the memory of the process is only q steps long. For ARMA models, the forecast follows the AR part after the MA terms have run out.
Every forecast has an error. The k-step forecast error is the sum of the future noise terms that you could not know. Its variance rises with k. For a stationary model it levels off at the variance of the process. For an ARIMA model with differencing it keeps growing without limit.
Exponential smoothing is a simpler, model-free method. It updates a smoothed level using a weight α between 0 and 1. It gives more weight to recent observations and is suited to series with no clear trend or seasonality. Its forecast for every future step is the latest smoothed level.
Key rules to remember
- Best forecast
- x̂(n, k) = E[X(n+k) | X(n), X(n−1), ...]
- Replace future e by 0, future X by forecasts, and keep past X and past e as known values.
- AR(1) k-step forecast
- x̂(n, k) = μ + α^k (x(n) − μ)
- For X(t) − μ = α(X(t−1) − μ) + e(t). It tends to μ as k grows if |α| < 1.
- MA(q) forecast
- x̂(n, k) = μ for k > q
- For k ≤ q, use the known past e terms with the correct θ coefficients, and set future e to 0.
- MA(∞) form and ψ weights
- X(t) = μ + Σ ψ(j) e(t−j), with ψ(0) = 1
- For ARMA(1,1) with parameters α and θ: ψ(1) = α + θ and ψ(j) = α ψ(j−1) for j ≥ 2.
- k-step forecast error
- e(n, k) = X(n+k) − x̂(n, k) = Σ (j = 0 to k−1) ψ(j) e(n+k−j)
- It has mean zero for the best forecast.
- Forecast error variance
- Var[e(n, k)] = σ² Σ (j = 0 to k−1) ψ(j)²
- For AR(1): σ²(1 − α^(2k)) ÷ (1 − α²). For a 1-step forecast it equals σ².
- ARIMA(p,1,q) forecasting
- Forecast the differenced series Y(t) = X(t) − X(t−1), then x̂(n, k) = x(n) + Σ (j = 1 to k) ŷ(n, j)
- For a random walk, x̂(n, k) = x(n) and the error variance is kσ².
- Exponential smoothing
- m̂(t) = α x(t) + (1 − α) m̂(t−1), and x̂(n, k) = m̂(n) for all k
- Equivalent weights on past observations are α(1 − α)^j. A larger α reacts faster.
How to solve Forecasting with Time Series Models questions
Use this routine for any forecasting question on AR, MA, ARMA or ARIMA models.
- 1Write the model in the form given. Identify μ (or whether the mean is zero), the parameters and σ². Check if the series is differenced.
- 2If the model is ARIMA, define the differenced series Y and write the ARMA model for Y. Keep x(n) to add back later.
- 3List what is known at time n: the latest observations and the latest noise terms e(n), e(n−1), ... Estimate noise terms from the model if they are not given.
- 4Compute the 1-step forecast by putting time n+1 into the equation. Set e(n+1) to 0 and use known values.
- 5Move forward step by step. In each new equation, replace unknown future X by the forecast you just found, and future e by 0.
- 6If the model is ARIMA, add the forecast differences back to x(n) to return to the original series.
- 7For the error variance, find the ψ weights up to ψ(k−1), then compute σ² Σ ψ(j)². For AR(1) use the closed form.
- 8State the result with its units, and say if the forecast converges to the mean or the variance grows without bound.
Quickest way: Forecast by recursion, error variance by ψ weights
When to use it: Use this in timed written questions or MCQs asking for a 1-, 2- or 3-step forecast or its variance.
- For AR(1), go straight to μ + α^k (x(n) − μ). Do not compute step by step.
- For MA(q), write the forecast as μ plus only the terms with e at or before time n. All other terms vanish.
- For ARMA(1,1), compute the 1-step forecast with the known e(n). After that the forecast is x̂(n, k) = α x̂(n, k−1) with the mean taken out.
- For variance, list ψ(0) = 1, ψ(1), ψ(2) and so on, square and add, then multiply by σ².
- Sanity check: the 1-step variance must be σ², and the variance must not fall as k increases.
Common mistakes in Forecasting with Time Series Models
Forgetting to subtract the mean before applying the AR decay factor.
Students remember α^k x(n) from zero-mean examples.
Fix: Always use μ + α^k (x(n) − μ). The decay applies to the deviation from the mean, not to the level.
Replacing a past noise term e(n) by zero in an MA or ARMA forecast.
Students remember that noise has mean zero but forget it only applies to future noise.
Fix: Only e(n+1), e(n+2), ... become zero. Terms up to time n are known and must be kept.
Summing too many or too few ψ weights in the error variance.
Confusion between k and k−1 in the sum limits.
Fix: A k-step error has k terms: ψ(0) to ψ(k−1). Check that the 1-step variance is σ².
Forecasting an ARIMA model as if the original series were stationary.
Students forget to difference first, or forget to add the forecast differences back.
Fix: Forecast the differenced series, then cumulate the forecasts onto x(n). Expect the error variance to grow without bound.
Using the 2-step forecast formula for AR(1) as x̂ = α x(n) instead of α² (x(n) − μ) + μ.
Students stop after one step of the recursion.
Fix: Each extra step multiplies the deviation by α once more. The power of α equals k.
Thinking exponential smoothing forecasts change with k.
Students mix it up with AR forecasts that decay.
Fix: In simple exponential smoothing the forecast for every future step is the same latest smoothed level m̂(n).
Worked examples
Example 1
A stationary AR(1) process satisfies X(t) − 50 = 0.6 (X(t−1) − 50) + e(t), where e(t) is white noise with variance 16. The latest observation is x(n) = 58. Find the 1-step and 3-step forecasts, and the variance of the 3-step forecast error.
Show the solution
- The deviation from the mean at time n is 58 − 50 = 8.
- 1-step forecast: 50 + 0.6 × 8 = 54.8.
- 3-step forecast: 50 + 0.6³ × 8 = 50 + 0.216 × 8 = 50 + 1.728 = 51.728.
- For the variance, ψ(j) = 0.6^j, so ψ(0) = 1, ψ(1) = 0.6, ψ(2) = 0.36.
- Sum of squares: 1 + 0.36 + 0.1296 = 1.4896.
- Variance = 16 × 1.4896 = 23.8336. Check with the closed form: 16 × (1 − 0.6⁶) ÷ (1 − 0.36) = 16 × 0.953344 ÷ 0.64 = 23.8336.
Answer: 1-step forecast = 54.8; 3-step forecast = 51.728; 3-step error variance = 23.8336 (about 23.83).
Example 2
A zero-mean ARMA(1,1) process is X(t) = 0.5 X(t−1) + e(t) + 0.4 e(t−1), with Var[e(t)] = 9. At time n, x(n) = 2 and the estimated noise term is e(n) = 0.5. Find the 1-step and 2-step forecasts and the variance of the 2-step forecast error.
Show the solution
- 1-step forecast: x̂(n, 1) = 0.5 × 2 + 0.4 × 0.5, with e(n+1) set to 0. This gives 1 + 0.2 = 1.2.
- 2-step forecast: x̂(n, 2) = 0.5 × x̂(n, 1) + 0 + 0.4 × 0, because both e(n+2) and e(n+1) are future terms and become 0. This gives 0.5 × 1.2 = 0.6.
- ψ weights: ψ(0) = 1 and ψ(1) = α + θ = 0.5 + 0.4 = 0.9.
- A 2-step error has terms ψ(0) e(n+2) and ψ(1) e(n+1).
- Variance = 9 × (1 + 0.9²) = 9 × 1.81 = 16.29.
Answer: 1-step forecast = 1.2; 2-step forecast = 0.6; 2-step error variance = 16.29.
Exam tips
- Write down the forecast rule in one line before you calculate: future e = 0, future X = forecast. Examiners give method marks for this.
- For written ARMA questions, you may need to estimate e(n) from earlier data first. Compute it recursively from the model equation and state the starting assumption.
- Show the ψ weights in a short list when asked for error variance. The final number alone can lose method marks if it is wrong.
- In ARIMA questions, say clearly that the variance grows without bound and that prediction intervals widen. Link this to a practical point such as long-term projections being unreliable.
- In Paper B (R), check that forecasts from your fitted model agree with a hand calculation for 1-step, and report the order, parameters and forecasts together.
Practice questions from Applications of time series models
- Consider the process Y_t = 1.5 Y_{t-1} - 0.5 Y_{t-2} + e_t, with e_t white noise. Which statement is correct?
- An ARMA model is fitted to a series of 200 observations and the residual autocorrelations at lags 1 to 10 are computed. A Ljung-Box (portman…
- The sample partial autocorrelation function of a stationary series of 400 observations shows significant spikes at lags 1 and 2 and is negli…
- A time series follows the process X_t = 0.6 X_{t-1} + e_t, where e_t is white noise with variance 4. Which statement about this process is c…
- A sample ACF of a monthly premium series decays very slowly from a value near 0.95 at lag 1 and stays large for many lags. What is the most …
Forecasting with Time Series Models in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Forecasting with Time Series Models: frequently asked questions
How do I forecast an AR(1) process k steps ahead?
Use x̂(n, k) = μ + α^k (x(n) − μ). The forecast moves from the last observation towards the mean, and the gap shrinks by a factor α each step. If |α| < 1 it converges to μ.
What is the forecast error variance for an ARIMA model?
It is σ² times the sum of the first k squared ψ weights, where the ψ weights come from the MA(∞) representation of the integrated process. With differencing, the ψ weights do not die away, so the variance keeps growing as k increases.
What is the exponential smoothing forecast formula?
The smoothed level is m̂(t) = α x(t) + (1 − α) m̂(t−1), and the forecast for any future step is the latest level m̂(n). For example, with α = 0.2, previous level 100 and new observation 110, the new level is 0.2 × 110 + 0.8 × 100 = 102.
Why does an MA(q) forecast equal the mean after q steps?
Once you look more than q steps ahead, every noise term in the equation is in the future and has expected value zero. Only μ is left. The process has no memory beyond q periods.