Actuarial Statistics · Hypothesis testing and goodness of fit
Kolmogorov-Smirnov, Runs and Other Goodness of Fit Tests
Updated 11 October 2026 · Fact-checked
The Kolmogorov-Smirnov test checks whether data come from a stated distribution. Find D, the largest gap between the empirical CDF and the fitted CDF, and compare it with a critical value. Runs, signs, serial correlation and Q-Q checks test whether residuals look random and normal. A large D or too few runs means poor fit.
Understand Kolmogorov-Smirnov, Runs and Other Fit Tests
A fit test asks one question: do the data look like they came from the model I have chosen? The chi-square test answers this by grouping data. Other tests use the individual values and so lose less information.
The Kolmogorov-Smirnov (KS) test compares two curves. The empirical distribution function F_n(x) is the proportion of sample values ≤ x. It is a staircase. The fitted CDF F0(x) is the smooth curve from your model. The test statistic D is the biggest vertical distance between them. If D is large, the model does not fit. The null hypothesis is that the sample comes from F0. The test applies to continuous distributions. If you estimate parameters from the same data, the standard critical values are too large, so the test becomes conservative. State this caveat in an answer.
The Anderson-Darling test has the same idea but weights the gaps by how far into the tails they are. It is more sensitive to poor fit in the tails than KS, which is often most sensitive near the middle. For actuarial work, tails matter, so this is a real advantage. You are expected to know the idea, not to compute it in most questions.
After fitting a model such as a regression or time series, you test the residuals. If the model is right, residuals should look like independent random noise with mean zero and, usually, a normal shape. The signs test counts positive residuals. The runs test counts changes of sign. Serial correlation checks whether neighbouring residuals are related. A Q-Q plot checks the shape against a normal distribution.
The common thread is this: each test looks for one kind of pattern. Too many positives, too few runs, a large lag-1 correlation or a bent Q-Q line each point to a specific model fault.
Key rules to remember
- Empirical distribution function
- F_n(x) = (number of sample values ≤ x) ÷ n
- A step function. It jumps by 1/n at each distinct observation.
- KS statistic
- D = max over x of |F_n(x) − F0(x)|
- Check at each ordered value x(i) both |i/n − F0(x(i))| and |(i−1)/n − F0(x(i))|. Take the largest of all.
- KS decision rule
- Reject H0 if D > critical value for n and the chosen significance level
- Use the tables you are given. For large n the 5% critical value is roughly 1.36 ÷ √n. Check the tables book before relying on this.
- Anderson-Darling statistic
- A² = −n − (1/n) Σ (2i − 1)[ln z_i + ln(1 − z_(n+1−i))], where z_i = F0(x(i))
- Ordered data x(1) ≤ … ≤ x(n). Larger A² means worse fit, with extra weight on the tails.
- Signs test
- Under H0, number of positive residuals ~ Binomial(n, ½)
- Drop zero residuals and reduce n. Use a normal approximation with mean n/2 and variance n/4 for large n.
- Runs test (mean)
- E(R) = 2·n1·n2 ÷ (n1 + n2) + 1
- n1 positive and n2 negative residuals. A run is an unbroken sequence of the same sign.
- Runs test (variance)
- Var(R) = 2·n1·n2·(2·n1·n2 − n1 − n2) ÷ [(n1 + n2)²·(n1 + n2 − 1)]
- For large n1 and n2, R is approximately normal. Use z = (R ± 0.5 − E(R)) ÷ √Var(R) with a continuity correction.
- Lag-1 serial correlation
- r1 = Σ (e_t − ē)(e_(t+1) − ē) ÷ Σ (e_t − ē)²
- Numerator sum runs over t = 1 to n−1. Under independence, r1 is approximately N(0, 1/n), so compare |r1| with 1.96 ÷ √n at 5%.
- Q-Q plot
- Plot ordered sample values x(i) against theoretical quantiles of the fitted distribution
- Points near a straight line mean good fit. For normality, plot against standard normal quantiles.
How to solve Kolmogorov-Smirnov, Runs and Other Fit Tests questions
Use this method for any question on KS, runs, signs, serial correlation or Q-Q plots.
- 1Identify what is being tested: a distribution (KS, Anderson-Darling, Q-Q) or the randomness of residuals (signs, runs, serial correlation).
- 2State H0 and H1 in words. For example, H0: the sample comes from the stated distribution, or H0: residuals are independent with median zero.
- 3Order the data or list the residual signs in time order. Count n, n1, n2 or runs as needed.
- 4Compute the statistic. For KS, find F_n and F0 at every point and take the largest gap, checking both sides of each step. For runs, find E(R) and Var(R), then z.
- 5Compare with the critical value or normal quantile. State the significance level you use.
- 6State the conclusion in context. Say what the result implies for the model, for example autocorrelation, wrong distribution or heavy tails.
- 7Add a caveat if relevant: parameters estimated from the data, small sample size, or low power of the test.
Quickest way: Fast KS calculation with a table
When to use it: Use this when you are given a small sample and a fully specified distribution and asked to carry out the KS test.
- Sort the data. Write columns: i, x(i), F0(x(i)), i/n, (i−1)/n.
- Add two difference columns: |i/n − F0| and |(i−1)/n − F0|.
- Take the single largest entry in both columns. That is D.
- Read the critical value from the tables for your n and level. Reject if D exceeds it.
- Write one sentence of interpretation. Mention that the test is for continuous distributions and that estimating parameters makes it conservative.
Common mistakes in Kolmogorov-Smirnov, Runs and Other Fit Tests
Taking the difference only at i/n and forgetting (i−1)/n.
The empirical CDF jumps at each point, so the gap can be largest just before the jump. Students use only the top of each step.
Fix: Compute both gaps at every ordered value. The maximum may be on either side.
Using the unsorted data in the KS table.
Rushing. F_n only makes sense with values in increasing order.
Fix: Sort first and number the values 1 to n before computing anything.
Using KS critical values as exact when parameters were estimated from the same data.
Students treat the table as always valid.
Fix: State that the test is conservative in this case and that the real p-value is smaller than the table suggests.
Counting runs wrongly, for example counting the number of sign changes instead of runs.
A run is a block, not a change. The two differ by one.
Fix: Number of runs = number of sign changes + 1. Write the sign sequence and underline each block.
Concluding the model is correct when the test does not reject.
Failing to reject is read as proof.
Fix: Say there is no evidence against the model at this level. Small samples give tests low power.
Misreading a Q-Q plot, for example thinking any curve means non-normal data without describing it.
Students memorise 'straight line is good' but not what the departures mean.
Fix: Describe the shape. Ends bending away from the line on both sides suggest heavy tails. A curve that bends one way suggests skewness.
Worked examples
Example 1
A sample of five observations is 0.12, 0.35, 0.48, 0.77 and 0.91. Test at the 5% level whether it comes from the uniform distribution on (0, 1) using the Kolmogorov-Smirnov test. The tables give a 5% critical value of about 0.56 for n = 5.
Show the solution
- H0: the sample comes from U(0, 1). H1: it does not. Here F0(x) = x for 0 ≤ x ≤ 1, and n = 5.
- The data are already ordered. For each x(i), compute |(i−1)/n − F0| and |i/n − F0|.
- i = 1, x = 0.12: |0 − 0.12| = 0.12 and |0.2 − 0.12| = 0.08.
- i = 2, x = 0.35: |0.2 − 0.35| = 0.15 and |0.4 − 0.35| = 0.05.
- i = 3, x = 0.48: |0.4 − 0.48| = 0.08 and |0.6 − 0.48| = 0.12.
- i = 4, x = 0.77: |0.6 − 0.77| = 0.17 and |0.8 − 0.77| = 0.03.
- i = 5, x = 0.91: |0.8 − 0.91| = 0.11 and |1 − 0.91| = 0.09.
- The largest gap is D = 0.17.
- Compare: 0.17 is less than about 0.56, so we do not reject H0.
Answer: D = 0.17, which is below the critical value of about 0.56. There is no evidence at the 5% level against the uniform distribution. With only five points, the test has low power.
Example 2
A time series model is fitted and 20 residuals are examined. There are 12 positive and 8 negative residuals in time order, forming 5 runs. Use the runs test with a normal approximation and continuity correction to test at the 5% level whether the signs are random.
Show the solution
- H0: the order of signs is random. H1: it is not. Here n1 = 12, n2 = 8, and the number of runs is R = 5.
- Mean: E(R) = 2 × 12 × 8 ÷ 20 + 1 = 192 ÷ 20 + 1 = 10.6.
- Variance: Var(R) = 2 × 96 × (192 − 20) ÷ (20² × 19) = 192 × 172 ÷ 7,600 = 33,024 ÷ 7,600 = 4.345.
- Standard deviation = √4.345 = 2.085.
- Because R is small, use the continuity correction R + 0.5 = 5.5. Then z = (5.5 − 10.6) ÷ 2.085 = −2.45.
- For a two-sided 5% test the critical values are ±1.96. Since −2.45 < −1.96, reject H0.
- Too few runs means the residuals stay on the same side for long stretches.
Answer: z ≈ −2.45, so reject H0 at the 5% level. The residuals show fewer runs than expected, which points to positive serial correlation or a missing trend. The model should be revised.
Exam tips
- In a KS calculation, show a table. Marks are for the working columns and the maximum, not only for D.
- Always write H0 and H1 and a one-line conclusion in context. Examiners reward interpretation.
- Know what each residual test detects: signs for bias, runs for patterns in order, serial correlation for dependence, Q-Q for shape.
- If asked to compare KS with chi-square or Anderson-Darling, state that KS suits continuous data without grouping, and Anderson-Darling is more sensitive in the tails.
- For computer-based work, state which function or command you used and show its output together with your interpretation.
Practice questions from Hypothesis testing and goodness of fit
- A Kolmogorov-Smirnov test compares an empirical distribution function with a fully specified continuous distribution. Which statement about …
- A test of H0: μ = 50 against H1: μ > 50 is carried out at the 5% significance level. Which statement correctly describes the probability of …
- A die is rolled 120 times to test whether it is fair. The observed frequencies for faces 1 to 6 are 15, 25, 20, 18, 22 and 20. What is the v…
- In a chi-square goodness of fit test, an expected frequency in one tail cell is only 1.2 while all others exceed 5. What is the standard rec…
- A chi-square goodness-of-fit test checks whether the number of claims per policy follows a Poisson distribution whose mean is estimated from…
Kolmogorov-Smirnov, Runs and Other Fit Tests in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Kolmogorov-Smirnov, Runs and Other Fit Tests: frequently asked questions
How do I calculate the Kolmogorov-Smirnov statistic by hand?
Sort the data and compute F0 at each value. For each ordered value, find the gaps to i/n and to (i−1)/n. The largest of these gaps is D. Compare D with the critical value in the tables.
What is the difference between the runs test and the signs test?
The signs test counts how many residuals are positive and checks the count against Binomial(n, ½). It ignores order. The runs test counts blocks of equal sign and so detects patterns in the order of residuals.
How do I read a Q-Q plot to check normality?
Plot the ordered residuals against standard normal quantiles. If the points lie close to a straight line, the normal assumption is reasonable. Curves or ends that bend away from the line show skewness or heavy tails.
Why is the Anderson-Darling test preferred to KS for tails?
Anderson-Darling gives more weight to differences in the tails of the distribution. KS treats all gaps equally and is often most sensitive near the centre. For risk work, tail fit is important.