Poisson Goodness of Fit: From Intuition to Precision
Imagine you manage a small call centre. You record the number of calls arriving each minute for 100 minutes. You get a list: 0 calls in some minutes, 1 call in others, 2 calls in a few, and so on. Your hunch is that calls arrive randomly and independently — a classic Poisson process. But is your data actually consistent with a Poisson distribution? That is exactly what a Poisson Goodness of Fit test answers.
The Intuition: Does the Pattern Fit the Model?
A Poisson distribution has one parameter, λ (the average rate). If your data truly comes from a Poisson process, then the proportion of minutes with 0 calls should be close to e−λ, the proportion with 1 call should be close to λe−λ, and so on. The test compares what you observed in your data to what you expected under the Poisson model.
If the observed and expected counts are very close, the Poisson model is a good fit. If they are far apart, something is wrong — maybe calls are not independent (e.g., a TV ad causes a burst), or the rate is not constant.
A common mistake is to think this test proves the data is Poisson. It only says: "If the data were Poisson, how unlikely would these differences be?" A small p-value means the data is inconsistent with Poisson; a large p-value means we cannot reject the Poisson model — not that it is definitely correct.
The Precise Statement
We have n independent observations, each a count (0, 1, 2, …). We want to test:
- H0: The data come from a Poisson distribution (with unknown λ).
- H1: The data do not come from a Poisson distribution.
Step 1: Estimate λ.
Since λ is unknown, we estimate it from the data: λ^=xˉ, the sample mean.
Step 2: Group the data into categories.
Because counts can go up to large numbers, we group them into k bins. Typically: 0, 1, 2, …, up to some maximum, and then a final bin for "all counts ≥m". Each bin must have an expected count of at least 5 (a standard rule for the chi-square approximation to be valid).
Step 3: Compute expected frequencies under H0.
For each bin i, the expected probability under Poisson(λ^) is:
- For bin "0": P(X=0)=e−λ^
- For bin "1": P(X=1)=λ^e−λ^
- For bin "2": P(X=2)=2!λ^2e−λ^
- …
- For the last bin "≥m": P(X≥m)=1−∑j=0m−1P(X=j)
Then multiply each probability by n to get the expected count Ei.
Step 4: Compute the test statistic.
χ2=∑i=1kEi(Oi−Ei)2
where Oi is the observed count in bin i, and Ei is the expected count.
Step 5: Compare to a chi-square distribution.
Under H0, this statistic approximately follows a χ2 distribution with degrees of freedom:
df=k−1−p
where p is the number of parameters estimated from the data. For Poisson, p=1 (we estimated λ), so df=k−2.
χ2=∑E(O−E)2with df=(number of bins)−2
Step 6: Make a decision.
If the calculated χ2 is larger than the critical value from the χ2 table (at your chosen significance level, say 0.05), reject H0 — the data does not fit a Poisson distribution. Otherwise, fail to reject H0.
A Worked Example (Brief)
Suppose you observe 100 minutes of call data:
| Calls per minute | Observed count |
|---|
| 0 | 30 |
| 1 | 40 |
| 2 | 20 |
| 3 or more | 10 |
Sample mean λ^=1000⋅30+1⋅40+2⋅20+3⋅10=1.1.
Expected probabilities:
P(0)=e−1.1≈0.3329, so E0=33.29
P(1)=1.1e−1.1≈0.3662, so E1=36.62
P(2)=21.12e−1.1≈0.2014, so E2=20.14
P(≥3)=1−(0.3329+0.3662+0.2014)=0.0995, so E3=9.95
Compute χ2:
χ2=33.29(30−33.29)2+36.62(40−36.62)2+20.14(20−20.14)2+9.95(10−9.95)2≈0.325+0.312+0.001+0.000≈0.638
Degrees of freedom: 4−2=2.
Critical value at α=0.05: about 5.99.
Since 0.638<5.99, we fail to reject H0 — the data is consistent with a Poisson distribution.
Always check that no expected count is below 5. If one is, combine that bin with a neighbour. This preserves the validity of the chi-square approximation.