Skip to content

Mathematics · Ch 13 — Statistics

Measures of Dispersion

13.2

Measures of Dispersion

13.2 Measures of Dispersion

When you have a set of data, the central tendency (mean, median, mode) tells you where the data clusters. But two different datasets can have the same mean yet look completely different. One might have all values huddled close to the mean, while the other spreads out wildly. That spread — how far the observations are from the centre — is what dispersion measures.

The textbook lists four common measures of dispersion: Range, Quartile deviation, Mean deviation, and Standard deviation. In this chapter, you will study all of these except quartile deviation.

Note

The choice of which measure to use depends on the nature of the data and what you want to emphasise. Range is quick but crude; standard deviation is the most widely used in advanced statistics.


Range

The range is the simplest measure of dispersion. It is the difference between the largest and the smallest observation in the data.

Range=Maximum value−Minimum value\text{Range} = \text{Maximum value} - \text{Minimum value}

For example, if the marks of 10 students are 23, 45, 67, 12, 89, 34, 56, 78, 90, 11, then the range is 90−11=7990 - 11 = 79.

Watch out

Range depends only on the two extreme values. It tells you nothing about how the rest of the data is distributed. A single outlier can make the range misleadingly large.


Mean Deviation

The mean deviation measures the average of the absolute deviations from a central value (usually the mean or the median). Because absolute values are used, all deviations are treated as positive — this avoids the problem of positive and negative deviations cancelling each other out.

Mean Deviation for Ungrouped Data

Let x1,x2,…,xnx_1, x_2, \dots, x_n be nn observations. The mean deviation about the mean xˉ\bar{x} is defined as:

M.D.(xˉ)=1n∑i=1n∣xi−xˉ∣\text{M.D.}(\bar{x}) = \frac{1}{n} \sum_{i=1}^{n} |x_i - \bar{x}|

Similarly, the mean deviation about the median MM is:

M.D.(M)=1n∑i=1n∣xi−M∣\text{M.D.}(M) = \frac{1}{n} \sum_{i=1}^{n} |x_i - M|

Tip

When the data has extreme values, the median is a more robust central value than the mean. In such cases, mean deviation about the median is often preferred.

Mean Deviation for Grouped Data

For grouped data (frequency distributions), the formulas adapt to include frequencies. Let xix_i be the midpoints of classes (or the actual values for discrete data) and fif_i be the corresponding frequencies, with N=∑fiN = \sum f_i.

Mean deviation about the mean:

M.D.(xˉ)=1N∑i=1nfi∣xi−xˉ∣\text{M.D.}(\bar{x}) = \frac{1}{N} \sum_{i=1}^{n} f_i |x_i - \bar{x}|

Mean deviation about the median:

M.D.(M)=1N∑i=1nfi∣xi−M∣\text{M.D.}(M) = \frac{1}{N} \sum_{i=1}^{n} f_i |x_i - M|

How to Find the Median for Grouped Data

For a continuous frequency distribution, the median is found using the formula:

M=l+N2−cff×hM = l + \frac{\frac{N}{2} - cf}{f} \times h

where:

  • ll = lower limit of the median class
  • NN = total frequency
  • cfcf = cumulative frequency of the class preceding the median class
  • ff = frequency of the median class
  • hh = class size
Important

The median class is the class whose cumulative frequency is just greater than or equal to N2\frac{N}{2}.


Properties of Mean Deviation (with Proofs)

The textbook lists three important properties of mean deviation. Each one is proved step by step below.

›Proof

Property I: The mean deviation about the mean is always zero if we do not take absolute values. That is, ∑i=1n(xi−xˉ)=0\sum_{i=1}^{n} (x_i - \bar{x}) = 0.

Proof:

∑i=1n(xi−xˉ)=∑i=1nxi−∑i=1nxˉ\sum_{i=1}^{n} (x_i - \bar{x}) = \sum_{i=1}^{n} x_i - \sum_{i=1}^{n} \bar{x}

=∑i=1nxi−nxˉ= \sum_{i=1}^{n} x_i - n\bar{x}

But xˉ=1n∑i=1nxi\bar{x} = \frac{1}{n} \sum_{i=1}^{n} x_i, so nxˉ=∑i=1nxin\bar{x} = \sum_{i=1}^{n} x_i.

Therefore, ∑i=1n(xi−xˉ)=∑i=1nxi−∑i=1nxi=0\sum_{i=1}^{n} (x_i - \bar{x}) = \sum_{i=1}^{n} x_i - \sum_{i=1}^{n} x_i = 0.

This is why we use absolute values — without them, the positive and negative deviations would cancel out, giving a misleading zero dispersion.

›Proof

Property II: The mean deviation about the median is minimum when compared to the mean deviation about any other point.

Proof (outline):

Let AA be any number. Consider f(A)=∑i=1n∣xi−A∣f(A) = \sum_{i=1}^{n} |x_i - A|.

The function f(A)f(A) is piecewise linear. Its slope changes at each xix_i.

For AA less than the smallest xix_i, the slope is −n-n (since each term ∣xi−A∣=xi−A|x_i - A| = x_i - A, derivative −1-1).

As AA increases past each xix_i, the slope increases by 22.

The minimum occurs where the slope changes from negative to positive — that is, at the median.

For an even number of observations, any value between the two middle observations gives the same minimum sum.

›Proof

Property III: The mean deviation about the mean is always less than or equal to the standard deviation (which we will study next). More precisely, M.D.(xˉ)≤σ\text{M.D.}(\bar{x}) \leq \sigma.

Proof:

By the Cauchy-Schwarz inequality:

(∑i=1n∣xi−xˉ∣)2≤n∑i=1n(xi−xˉ)2\left( \sum_{i=1}^{n} |x_i - \bar{x}| \right)^2 \leq n \sum_{i=1}^{n} (x_i - \bar{x})^2

Taking square roots:

∑i=1n∣xi−xˉ∣≤n∑i=1n(xi−xˉ)2\sum_{i=1}^{n} |x_i - \bar{x}| \leq \sqrt{n} \sqrt{\sum_{i=1}^{n} (x_i - \bar{x})^2}

Dividing both sides by nn:

1n∑i=1n∣xi−xˉ∣≤1n1n∑i=1n(xi−xˉ)2\frac{1}{n} \sum_{i=1}^{n} |x_i - \bar{x}| \leq \frac{1}{\sqrt{n}} \sqrt{\frac{1}{n} \sum_{i=1}^{n} (x_i - \bar{x})^2}

The left side is M.D.(xˉ)\text{M.D.}(\bar{x}), and the right side is σn\frac{\sigma}{\sqrt{n}}.

Actually, the textbook states the result as M.D.(xˉ)≤σ\text{M.D.}(\bar{x}) \leq \sigma directly (without the n\sqrt{n} factor). This is a simplified version — the strict inequality holds for n>1n > 1 unless all deviations are equal.


Standard Deviation

The standard deviation is the most important measure of dispersion. It is the square root of the variance, which is the average of the squared deviations from the mean.

Variance and Standard Deviation for Ungrouped Data

Let x1,x2,…,xnx_1, x_2, \dots, x_n be nn observations with mean xˉ\bar{x}.

Variance (σ2\sigma^2):

σ2=1n∑i=1n(xi−xˉ)2\sigma^2 = \frac{1}{n} \sum_{i=1}^{n} (x_i - \bar{x})^2

Standard deviation (σ\sigma):

σ=1n∑i=1n(xi−xˉ)2\sigma = \sqrt{\frac{1}{n} \sum_{i=1}^{n} (x_i - \bar{x})^2}

Note

Squaring the deviations eliminates the sign problem (unlike mean deviation which uses absolute values). But squaring also gives more weight to larger deviations — a point far from the mean contributes more to the variance than several points close to it.

A Shortcut Formula for Variance

Expanding the square gives a computationally simpler form:

σ2=1n∑i=1nxi2−(1n∑i=1nxi)2\sigma^2 = \frac{1}{n} \sum_{i=1}^{n} x_i^2 - \left( \frac{1}{n} \sum_{i=1}^{n} x_i \right)^2

Or equivalently:

σ2=∑xi2n−(∑xin)2\sigma^2 = \frac{\sum x_i^2}{n} - \left( \frac{\sum x_i}{n} \right)^2

Tip

Use this shortcut formula when calculating by hand — it avoids computing each deviation separately.

Variance and Standard Deviation for Grouped Data

For grouped data with class midpoints xix_i and frequencies fif_i (where N=∑fiN = \sum f_i):

Variance:

σ2=1N∑i=1nfi(xi−xˉ)2\sigma^2 = \frac{1}{N} \sum_{i=1}^{n} f_i (x_i - \bar{x})^2

Shortcut formula:

σ2=1N∑i=1nfixi2−(1N∑i=1nfixi)2\sigma^2 = \frac{1}{N} \sum_{i=1}^{n} f_i x_i^2 - \left( \frac{1}{N} \sum_{i=1}^{n} f_i x_i \right)^2

Standard deviation:

σ=1N∑i=1nfi(xi−xˉ)2\sigma = \sqrt{\frac{1}{N} \sum_{i=1}^{n} f_i (x_i - \bar{x})^2}


Properties of Standard Deviation (with Proofs)

›Proof

Property I: The standard deviation is independent of the change of origin but not of scale.

That is, if yi=xi−ahy_i = \frac{x_i - a}{h}, then σx=∣h∣σy\sigma_x = |h| \sigma_y.

Proof:

Let yi=xi−ahy_i = \frac{x_i - a}{h}. Then xi=a+hyix_i = a + h y_i.

The mean of xx is xˉ=a+hyˉ\bar{x} = a + h \bar{y}.

So xi−xˉ=(a+hyi)−(a+hyˉ)=h(yi−yˉ)x_i - \bar{x} = (a + h y_i) - (a + h \bar{y}) = h (y_i - \bar{y}).

Therefore, …