Mathematics · Ch 13 — Statistics
Measures of Dispersion
Measures of Dispersion
13.2 Measures of Dispersion
When you have a set of data, the central tendency (mean, median, mode) tells you where the data clusters. But two different datasets can have the same mean yet look completely different. One might have all values huddled close to the mean, while the other spreads out wildly. That spread — how far the observations are from the centre — is what dispersion measures.
The textbook lists four common measures of dispersion: Range, Quartile deviation, Mean deviation, and Standard deviation. In this chapter, you will study all of these except quartile deviation.
The choice of which measure to use depends on the nature of the data and what you want to emphasise. Range is quick but crude; standard deviation is the most widely used in advanced statistics.
Range
The range is the simplest measure of dispersion. It is the difference between the largest and the smallest observation in the data.
For example, if the marks of 10 students are 23, 45, 67, 12, 89, 34, 56, 78, 90, 11, then the range is .
Range depends only on the two extreme values. It tells you nothing about how the rest of the data is distributed. A single outlier can make the range misleadingly large.
Mean Deviation
The mean deviation measures the average of the absolute deviations from a central value (usually the mean or the median). Because absolute values are used, all deviations are treated as positive — this avoids the problem of positive and negative deviations cancelling each other out.
Mean Deviation for Ungrouped Data
Let be observations. The mean deviation about the mean is defined as:
Similarly, the mean deviation about the median is:
When the data has extreme values, the median is a more robust central value than the mean. In such cases, mean deviation about the median is often preferred.
Mean Deviation for Grouped Data
For grouped data (frequency distributions), the formulas adapt to include frequencies. Let be the midpoints of classes (or the actual values for discrete data) and be the corresponding frequencies, with .
Mean deviation about the mean:
Mean deviation about the median:
How to Find the Median for Grouped Data
For a continuous frequency distribution, the median is found using the formula:
where:
- = lower limit of the median class
- = total frequency
- = cumulative frequency of the class preceding the median class
- = frequency of the median class
- = class size
The median class is the class whose cumulative frequency is just greater than or equal to .
Properties of Mean Deviation (with Proofs)
The textbook lists three important properties of mean deviation. Each one is proved step by step below.
›Proof
Property I: The mean deviation about the mean is always zero if we do not take absolute values. That is, .
Proof:
But , so .
Therefore, .
This is why we use absolute values — without them, the positive and negative deviations would cancel out, giving a misleading zero dispersion.
›Proof
Property II: The mean deviation about the median is minimum when compared to the mean deviation about any other point.
Proof (outline):
Let be any number. Consider .
The function is piecewise linear. Its slope changes at each .
For less than the smallest , the slope is (since each term , derivative ).
As increases past each , the slope increases by .
The minimum occurs where the slope changes from negative to positive — that is, at the median.
For an even number of observations, any value between the two middle observations gives the same minimum sum.
›Proof
Property III: The mean deviation about the mean is always less than or equal to the standard deviation (which we will study next). More precisely, .
Proof:
By the Cauchy-Schwarz inequality:
Taking square roots:
Dividing both sides by :
The left side is , and the right side is .
Actually, the textbook states the result as directly (without the factor). This is a simplified version — the strict inequality holds for unless all deviations are equal.
Standard Deviation
The standard deviation is the most important measure of dispersion. It is the square root of the variance, which is the average of the squared deviations from the mean.
Variance and Standard Deviation for Ungrouped Data
Let be observations with mean .
Variance ():
Standard deviation ():
Squaring the deviations eliminates the sign problem (unlike mean deviation which uses absolute values). But squaring also gives more weight to larger deviations — a point far from the mean contributes more to the variance than several points close to it.
A Shortcut Formula for Variance
Expanding the square gives a computationally simpler form:
Or equivalently:
Use this shortcut formula when calculating by hand — it avoids computing each deviation separately.
Variance and Standard Deviation for Grouped Data
For grouped data with class midpoints and frequencies (where ):
Variance:
Shortcut formula:
Standard deviation:
Properties of Standard Deviation (with Proofs)
›Proof
Property I: The standard deviation is independent of the change of origin but not of scale.
That is, if , then .
Proof:
Let . Then .
The mean of is .
So .
Therefore, …