Skip to content

Informatics Practices · Ch 5 — Understanding Data

Measures of Variability

5.5.2

Measures of Variability

A single central value never tells the whole story. Two different datasets can have exactly the same mean, median or mode and yet be completely different in how their values are scattered — or, equally, share the same spread while having different centres. Measures of variability capture this second dimension: they describe the spread or variation of the values around the mean. They are also called measures of dispersion, because they indicate the degree of diversity in a dataset — how much difference exists within the group. The two common measures of dispersion are the range and the standard deviation.

Note

Think and reflect: out of the mean and the median, which is more sensitive to outliers? The mean — every value enters its calculation, so one extreme value pulls the average towards itself. The median depends only on the middle position of the sorted data, so an extreme value at either end barely moves it.

(A) Range

The range is the difference between the maximum and minimum values of the data — the largest value minus the smallest. If M is the maximum value and S is the minimum, then

Range = M − S (Maximum − Minimum)

The range can be calculated only for numerical data. As a measure of dispersion it tells you about the coverage or spread of the data values — for example, the difference in salaries of employees, in the marks of a student, or in the prices of toys.

Its weakness follows directly from its definition: since the range is computed from just the two extreme values, any outlier in the data badly influences the result. One abnormally large or small value changes the range completely, even if every other value is tightly clustered.

Example 5.4. In the height data of the previous section, the minimum height is 85 cm and the maximum is 115 cm. The range is therefore 115 − 85 = 30 cm.

(B) Standard deviation

The standard deviation measures the differences within a group or set of data of a variable. Like the range, it measures the spread of the data — but where the range looks at only the two extreme values, the standard deviation takes all the given data into account.

It is calculated as the positive square root of the average of the squared differences of each value from the mean. In steps: subtract the mean from each value, square each of these differences, average the squares over all n values, and take the positive square root of that average. Given n values x₁, x₂, x₃, …, xₙ with mean x̄, the standard deviation — written σ, the Greek letter sigma — is

σ = √( Σ(xᵢ − x̄)² / n )

Reading the result is straightforward: a smaller standard deviation means the data are less spread out (values sit close to the mean), while a larger standard deviation means the data are more spread.

Example 5.5. Compute the standard deviation of the heights of the nine students used earlier. Their mean x̄ was 101.33 cm. Subtract the mean from each value, square the result, and tabulate:

Height x (cm)x − x̄(x − x̄)²
90−11.33128.37
1020.670.36
1108.6775.17
11513.67186.87
85−16.33266.67
90−11.33128.37
100−1.331.77
1108.6775.17
1108.6775.17

Here n = 9 and x̄ = 101.33. The plain deviations (x − x̄) cancel out to practically zero — which is exactly why they are squared before averaging. The squared deviations total 938.00. Dividing by the number of values gives 938 / 9 = 104.22, and taking the positive square root gives

σ = √104.22 = 10.2 cm

Choosing the right technique

The point of learning these techniques is knowing which one to apply to reach a decision. Match each problem below with the suitable statistical method (mean / median / mode / range / standard deviation):

Problem statementSuitable technique — and why
The management of a company wants to know about the disparity in salaries of all employeesStandard deviation — disparity is variation within the group, and it should reflect every salary, not just the two extremes
A teacher wants to know the average performance of the whole class in a testMean — the average of all the marks
Compare the height of residents of two citiesMean — one representative height per city, then compare
Find the dominant value from a set of valuesMode — the value with the highest frequency of occurrence
Table 5.3Table 5.3 Standard deviation of attendance of 9 students
Height (x) in cmx − x̄(x − x̄)²
90-11.33128.37
1020.670.36
1108.6775.17
11513.67186.87
85-16.33266.67
90-11.33128.37
100-1.331.77
1108.6775.17
1108.6775.17