Skip to content

Mathematics · Ch 13 — Statistics

Introduction to Dispersion

1

Introduction to Dispersion

Dispersion (also called variability or spread) measures how scattered or spread out the observations of a data set are around a central value such as the mean or median. Two data sets can have exactly the same average and yet be very different in how consistent their individual values are — averages alone hide this difference, which is exactly why measures of dispersion are needed alongside measures of central tendency.

A motivating example. Consider the marks (out of 1010) of two students across 55 tests:

Student P:5,5,5,5,5Student Q:0,3,5,7,10\text{Student } P: \quad 5, 5, 5, 5, 5 \qquad\qquad \text{Student } Q: \quad 0, 3, 5, 7, 10

Both have the same mean, 55. Yet Student PP's performance is perfectly consistent test after test, while Student QQ's marks swing wildly from 00 to 1010. The mean alone cannot tell these two situations apart — a measure of dispersion is needed to capture how spread out the values are around that common mean of 55.

This chapter studies four measures of dispersion, in increasing order of how much information about the data they use:

  • Range — the simplest measure, using only the largest and smallest observations.
  • Mean deviation — the average of the absolute distances of every observation from a central value (the mean or the median).
  • Variance — the average of the squared distances of every observation from the mean.
  • Standard deviation — the (positive) square root of the variance, bringing the unit of measurement back in line with the original data.

Each of range, mean deviation, variance, and standard deviation is developed below first for ungrouped (raw) data — where every individual observation is listed separately — and then extended to grouped (frequency) data, both discrete frequency distributions (a value xix_i repeated fif_i times) and continuous frequency distributions (data organised into class intervals with a frequency for each class). A shortcut (step-deviation) method for grouped data closes the chapter, making these computations manageable even for large, inconveniently-valued data sets.