Skip to content

Informatics Practices · Ch 5 — Understanding Data

Statistical Techniques for Data Processing

5.5

Statistical Techniques for Data Processing

A pile of raw data values, however carefully collected and stored, answers no questions on its own. To turn those values into information, they must be processed — and the very first stage of processing is building a preliminary understanding of the data: what a typical value looks like, how the values behave as a group, and whether anything unusual is hiding inside them. Statistics provides a set of standard techniques for exactly this purpose.

The idea common to all of these techniques is summarisation. Instead of reading through every individual entry, summarisation methods are applied to tabular data — data arranged in rows and columns — so that the whole collection can be comprehended easily through a handful of computed values. A teacher with the marks of an entire class, or a company with the salaries of all its employees, does not need to inspect each record one by one; a few well-chosen summary figures convey the essential character of the data at a glance.

The commonly used statistical techniques for data summarisation fall into two groups, taken up in the subsections that follow:

  • Measures of central tendency — a single value that represents the data as a whole. The three most common are the mean (the average), the median (the middle value when the data are sorted) and the mode (the value that occurs most frequently).
  • Measures of variability — values that describe how spread out the data are around the centre. The common ones are the range (the gap between the largest and smallest values) and the standard deviation (a spread measure that takes every value into account). …