Skip to content

Informatics Practices · Ch 5 — Understanding Data

Measures of Central Tendency

5.5.1

Measures of Central Tendency

A dataset with dozens or hundreds of entries cannot be judged by staring at each value individually. A measure of central tendency solves this by condensing the whole collection into a single value that gives an idea about the data. The three most common measures are the mean, the median and the mode — and each answers a different question. Rather than examining every data value, computing these three tells you, respectively, the average of the data, its middle value, and the value that occurs most frequently. Which of the three is the suitable choice is not arbitrary: it depends on certain characteristics of the data itself, as the discussion below makes clear.

(A) Mean

The mean is simply the average of the numeric values of an attribute — the two words, mean and average, name the same thing. Suppose weight data exist for 40 students of a class. There is no need to read all 40 entries; calculating the average immediately gives an idea of the typical weight of a student in that class.

Given n values x₁, x₂, x₃, …, xₙ, the mean is computed by adding all the values and dividing by how many there are:

x̄ = (x₁ + x₂ + x₃ + … + xₙ) / n

Example 5.1. The heights (in cm) of the students in a class are [90, 102, 110, 115, 85, 90, 100, 110, 110]. The mean height of the class is

(90 + 102 + 110 + 115 + 85 + 90 + 100 + 110 + 110) / 9 = 912 / 9 = 101.33 cm

The mean has one significant weakness: it is not a suitable choice when the data contain outliers. An outlier is an exceptionally large or exceptionally small value in comparison to the other values of the data. Usually outliers are treated as errors, because they can influence the average — and any other statistical calculation based on the data. If outliers are present, they should be removed first, and the mean then calculated from the remaining data.

Watch out

A single extreme value can drag the mean far away from where most of the data actually lie. Before trusting an average, check the data for outliers — an unusually huge or tiny value distorts the mean even though every other value is ordinary.

(B) Median

The median, like the mean, is computed for a single attribute or variable at a time. The procedure is different: first sort all the values in ascending or descending order; the value that then sits in the middle is the median.

Two cases arise, depending on how many values there are:

  • Odd number of values — the median is the value at the exact middle position of the sorted list.
  • Even number of values — there is no single middle position, so the median is the average of the two middle values.

What the median represents is the actual central value at which the given data is divided into two equal parts — half the values lie on one side of it and half on the other.

Example 5.2. Take the same height data used for the mean. Sorting it in ascending order gives [85, 90, 90, 100, 102, 110, 110, 110, 115]. There are 9 values in total — an odd number — so the median is the value at position 5, which is 102 cm. Notice that position 5 is the middle whether you count from left to right or from right to left; the median splits the data evenly either way.

(C) Mode

The mode is the value that appears the most number of times in the given data of an attribute or variable. It is computed from the frequency of occurrence of each distinct value in the data — count how often every distinct value appears, and the one with the highest count is the mode.

Three points about the mode are worth fixing clearly:

  • A dataset has no mode if every value occurs exactly once — no value is more frequent than any other.
  • A dataset can have multiple modes if more than one value shares the same highest frequency.
  • The mode can be found for non-numeric data as well as numeric data — unlike the mean and the median, which need numbers, the mode only needs counting. The most popular car colour in a survey, for instance, is a mode.

Example 5.3. In the height list [90, 102, 110, 115, 85, 90, 100, 110, 110], the value 110 occurs 3 times — a higher frequency than any other value. The mode is therefore 110.

The three measures side by side

| Measure | What it gives | How it is found | Points to remember |

|---|---|---|---| …