Skip to content

Mathematics · Ch 13 — Statistics

Variance and Standard Deviation

13.5

Variance and Standard Deviation

Why Squaring the Deviations?

When we calculated the mean deviation, we took absolute values of the deviations to prevent positive and negative deviations from cancelling each other out. But absolute values are mathematically awkward — they are not differentiable at zero and don't behave nicely in further calculations.

A much cleaner approach is to square each deviation. Since a square is always non-negative, the sign problem vanishes automatically. If we have nn observations x1,x2,…,xnx_1, x_2, \dots, x_n with mean xˉ\bar{x}, then the sum

∑i=1n(xi−xˉ)2\sum_{i=1}^{n} (x_i - \bar{x})^2

is always zero or positive. If this sum is exactly zero, then every single deviation must be zero — meaning all observations are identical to the mean, and there is no dispersion at all. If the sum is small, the observations cluster close to the mean (low dispersion). If it is large, the observations are spread far from the mean (high dispersion).

So far, this sum seems like a reasonable measure of scatter. But there is a serious flaw, as the textbook demonstrates with two contrasting data sets.


The Flaw in Using the Sum of Squared Deviations Directly

Consider Set A: six observations — 5, 15, 25, 35, 45, 55. Their mean is xˉ=30\bar{x} = 30. The sum of squared deviations is

∑i=16(xi−30)2=(5−30)2+(15−30)2+(25−30)2+(35−30)2+(45−30)2+(55−30)2\sum_{i=1}^{6} (x_i - 30)^2 = (5-30)^2 + (15-30)^2 + (25-30)^2 + (35-30)^2 + (45-30)^2 + (55-30)^2

=625+225+25+25+225+625=1750= 625 + 225 + 25 + 25 + 225 + 625 = 1750

Now consider Set B: 31 observations — 15, 16, 17, …, 45. Their mean is also yˉ=30\bar{y} = 30. The sum of squared deviations is

∑i=131(yi−30)2=(15−30)2+(16−30)2+⋯+(45−30)2\sum_{i=1}^{31} (y_i - 30)^2 = (15-30)^2 + (16-30)^2 + \cdots + (45-30)^2

=(−15)2+(−14)2+⋯+(−1)2+02+12+⋯+142+152= (-15)^2 + (-14)^2 + \cdots + (-1)^2 + 0^2 + 1^2 + \cdots + 14^2 + 15^2

This is twice the sum of squares of the first 15 natural numbers (since 020^2 contributes nothing):

=2×15(15+1)(2×15+1)6=2×15×16×316=2×5×16×31=2480= 2 \times \frac{15(15+1)(2\times 15 + 1)}{6} = 2 \times \frac{15 \times 16 \times 31}{6} = 2 \times 5 \times 16 \times 31 = 2480

If we judged dispersion by the raw sum of squared deviations, we would say Set A (sum = 1750) has less dispersion than Set B (sum = 2480). But look at the actual spread: Set A's deviations range from −25-25 to +25+25, while Set B's deviations only range from −15-15 to +15+15. The observations in Set A are clearly more scattered from the mean. The sum of squared deviations is misleading here because it depends on the number of observations — Set B has many more terms, so its sum is larger even though each individual deviation is smaller.

Watch out

The raw sum ∑(xi−xˉ)2\sum (x_i - \bar{x})^2 is not a proper measure of dispersion because it grows with the sample size. A set with many tightly clustered observations can have a larger sum than a set with fewer, more scattered observations.


The Solution: Variance — the Mean of the Squared Deviations

To remove the dependence on the number of observations, we take the average of the squared deviations. This quantity is called the variance, denoted by σ2\sigma^2 (sigma squared).

For Set A, the mean of the squared deviations is

16×1750=291.67\frac{1}{6} \times 1750 = 291.67

For Set B, it is

131×2480=80\frac{1}{31} \times 2480 = 80 …

Figure 13.5Set A (5,15,25,35,45,55) about its mean 30 (large scatter)
Fig. 13.5 — Set A (5,15,25,35,45,55) about its mean 30 (large scatter)

Drawn by us to help you understand the concept clearly, and verified to make sure it's accurate. For exams, practice from your textbook's own diagram.

Fig. 13.5 is a simple number-line diagram that makes a subtle but critical point about how we measure scatter. The line runs from 0 to 60, with tick marks every 5 units. Six dots sit at 5, 15, 25, 35, 45, and 55 — evenly spaced, far apart. An upward arrow labelled "Mean" points to 30, right in the centre. The caption calls this "large scatter", and the visual is unmistakable: the points are spread over a range of 50 units, from 5 to 55, with deviations as large as –25 and +25 from the mean.

The textbook uses this figure to expose a flaw in a naive measure of dispersion. If you simply add up the squared deviations from the mean — ∑(xi−xˉ)2\sum (x_i - \bar{x})^2 — you get 1750 for this set of six numbers. That sum is large in absolute terms. But compare it to Set B (31 observations from 15 to 45, also centred at 30), whose sum of squared deviations is 2480 — even larger. If you judged dispersion by the raw sum alone, you'd wrongly conclude that Set B is more spread out. The diagram shows the opposite: the six points in Fig. 13.5 are clearly more scattered than the tightly packed 31 points of Set B (shown in a separate figure, Fig. 13.6). The raw sum is misleading because it depends on the number of observations, not just the spread.

The solution the textbook develops is to take the mean of the squared deviations — the variance. For the six points in Fig. 13.5:

σ2=16∑i=16(xi−xˉ)2=17506≈291.67\sigma^2 = \frac{1}{6} \sum_{i=1}^{6} (x_i - \bar{x})^2 = \frac{1750}{6} \approx 291.67

For Set B, the variance is 2480/31≈802480/31 \approx 80. Now the numbers match the visual: the variance for the widely spaced six points is much larger, correctly reflecting the "large scatter" the figure shows.

σ2=1n∑i=1n(xi−xˉ)2\sigma^2 = \frac{1}{n} \sum_{i=1}^{n} (x_i - \bar{x})^2

where σ2\sigma^2 is the variance, nn is the number of observations, xix_i are the individual observations, and xˉ\bar{x} is their mean. …

Figure 13.6Set B (31 integers 15..45) about its mean 30 (small scatter)
Fig. 13.6 — Set B (31 integers 15..45) about its mean 30 (small scatter)

Drawn by us to help you understand the concept clearly, and verified to make sure it's accurate. For exams, practice from your textbook's own diagram.

What Fig. 13.6 Shows

The figure presents a number line from 0 to 60. Thirty-one indigo dots sit at every integer from 15 to 45 — a dense, continuous block of points. An upward arrow labelled "Mean" points to the value 30, which is the centre of this block. The visual message is immediate: the observations are packed tightly around the mean, with the farthest points only 15 units away on either side.

This is the companion to Fig. 13.5, which showed Set A — six widely scattered points (5, 15, 25, 35, 45, 55) also centred at 30. Together, the two figures make a single pedagogical point: raw scatter looks different from what the sum of squared deviations alone suggests.

The Physical Idea

The textbook has just introduced the sum of squared deviations from the mean, ∑i=1n(xi−xˉ)2\sum_{i=1}^n (x_i - \bar{x})^2, as a candidate measure of dispersion. For Set A (six points), this sum is 1750. For Set B (31 points), it is 2480. If you used only the raw sum, you would conclude that Set B has more dispersion — yet the figure shows the opposite: Set B's points are far more concentrated near the mean.

The problem is that the raw sum grows with the number of observations. A larger dataset naturally produces a larger total, even if each individual point is closer to the mean. The figure makes this contradiction visually obvious: the dense cluster of 31 points clearly has less scatter than the sparse six-point set, but the raw sum says otherwise.

The Resolution: Variance

To correct for the number of observations, we take the mean of the squared deviations. This quantity is called the variance, denoted σ2\sigma^2:

σ2=1n∑i=1n(xi−xˉ)2\sigma^2 = \frac{1}{n} \sum_{i=1}^n (x_i - \bar{x})^2

For Set A: σ2=16×1750=291.67\sigma^2 = \frac{1}{6} \times 1750 = 291.67

For Set B: σ2=131×2480=80\sigma^2 = \frac{1}{31} \times 2480 = 80

Now the numbers match what the figure shows: Set A has nearly four times the variance of Set B, confirming that its points are indeed more scattered.

Watch out

Do not confuse the sum of squared deviations with the variance. The sum is useful only as an intermediate step; the variance divides by nn to give a per-observation measure. Without this division, comparing datasets of different sizes is meaningless.

What Each Symbol Means

  • xix_i — the ii-th observation (each integer dot on the number line)
  • xˉ\bar{x} — the arithmetic mean of all observations (30, marked by the arrow)
  • xi−xˉx_i - \bar{x} — the deviation of an observation from the mean (the horizontal distance from a dot to the arrow)
  • (xi−xˉ)2(x_i - \bar{x})^2 — the squared deviation (always non-negative, solving the sign problem that plagued mean deviation) …