Imagine you're a teacher with two classes that both scored an average of 75 on a test. In Class A, every student scored between 73 and 77 — almost everyone got nearly the same mark. In Class B, half the students scored 95 and the other half scored 55. The average is still 75, but the two classes are completely different.
The average tells you the center of the data. But it says nothing about how spread out the numbers are. That's where standard deviation comes in.
Standard deviation measures how far the typical data point lies from the mean. A small SD means the data is tightly clustered around the average. A large SD means the data is widely scattered.
The Core Idea in One Sentence
Important
Standard deviation is the average distance of each data point from the mean, but calculated in a special way that gives more weight to points that are far away.
Why Not Just Use the Simple Average Distance?
You might think: why not just take each data point, subtract the mean, and average those distances? That's a natural idea — it's called the mean absolute deviation. But it has a mathematical problem: the positive and negative deviations cancel out. For any dataset, the sum of (value − mean) is always zero.
So we square the deviations first. Squaring makes everything positive, and it also penalizes extreme values more heavily — a point 10 units away contributes 100 to the sum, while a point 2 units away contributes only 4. This is often desirable because outliers matter.
The Precise Definition
For a set of n values x1,x2,…,xn with mean xˉ:
SD=n∑i=1n(xi−xˉ)2
Let's break that down step by step:
Find the meanxˉ of all values.
Subtract the mean from each value to get the deviation xi−xˉ.
Square each deviation to make it positive.
Average these squared deviations — this gives the variance.
Take the square root to return to the original units.
The square root at the end is crucial. If your data is in marks, the variance is in "squared marks" — meaningless. Taking the square root brings the measure back to marks.
A Worked Example
Take the numbers: 2, 4, 6, 8, 10.
Step 1: Mean xˉ=52+4+6+8+10=6.
Step 2 & 3: Compute squared deviations:
(2−6)2=16
(4−6)2=4
(6−6)2=0
(8−6)2=4
(10−6)2=16
Step 4: Variance =516+4+0+4+16=540=8.
Step 5: SD =8≈2.83.
Interpretation: the typical data point is about 2.83 units away from the mean of 6.
Population vs. Sample
There's a subtle but important variation. When your data is the entire population (all students in a school, all planets in the solar system), you divide by n.
When your data is a sample from a larger population (a random selection of 100 voters), you divide by n−1 instead. This is called Bessel's correction, and it gives a slightly larger SD that better estimates the true population SD.
Note
For exam purposes: if the problem says "population" or "all", use n. If it says "sample" or "a random sample", use n−1.
What SD Tells You (and Doesn't)
A small SD doesn't mean the data is "good" — it means the data is consistent. A large SD doesn't mean "bad" — it means the data is spread out.
In many real-world contexts, SD helps you make decisions. If a machine produces parts with a small SD in diameter, it's reliable. If stock returns have a large SD, they're volatile (risky).
The mean gives you the center. The SD gives you the spread. Together, they describe the entire distribution in two numbers — which is remarkably powerful for such a simple idea.
For 9,3,8,8,9,8,9,18 the variance is 15 and the standard deviation is 15≈3.87.
σ2=n∑(xi−xˉ)2, σ=σ2
n=8, ∑xi=72⇒xˉ=9
(xi−9)2: 0,36,1,1,0,1,0,81⇒ sum =120
σ2=120/8=15⇒σ=15≈3.87
✓Final answer
Standard deviation ≈3.87
For the data set 9,3,8,8,9,8,9,18, the variance is 15 and the standard deviation is 15≈3.87.