Imagine you're a teacher with two classes that both scored an average of 75 on a test. In Class A, every student scored between 73 and 77 — almost everyone got nearly the same mark. In Class B, half the students scored 95 and the other half scored 55. The average is still 75, but the two classes are completely different.
The average tells you the center of the data. But it says nothing about how spread out the numbers are. That's where standard deviation comes in.
Standard deviation measures how far the typical data point lies from the mean. A small SD means the data is tightly clustered around the average. A large SD means the data is widely scattered.
The Core Idea in One Sentence
Important
Standard deviation is the average distance of each data point from the mean, but calculated in a special way that gives more weight to points that are far away.
Why Not Just Use the Simple Average Distance?
You might think: why not just take each data point, subtract the mean, and average those distances? That's a natural idea — it's called the mean absolute deviation. But it has a mathematical problem: the positive and negative deviations cancel out. For any dataset, the sum of (value − mean) is always zero.
So we square the deviations first. Squaring makes everything positive, and it also penalizes extreme values more heavily — a point 10 units away contributes 100 to the sum, while a point 2 units away contributes only 4. This is often desirable because outliers matter.
The Precise Definition
For a set of n values x1,x2,…,xn with mean xˉ:
SD=n∑i=1n(xi−xˉ)2
Let's break that down step by step:
Find the meanxˉ of all values.
Subtract the mean from each value to get the deviation xi−xˉ.
Square each deviation to make it positive.
Average these squared deviations — this gives the variance.
Take the square root to return to the original units.
The square root at the end is crucial. If your data is in marks, the variance is in "squared marks" — meaningless. Taking the square root brings the measure back to marks.