Abnormality Definition — First Principles
Imagine you are looking at a set of data points — say, the heights of all students in your class. Most heights cluster around some average value. A few students are taller, a few shorter, but almost everyone falls within a certain range. Now suppose one student is 8 feet tall. That student's height is abnormal — it lies far outside the pattern of the rest of the group.
That is the core intuition behind abnormality in statistics: an observation that deviates so much from the rest of the data that it stands out as unusual.
The Precise Definition
In formal terms, an abnormality (often called an outlier) is a data point that lies at an abnormal distance from other values in a random sample from a population. There is no single universal cutoff — it depends on the context — but the most common working definition in Indian exam syllabi (especially for statistics and probability) is:
An observation x is abnormal if ∣x−μ∣>3σ
where μ is the population mean and σ is the population standard deviation.
This is the 3-sigma rule. The reasoning: for a normal (bell-shaped) distribution, about 99.7% of all data lies within three standard deviations of the mean. Anything beyond that has less than a 0.3% chance of occurring naturally — so we call it abnormal.
Why 3-Sigma? The Intuition
The number 3 is not arbitrary. In a normal distribution:
- Within 1σ of μ: ~68% of data
- Within 2σ of μ: ~95% of data
- Within 3σ of μ: ~99.7% of data
A point beyond 3σ is so rare that it is more likely to be a mistake, a special cause, or a genuinely exceptional case — hence "abnormal."
The 3-sigma rule assumes the data follows a normal distribution. If the distribution is skewed or has heavy tails, this rule can misclassify normal points as abnormal. Always check the shape of your data first.
A Concrete Example
Suppose the mean height of Indian adult males is 165 cm with a standard deviation of 7 cm. A person who is 190 cm tall:
z=7190−165=725≈3.57 …