Mathematics · Ch 13 — Statistics
Variance and Standard Deviation (Ungrouped Data)
Variance and Standard Deviation (Ungrouped Data)
Variance is another way of averaging the deviations of the observations from the mean while avoiding the identical-cancellation problem noted in the previous section — but instead of taking the absolute value of each deviation, variance takes its square.
Why square the deviations, rather than use the absolute value again? Squaring, like taking the absolute value, destroys the sign of each deviation (since ), so squared deviations never cancel to zero on averaging, exactly like absolute deviations. But the square has two further advantages that make it the preferred choice throughout higher statistics and probability: (i) it is a smooth, algebraically manipulable function (unlike , which has a sharp corner at zero and is awkward to differentiate or combine algebraically), which becomes essential once statistics is connected to calculus and probability theory later in the course; and (ii) squaring gives proportionately more weight to larger deviations — an observation twice as far from the mean contributes four times as much to the variance, which is often a desirable property when large deviations are considered especially significant.
Variance, for ungrouped data. If has mean , the variance (often denoted or ) is
Standard deviation. Because variance is expressed in the square of the original unit (e.g. marks, or kg), it is not directly comparable to the data itself. Taking the positive square root restores the original unit, giving the standard deviation:
Standard deviation is the single most widely used measure of dispersion in statistics.
A useful computational shortcut. Expanding the square in the variance formula gives an equivalent, often faster, way to compute it directly from and , without first subtracting the mean from every single observation:
…