Skip to content

Mathematics · Ch 8 — Measures of Dispersion

Change of Origin and Scale

8.2.3

Change of Origin and Scale

Variance and standard deviation respond in specific, predictable ways when the data is shifted or rescaled, and this property is used to make hand computation easier, especially for grouped data with large or awkward mid-values.

  1. Change of origin only: define d=x−Ad = x - A, where A is a constant (commonly the class mark of the middle class — if the number of classes is odd, A is the mid-value of the single middle class; if even, A is the mid-value of whichever of the two middle classes has the greater frequency). Since every value is shifted by the same constant amount, the spread of the data does not change: σd2=σx2,σd=σx\sigma_d^2 = \sigma_x^2, \qquad \sigma_d = \sigma_x In other words, the standard deviation of x1−A,x2−A,…,xn−Ax_1-A, x_2-A, \ldots, x_n-A is exactly the same as the standard deviation of x1,x2,…,xnx_1, x_2, \ldots, x_n — variance and S.D. are independent of a change of origin.

  2. Change of scale: define u=x−Ahu = \dfrac{x-A}{h}, where h is the class width (or, when the data has no class intervals, h is the common difference between consecutive values of xix_i), with h≠0h\neq 0. Unlike a pure shift, dividing by h does change the spread proportionally: σx=h σu,σx2=h2 σu2\sigma_x = h\,\sigma_u, \qquad \sigma_x^2 = h^2\,\sigma_u^2 So the standard deviation of x1−Ah,x2−Ah,…,xn−Ah\frac{x_1-A}{h}, \frac{x_2-A}{h}, \ldots, \frac{x_n-A}{h} is exactly 1h\frac{1}{h} times the standard deviation of the original values — variance and S.D. are NOT independent of a change of scale; they scale with h.

In practice this substitution converts the original x (or class mid-value) column into a column of small, easy integers u (…, −2, −1, 0, 1, 2, …), whose variance σu2\sigma_u^2 is computed the ordinary way using Var(u)=σu2=∑fiui2N−uˉ2\text{Var}(u)=\sigma_u^2=\frac{\sum f_iu_i^2}{N}-\bar{u}^2, after which the answer is converted back to the original scale using σx=hσu\sigma_x = h\sigma_u and Var(X)=h2Var(u)\text{Var}(X)=h^2\text{Var}(u).

Worked Example (Ex.4): Values 15, 20, 25, 30, 35, 40, 45 with frequencies 13, 12, 15, 18, 17, 10, 15 (N=100). Using u=(x−30)/5u=(x-30)/5, the u-values run from −3 to 3. Building the f.u and f.u² columns gives ∑fiui=4\sum f_iu_i = 4 and ∑fiui2=372\sum f_iu_i^2 = 372, so uˉ=4/100=0.04\bar{u}=4/100=0.04 and Var(u)=372/100−0.042=3.72−0.0016=3.7184\text{Var}(u) = 372/100 - 0.04^2 = 3.72-0.0016=3.7184, giving σu≈1.928\sigma_u \approx 1.928. Converting back with h=5h=5: σx=5×1.928≈9.64\sigma_x = 5\times1.928 \approx 9.64.

Worked Example (Ex.5): A grouped distribution with 8 classes of width 10 from 45-55 to 115-125, frequencies 7, 20, 27, 23, 13, 6, 3, 1 (N=100), mid-values 50 to 120. Using u=(x−90)/10u=(x-90)/10, the f.u and f.u² columns give ∑fiui=−150\sum f_iu_i = -150 and ∑fiui2=450\sum f_iu_i^2 = 450, so uˉ=−150/100=−1.5\bar{u}=-150/100=-1.5 and Var(u)=450/100−(−1.5)2=4.5−2.25=2.25\text{Var}(u) = 450/100 - (-1.5)^2 = 4.5-2.25 = 2.25. Converting back with h=10h=10: Var(X)=h2Var(u)=100×2.25=225\text{Var}(X) = h^2\text{Var}(u) = 100\times2.25 = 225, so S.D.=σx=225=15\text{S.D.} = \sigma_x = \sqrt{225} = 15. …

Table Ex.4Variance/S.D. via change of scale, ungrouped values 15-45
Xuff.uf.u²
15−313−39117
20−212−2448
25−115−1515
3001800
351171717
402102040
Table Ex.5Variance/S.D. via change of scale, grouped C.I. 45-125
Class-intervalsMid value (xix_i)fif_iuiu_ifiuif_iu_ifiui2f_iu_i^2
45-55507−4−28112
55-656020−3−60180
65-757027−2−54108
75-858023−1−2323
85-959013000
95-1051006166
105-11511032612
Table Ex.6S.D. of heights of 500 plants via change of scale
ClassMid value (xix_i)fif_iuiu_ifiuif_iu_ifiui2f_iu_i^2
20-2522.5145−2−290580
25-3027.5125−1−125125
30-3532.590000
35-4037.54014040
40-4542.545290180