Skip to content

Mathematics · Ch 7 — Statistics

Mean Deviation for Grouped Data

7.4.2

Mean Deviation for Grouped Data

Mean Deviation for Grouped Data

Data can be grouped in two ways: as a discrete frequency distribution (where each observation is a distinct value with its own frequency) or as a continuous frequency distribution (where data is divided into class intervals). The method for finding mean deviation differs slightly between the two, but the core idea remains the same — we measure the average absolute distance of the data from a central value.

(a) Discrete Frequency Distribution

Suppose we have nn distinct values x1,x2,…,xnx_1, x_2, \dots, x_n occurring with frequencies f1,f2,…,fnf_1, f_2, \dots, f_n respectively. The data is arranged as:

xxx1x_1x2x_2x3x_3…\dotsxnx_n
fff1f_1f2f_2f3f_3…\dotsfnf_n

Here N=∑i=1nfiN = \sum_{i=1}^{n} f_i is the total number of observations.

(i) Mean Deviation About the Mean

Step 1 — Find the mean.

The mean xˉ\bar{x} of the grouped data is:

xˉ=∑i=1nfixi∑i=1nfi=1N∑i=1nfixi\bar{x} = \frac{\sum_{i=1}^{n} f_i x_i}{\sum_{i=1}^{n} f_i} = \frac{1}{N} \sum_{i=1}^{n} f_i x_i

Step 2 — Find absolute deviations from the mean.

For each observation xix_i, compute ∣xi−xˉ∣|x_i - \bar{x}|.

Step 3 — Find the weighted mean of these absolute deviations.

Multiply each absolute deviation by its frequency, sum them up, and divide by NN:

M.D.(xˉ)=∑i=1nfi∣xi−xˉ∣∑i=1nfi=1N∑i=1nfi∣xi−xˉ∣\text{M.D.}(\bar{x}) = \frac{\sum_{i=1}^{n} f_i |x_i - \bar{x}|}{\sum_{i=1}^{n} f_i} = \frac{1}{N} \sum_{i=1}^{n} f_i |x_i - \bar{x}|

Watch out

The deviations are taken in absolute value — signs are ignored. If you sum the signed deviations fi(xi−xˉ)f_i(x_i - \bar{x}), you will always get zero. That is why we use absolute values.

Example 4 — Find the mean deviation about the mean for:

xix_i25681012
fif_i2810785

Solution. First compute the mean:

xˉ=∑fixiN=2(2)+5(8)+6(10)+8(7)+10(8)+12(5)2+8+10+7+8+5=4+40+60+56+80+6040=30040=7.5\bar{x} = \frac{\sum f_i x_i}{N} = \frac{2(2) + 5(8) + 6(10) + 8(7) + 10(8) + 12(5)}{2+8+10+7+8+5} = \frac{4+40+60+56+80+60}{40} = \frac{300}{40} = 7.5

Now compute ∣xi−xˉ∣|x_i - \bar{x}| and fi∣xi−xˉ∣f_i|x_i - \bar{x}|:

xix_ifif_ifixif_i x_i∣xi−7.5∣\lvert x_i - 7.5 \rvertfi∣xi−7.5∣f_i\lvert x_i - 7.5 \rvert
2245.511
58402.520
610601.515
87560.53.5
108802.520
125604.522.5
Total4030092

M.D.(xˉ)=1N∑fi∣xi−xˉ∣=140×92=2.3\text{M.D.}(\bar{x}) = \frac{1}{N} \sum f_i |x_i - \bar{x}| = \frac{1}{40} \times 92 = 2.3

(ii) Mean Deviation About the Median

Step 1 — Find the median.

Arrange the observations in ascending order. Compute cumulative frequencies. The median is the value of the observation whose cumulative frequency is equal to or just greater than N/2N/2.

Note

If NN is even, the median is the average of the N2\frac{N}{2}-th and (N2+1)\left(\frac{N}{2}+1\right)-th observations. Both lie in the same class when the cumulative frequency first reaches or exceeds N/2N/2.

Step 2 — Find absolute deviations from the median.

For each xix_i, compute ∣xi−M∣|x_i - M|, where MM is the median.

Step 3 — Find the weighted mean of these absolute deviations.

M.D.(M)=∑i=1nfi∣xi−M∣∑i=1nfi=1N∑i=1nfi∣xi−M∣\text{M.D.}(M) = \frac{\sum_{i=1}^{n} f_i |x_i - M|}{\sum_{i=1}^{n} f_i} = \frac{1}{N} \sum_{i=1}^{n} f_i |x_i - M|

Example 5 — Find the mean deviation about the median for:

xix_i3691213152122
fif_i34524543

Solution. The data is already in ascending order. Compute cumulative frequencies:

xix_i3691213152122
fif_i34524543
c.f.37121418232730

N=30N = 30, so N/2=15N/2 = 15. The 15th and 16th observations both lie in the cumulative frequency 18, which corresponds to xi=13x_i = 13. Therefore:

M=13+132=13M = \frac{13 + 13}{2} = 13

Now compute ∣xi−M∣|x_i - M| and fi∣xi−M∣f_i|x_i - M|:

xix_i∣xi−13∣\lvert x_i - 13 \rvertfif_ifi∣xi−13∣f_i\lvert x_i - 13 \rvert
310330
67428
94520
12122
13040
152510
218432
229327
Total30149

M.D.(M)=130×149=4.97\text{M.D.}(M) = \frac{1}{30} \times 149 = 4.97


(b) Continuous Frequency Distribution

In a continuous frequency distribution, data is grouped into class intervals without gaps. For example:

Marks0-1010-2020-3030-4040-5050-60
Students12182720176

The key assumption: the frequency in each class is centred at its mid-point. So we first find the mid-point xix_i of each class, then proceed exactly as for a discrete frequency distribution.

(i) Mean Deviation About the Mean

Step 1 — Find the mid-point xix_i of each class.

For a class a−ba-b, the mid-point is a+b2\frac{a+b}{2}.

Step 2 — Compute the mean xˉ\bar{x}.

Use the same formula as for discrete data:

xˉ=∑fixi∑fi=1N∑fixi\bar{x} = \frac{\sum f_i x_i}{\sum f_i} = \frac{1}{N} \sum f_i x_i

Step 3 — Compute absolute deviations ∣xi−xˉ∣|x_i - \bar{x}| and their weighted sum.

Then:

M.D.(xˉ)=1N∑fi∣xi−xˉ∣\text{M.D.}(\bar{x}) = \frac{1}{N} \sum f_i |x_i - \bar{x}|

Example 6 — Find the mean deviation about the mean for:

Marks10-2020-3030-4040-5050-6060-7070-80
Students23814832

Solution. Compute mid-points and the required columns:

Classfif_ixix_ifixif_i x_i∣xi−45∣\lvert x_i - 45 \rvertfi∣xi−45∣f_i\lvert x_i - 45 \rvert
10-20215303060
20-30325752060
30-408352801080
40-50144563000
50-608554401080
60-703651952060
70-802751503060
Total401800400

xˉ=180040=45\bar{x} = \frac{1800}{40} = 45

M.D.(xˉ)=140×400=10\text{M.D.}(\bar{x}) = \frac{1}{40} \times 400 = 10

Shortcut Method (Step-Deviation Method) for Mean Deviation About the Mean

When the mid-points are large, we can simplify the calculation of xˉ\bar{x} using step deviations. This involves shifting the origin (choosing an assumed mean aa) and changing the scale (dividing by a common factor hh).

Define a new variable:

di=xi−ahd_i = \frac{x_i - a}{h}

where aa is the assumed mean (usually a mid-point near the centre of the data) and hh is the common factor (the class width, if all classes have equal width).

Then the mean is:

xˉ=a+h⋅∑fidiN\bar{x} = a + h \cdot \frac{\sum f_i d_i}{N}

Tip

The step-deviation method only simplifies the calculation of xˉ\bar{x}. The rest of the procedure — finding ∣xi−xˉ∣|x_i - \bar{x}| and then M.D.(xˉ)\text{M.D.}(\bar{x}) — remains exactly the same. You still need to compute the actual absolute deviations from the actual mean.

Example 6 using step-deviation method.

Take a=45a = 45 (the mid-point of the 40-50 class) and h=10h = 10 (the class width).

Classfif_ixix_idi=xi−4510d_i = \frac{x_i - 45}{10}fidif_i d_i∣xi−45∣\lvert x_i - 45 \rvertfi∣xi−45∣f_i\lvert x_i - 45 \rvert
10-20215-3-63060
20-30325-2-62060
30-40835-1-81080
40-5014450000
50-60855181080
60-70365262060
70-80275363060
Total400400

xˉ=45+10×040=45\bar{x} = 45 + 10 \times \frac{0}{40} = 45

M.D.(xˉ)=40040=10\text{M.D.}(\bar{x}) = \frac{400}{40} = 10

The result is identical, as expected.

(ii) Mean Deviation About the Median

The process is similar to the discrete case, but we must first find the median using the formula for continuous data.

Step 1 — Find the median class.

Arrange the data in ascending order (the classes are already in order). Compute cumulative frequencies. The median class is the class interval whose cumulative frequency is just greater than or equal to N/2N/2.

Step 2 — Apply the median formula.

Median=l+N2−Cf×h\text{Median} = l + \frac{\frac{N}{2} - C}{f} \times h

where:

  • ll = lower limit of the median class
  • ff = frequency of the median class
  • hh = width of the median class
  • CC = cumulative frequency of the class just preceding the median class
  • NN = total frequency …
Figure 13.3Shifting the origin to the assumed mean (step-deviation method)
Fig. 13.3 — Shifting the origin to the assumed mean (step-deviation method)

Drawn by us to help you understand the concept clearly, and verified to make sure it's accurate. For exams, practice from your textbook's own diagram.

Fig. 13.3 is a visual of the core idea behind the step-deviation method: you can shift the origin of your number line to a convenient assumed mean, work with smaller numbers, and then shift back at the end.

The figure shows two aligned number lines, one above the other. The bottom line is the original data scale, running from 0 to 120. The top line is the deviation scale, running from −60 to 60. Each tick mark on the bottom line has a corresponding tick directly above it on the top line. The key alignment is at the value 60 on the bottom line: directly above it, on the top line, sits the deviation 0. An upward arrow labelled "Assumed Mean" points to 60 on the bottom line.

What this teaches is that subtracting the assumed mean aa from every observation xix_i is equivalent to sliding the entire number line leftwards so that the origin (0 on the deviation line) now sits at the assumed mean. The deviations xi−ax_i - a are the new coordinates on the shifted line. If all these deviations also share a common factor hh (like all being multiples of 10), you can further scale them down by dividing by hh, giving the step-deviations di=xi−ahd_i = \frac{x_i - a}{h}. This is a change of scale on the number line, shown separately in Fig. 13.4.

The textbook uses this figure to introduce the step-deviation formula for the mean:

xˉ=a+h⋅1N∑i=1nfidi\bar{x} = a + h \cdot \frac{1}{N} \sum_{i=1}^{n} f_i d_i

Here, aa is the assumed mean (the point to which you shift the origin), hh is the common factor (the scale divisor), di=xi−ahd_i = \frac{x_i - a}{h} are the step-deviations, fif_i are the frequencies, and N=∑fiN = \sum f_i is the total frequency. The formula says: compute the mean of the step-deviations, multiply by hh to undo the scaling, then add aa to undo the shift. The figure makes it clear that this is just a two-step translation of the number line — first a shift, then a rescaling — and that the mean follows the same transformation. …

Figure 13.4Change of scale to step-deviations
Fig. 13.4 — Change of scale to step-deviations

Drawn by us to help you understand the concept clearly, and verified to make sure it's accurate. For exams, practice from your textbook's own diagram.

Fig. 13.4 is a visual explanation of the step-deviation method — a shortcut for calculating the mean of grouped data when the class intervals are equal. The figure shows three horizontal number lines, one above the other, all aligned vertically so that the same tick marks line up.

The bottom number line is the original raw scale, running from 0 to 120. This is the scale on which the actual data (mid-points of classes) live. The middle number line shows "deviations from assumed mean" — it runs from −60 to +60. Each tick on this line corresponds to the difference between a raw value and the assumed mean, which is chosen as 60. The indigo dots on each tick mark the actual deviation values. An upward arrow labelled "Assumed Mean" points to the 60 mark on the bottom scale, making it clear that the origin of the middle line has been shifted to that point.

The top number line shows "step deviations" — it runs from −6 to +6. This is the result of dividing every deviation from the middle line by a common factor (here, 10). So the step deviation is simply the deviation divided by the class width.

The key idea the figure teaches is that you can simplify calculations by first shifting the origin (subtracting an assumed mean) and then changing the scale (dividing by the common factor). This transforms the original mid-points xix_i into step deviations did_i using:

di=xi−ahd_i = \frac{x_i - a}{h}

where aa is the assumed mean (60 in the figure) and hh is the common factor (10 in the figure). The mean of the original data can then be recovered by reversing both transformations:

xˉ=a+h⋅∑fidiN\bar{x} = a + h \cdot \frac{\sum f_i d_i}{N}

Here, N=∑fiN = \sum f_i is the total frequency. The textbook uses this method to compute the mean for the data in Example 6 (marks of students), where a=45a = 45 and h=10h = 10, and then finds the mean deviation about the mean using the usual formula:

M.D.(xˉ)=1N∑fi∣xi−xˉ∣\text{M.D.}(\bar{x}) = \frac{1}{N} \sum f_i |x_i - \bar{x}|

Note

The step-deviation method does not change the value of the mean deviation itself — it only simplifies the calculation of xˉ\bar{x}. Once xˉ\bar{x} is found, the deviations ∣xi−xˉ∣|x_i - \bar{x}| are computed directly from the original mid-points, not from the step deviations.

Watch out

A common mistake is to think that the step deviations did_i can be used directly in the mean deviation formula. They cannot — the mean deviation requires absolute differences from the actual mean xˉ\bar{x}, not from the assumed mean. The step-deviation trick works only for computing xˉ\bar{x} itself. …

Table 13.1Mean deviation about the mean (discrete distribution: x=2..12)
xix_ifif_ifixif_ix_i∣xi−xˉ∣\lvert x_i-\bar{x} \rvertfi∣xi−xˉ∣f_i\lvert x_i-\bar{x} \rvert
2245.511
58402.520
610601.515
87560.53.5
108802.520
Table 13.2Discrete data with cumulative frequency row
xix_i3691213152122
fif_i34524543
Table 13.3Absolute deviations from the median
∣xi−M∣\lvert x_i-M \rvert107410289
fif_i34524543
Table 13.4Mean deviation about the mean for grouped (continuous) data
Marks obtainedNumber of students fif_iMid-points xix_ifixif_ix_i∣xi−xˉ∣\lvert x_i-\bar{x} \rvertfi∣xi−xˉ∣f_i\lvert x_i-\bar{x} \rvert
10-20215303060
20-30325752060
30-408352801080
40-50144563000
50-608554401080
Table 13.5Mean deviation by the step-deviation (shortcut) method
Marks obtainedNumber of students fif_iMid-points xix_idi=xi−4510d_i=\dfrac{x_i-45}{10}fidif_id_i∣xi−xˉ∣\lvert x_i-\bar{x} \rvertfi∣xi−xˉ∣f_i\lvert x_i-\bar{x} \rvert
10-20215−3-3−6-63060
20-30325−2-2−6-62060
30-40835−1-1−8-81080
40-5014450000
50-60855181080
Table 13.6Mean deviation about the median for grouped (continuous) data
ClassFrequency fif_iCumulative frequency c.f.c.f.Mid-points xix_i∣xi−Med.∣\lvert x_i-\text{Med.} \rvertfi∣xi−Med.∣f_i\lvert x_i-\text{Med.} \rvert
0-1066523138
10-20713151391
20-30152825345
30-401644357112