Skip to content

Informatics Practices · Ch 3 — Data Handling using Pandas – II

Estimating Missing Values

3.8.3

Estimating Missing Values

Estimating Missing Values

Missing values in a dataset are a loss of information. Instead of simply dropping those rows, you can fill them with an estimated or approximated value. The idea is to make a reasonable guess so that the rest of your analysis can still proceed without losing data points.

The simplest approach is to replace every missing value with a single fixed number. For example, you might replace all NaN entries with zero, or with one. This is done using the fillna() function. You pass the replacement value as an argument: fillna(0) replaces every missing value with 0, and fillna(1) replaces them with 1.

The textbook shows a concrete example using a DataFrame called df that contains marks for a student named Raman. First, we extract Raman's data:

dfRaman = df.loc[df['Name']=='Raman']

Then we look at his Science marks across four unit tests.

The fourth test has a missing value. To fill it with zero:

dfFillZeroScience = dfScience.fillna(0)

The result is shown below.

Now we can compute Raman's percentage in Science. The total marks per test are 25, and there are 4 tests (the number of rows in dfRaman). So the calculation is:

dfFillZeroScience.sum() * 100 / (25 * row)

This gives 60.0%.

Note

The variable row here is the number of rows in dfRaman (which is 4), not the number of rows in the Science series. The formula uses 25 * row because each test is out of 25 marks and there are 4 tests.

Forward Fill and Backward Fill

A more context-aware method is to use the value just before or just after the missing entry. This is especially useful when data has a natural order, like test scores over time.

  • fillna(method='pad') (or method='ffill') replaces a missing value with the value that comes immediately before it.
  • fillna(method='bfill') replaces a missing value with the value that comes immediately after it.

The textbook demonstrates this with Raman's English marks.

Using forward fill:

dfFillPadEng = dfEng.fillna(method='pad')

The missing value in test 4 is replaced by the marks of test 3 (23.0).

Now the percentage in English is calculated the same way:

dfFillPadEng.sum() * 100 / (25 * row)

This yields 91.0%.

Watch out

Forward fill and backward fill only work sensibly when the data has a meaningful sequence. If the missing value is at the very beginning of the series, method='pad' cannot fill it (it remains NaN). Similarly, method='bfill' cannot fill a missing value at the very end. …

Table 3.50Output of dfScience -- Raman's Science marks before filling
IndexScience
021.0
120.0
219.0
3NaN
Table 3.51Output of dfFillZeroScience = dfScience.fillna(0)
IndexScience
021.0
120.0
219.0
30.0
Table 3.52Output of dfEng -- Raman's English marks before filling
IndexEng
021.0
124.0
223.0
3NaN
Table 3.53Output of dfFillPadEng = dfEng.fillna(method='pad')
IndexEng
021.0
124.0
223.0
323.0

Name: Eng, dtype: float64 …