Informatics Practices · Ch 3 — Data Handling using Pandas – II
Estimating Missing Values
Estimating Missing Values
Estimating Missing Values
Missing values in a dataset are a loss of information. Instead of simply dropping those rows, you can fill them with an estimated or approximated value. The idea is to make a reasonable guess so that the rest of your analysis can still proceed without losing data points.
The simplest approach is to replace every missing value with a single fixed number. For example, you might replace all NaN entries with zero, or with one. This is done using the fillna() function. You pass the replacement value as an argument: fillna(0) replaces every missing value with 0, and fillna(1) replaces them with 1.
The textbook shows a concrete example using a DataFrame called df that contains marks for a student named Raman. First, we extract Raman's data:
dfRaman = df.loc[df['Name']=='Raman']
Then we look at his Science marks across four unit tests.
The fourth test has a missing value. To fill it with zero:
dfFillZeroScience = dfScience.fillna(0)
The result is shown below.
Now we can compute Raman's percentage in Science. The total marks per test are 25, and there are 4 tests (the number of rows in dfRaman). So the calculation is:
dfFillZeroScience.sum() * 100 / (25 * row)
This gives 60.0%.
The variable row here is the number of rows in dfRaman (which is 4), not the number of rows in the Science series. The formula uses 25 * row because each test is out of 25 marks and there are 4 tests.
Forward Fill and Backward Fill
A more context-aware method is to use the value just before or just after the missing entry. This is especially useful when data has a natural order, like test scores over time.
fillna(method='pad')(ormethod='ffill') replaces a missing value with the value that comes immediately before it.fillna(method='bfill')replaces a missing value with the value that comes immediately after it.
The textbook demonstrates this with Raman's English marks.
Using forward fill:
dfFillPadEng = dfEng.fillna(method='pad')
The missing value in test 4 is replaced by the marks of test 3 (23.0).
Now the percentage in English is calculated the same way:
dfFillPadEng.sum() * 100 / (25 * row)
This yields 91.0%.
Forward fill and backward fill only work sensibly when the data has a meaningful sequence. If the missing value is at the very beginning of the series, method='pad' cannot fill it (it remains NaN). Similarly, method='bfill' cannot fill a missing value at the very end. …
| Index | Science |
|---|---|
| 0 | 21.0 |
| 1 | 20.0 |
| 2 | 19.0 |
| 3 | NaN |
| Index | Science |
|---|---|
| 0 | 21.0 |
| 1 | 20.0 |
| 2 | 19.0 |
| 3 | 0.0 |
| Index | Eng |
|---|---|
| 0 | 21.0 |
| 1 | 24.0 |
| 2 | 23.0 |
| 3 | NaN |
| Index | Eng |
|---|---|
| 0 | 21.0 |
| 1 | 24.0 |
| 2 | 23.0 |
| 3 | 23.0 |
Name: Eng, dtype: float64 …