Skip to content

Informatics Practices · Ch 3 — Data Handling using Pandas – II

Dropping Missing Values

3.8.2

Dropping Missing Values

Dropping Missing Values

When a dataset contains missing values (shown as NaN in Pandas), you have two broad choices: either drop the rows that contain those missing values, or replace them with some other value. Dropping is the simpler, more aggressive approach — it removes the entire row (or object) that has at least one missing value.

Because dropping a row reduces the size of your dataset, you should only use this strategy when the number of missing values is small and confined to a few rows. If many rows have missing data, dropping them would shrink your dataset too much and could distort your analysis.

The dropna() function

Pandas provides the dropna() function to drop rows that contain missing values. By default, dropna() returns a new DataFrame with all rows that had any NaN removed. It does not modify the original DataFrame unless you explicitly tell it to.

Example from the textbook:

Suppose you have a DataFrame df containing student marks across multiple subjects and unit tests. One row (the 4th row) has a NaN value. Calling df.dropna() produces a new DataFrame df1 that excludes that row entirely.

>>> df1 = df.dropna()
>>> print(df1)

The output shows all rows except the one with the missing value — the row for Raman's 4th unit test is gone. The resulting DataFrame now has 16 rows (indexed 0 through 15, skipping the dropped row).

Using inplace=True and the how parameter

You can also apply dropna() directly to a DataFrame without creating a new variable by setting inplace=True. This modifies the original DataFrame in place.

The how parameter controls when a row is dropped:

  • how='any' (the default) — drops a row if any of its values are missing.
  • how='all' — drops a row only if all of its values are missing.

Textbook example — filtering and dropping:

The book shows a practical scenario: you want to calculate Raman's percentage in Maths across his unit tests.

First, you create a subset DataFrame containing only Raman's rows:

dfRaman = df[df.Name == 'Raman']

Then you drop any row in dfRaman that has a missing value, and you do it in place:

dfRaman.dropna(inplace=True, how='any')

After this, dfRaman has only 3 rows (Raman's 1st, 2nd, and 3rd unit tests) — the 4th row, which had a NaN in the Maths column, is removed.

Next, you extract the Maths column:

dfMaths = dfRaman['Maths']

As the textbook itself prints it, this listing still shows all four rows, including the row-3 NaN (the printed reprint reflects the DataFrame just before the drop is fully applied) — the same output already shown as a table above.

Finally, the percentage is calculated:

row = len(dfMaths)
print(dfMaths.sum() * 100 / (25 * row), "%")

Since row is 3 (not 4), the percentage is computed from only the three valid test scores. The result is 76.0%. …

Table 3.48Output of df1 = df.dropna() -- Raman's Unit Test 4 row removed
IndexNameUTMathsScienceS.StHindiEng
0Raman122.021.0182021.0
1Raman221.020.0172224.0
2Raman314.019.0152423.0
4Zuhaire120.017.0222419.0
5Zuhaire223.015.0212515.0
6Zuhaire322.018.0192313.0
7Zuhaire419.020.0172116.0
8Ashravy123.019.0201522.0
9Ashravy224.022.0241721.0
10Ashravy312.025.0192123.0
11Ashravy415.020.0202017.0
12Mishti115.022.0252222.0
13Mishti218.021.0252423.0
14Mishti317.018.0202520.0
Table 3.49Output of dfMaths after dfRaman.dropna(inplace=True, how='any')
IndexMaths
022.0
121.0
214.0
3NaN

Name: Maths, dtype: float64

*(As printed in the book -- dropna(how='any') removes a row only where THAT row has a missing value; row 3 still shows NaN here since this column-slice reprint happen …