Q.Find the attributes which have missing values. Handle the missing values using following two ways:
You're viewing a preview — the full solution, concept, methods & PYQ mapping are locked.
Start your 14-day free trial to unlock the full solution →Missing data in the horsepower column of the auto-mpg dataset is handled by (i) forward-fill using .ffill() and (ii) row-wise deletion using .dropna(), with the original DataFrame preserved for the second operation.
This is a data cleaning / imputation question — you are given a real-world dataset (the classic UCI auto-mpg) and asked to apply two standard missing-value strategies. The core idea is that real data is rarely perfect; you must detect where values are absent and then decide how to handle them. The two approaches here represent opposite philosophies: (i) fill the gap with a reasonable guess (the previous row's value), and (ii) remove the entire row containing the gap.
Let's first load the dataset and inspect it. The auto-mpg dataset is available through seaborn or can be read directly from a CSV. We'll assume it's already in a DataFrame called autodf.
import pandas as pd
import seaborn as sns
# Load the auto-mpg dataset
autodf = sns.load_dataset('mpg')
# Note: seaborn's 'mpg' dataset is the same as the UCI auto-mpg
Now, step one: find which attributes have missing values.
# Check for missing values in each column
missing_counts = autodf.isnull().sum()
print("Missing values per column:")
print(missing_counts[missing_counts > 0])
Expected output:
horsepower 6
dtype: int64
Only the horsepower column has missing values — exactly 6 rows out of 398. The other eight columns (mpg, cylinders, displacement, weight, acceleration, model year, origin, car name) are complete.
A common mistake is to assume that NaN values appear in multiple columns. Always verify with .isnull().sum() before applying any imputation — you might be fixing a problem that doesn't exist.
Now, task (i): Replace missing values by the value before that — this is forward fill (also called "last observation carried forward"). The idea: if a car's horsepower is missing, we assume it is similar to the car listed just before it in the dataset. This makes sense only if the data is sorted in some meaningful order (e.g., by model year or by manufacturer), but the question explicitly asks for this method, so we apply it.
# Create a copy to avoid modifying the original
autodf_ffill = autodf.copy()
autodf_ffill['horsepower'] = autodf_ffill['horsepower'].ffill()
# Verify no missing values remain
print("Missing after forward fill:", autodf_ffill['horsepower'].isnull().sum())
Expected output:
Missing after forward fill: 0
The .ffill() method propagates the last valid value forward. If the very first row had a missing value, it would remain NaN (since there is no "before" value), but in this dataset the first row's horsepower is present, so all 6 gaps are filled.
.ffill() is a shortcut for fillna(method='ffill'). Both do the same thing — use whichever reads more clearly in your code. …
Unlock everything free for 14 days
- Full step-by-step solutions
- Concept-first explanations
- Methods, shortcuts & mistakes
- PYQ mapping + timed mock tests
Full access for 14 days. No credit card required.