Q.What are missing values? What are the strategies to handle them?
You're viewing a preview — the full solution, concept, methods & PYQ mapping are locked.
Start your 14-day free trial to unlock the full solution →Missing values (NaN) are unrecorded data points; you handle them either by deleting the incomplete rows/columns or by imputing -- filling them with an estimated value. Which strategy you pick genuinely changes your results, not just your code -- as the chapter's own Raman example shows.
What Are Missing Values?
Missing values are entries where no data was recorded for a particular row/column combination -- shown in pandas as NaN (Not a Number). They show up for many real reasons: a student was absent for a test, a sensor failed to log a reading, a form field was left blank, or a value genuinely doesn't apply to that record. Left unhandled, missing values break most statistical functions and machine-learning algorithms, which is why Section 3.8 of this chapter treats "detect the gaps, then decide a strategy, then apply it" as a standard workflow.
The Two Families of Strategy
1. Deletion -- remove the incomplete rows or columns.
- Row-wise (
dropna()): drop any row containing a missing value. Simple, but every dropped row is lost information -- fine when only a handful of rows are affected, risky when many are. - Column-wise: drop an entire column if too much of it is missing (rule of thumb: more than 70-80%) -- you lose the whole feature, but keep every row.
2. Imputation -- keep every row, but fill the gap with an estimate.
- Constant fill (
fillna(0),fillna(value)): replace with a fixed number -- appropriate when "missing" genuinely means a specific value (e.g. a test not attempted counted as 0 marks). - Mean/median fill: replace with the column's average -- median is safer when the column has outliers.
- Mode fill: replace with the most frequent value -- the standard choice for categorical columns.
- Forward/backward fill (
fillna(method='pad')/fillna(method='bfill')): carry the previous (or next) valid value forward -- natural for ordered data, such as a student's test-to-test performance. - Interpolation (
interpolate()): estimate a value that fits smoothly between the surrounding data points. - Model-based (KNN, regression): predict the missing value from the OTHER columns -- more accurate, but heavier to set up, and not covered in this chapter.
Seeing the Difference a Strategy Makes, With the Chapter's Own Data
This chapter's case study (Section 3.8) gives a perfect illustration: Raman was absent for Maths, Science and English in Unit Test 4 (S.St and Hindi ARE present that test). Applying three different, equally legitimate strategies to essentially the same problem -- "what percentage did Raman score?" -- gives three different, honest answers:
Strategy 1 -- delete the incomplete row (dropna()):
dfRaman = df[df['Name']=='Raman']
dfRaman.dropna(inplace=True, how='any') # UT4's row is removed entirely
With UT4 gone, only 3 Unit Tests remain, so Maths' percentage is computed out of 75 (25x3) instead of 100 -- and comes out to 76.0%.
Strategy 2 -- fill the gap with zero (fillna(0)):
dfRaman = df.loc[df['Name']=='Raman']
(row, col) = dfRaman.shape # row = 4 (all 4 Unit Tests, NaN included)
dfScience = dfRaman.loc[:, 'Science']
dfFillZeroScience = dfScience.fillna(0)
print(dfFillZeroScience.sum() * 100 / (25 * row), "%")
Output:
60.0 %
``` …
Unlock everything free for 14 days
- Full step-by-step solutions
- Concept-first explanations
- Methods, shortcuts & mistakes
- PYQ mapping + timed mock tests
Full access for 14 days. No credit card required.