Skip to content
Exercises · Q5

Q.What are missing values? What are the strategies to handle them?

CBSENCERTSubjective· 3mImportance★★★★★est
47% · 25/53 Questions
🔒 Locked · start free trial →

You're viewing a preview — the full solution, concept, methods & PYQ mapping are locked.

Start your 14-day free trial to unlock the full solution →

Missing values (NaN) are unrecorded data points; you handle them either by deleting the incomplete rows/columns or by imputing -- filling them with an estimated value. Which strategy you pick genuinely changes your results, not just your code -- as the chapter's own Raman example shows.

What Are Missing Values?

Missing values are entries where no data was recorded for a particular row/column combination -- shown in pandas as NaN (Not a Number). They show up for many real reasons: a student was absent for a test, a sensor failed to log a reading, a form field was left blank, or a value genuinely doesn't apply to that record. Left unhandled, missing values break most statistical functions and machine-learning algorithms, which is why Section 3.8 of this chapter treats "detect the gaps, then decide a strategy, then apply it" as a standard workflow.

The Two Families of Strategy

1. Deletion -- remove the incomplete rows or columns.

  • Row-wise (dropna()): drop any row containing a missing value. Simple, but every dropped row is lost information -- fine when only a handful of rows are affected, risky when many are.
  • Column-wise: drop an entire column if too much of it is missing (rule of thumb: more than 70-80%) -- you lose the whole feature, but keep every row.

2. Imputation -- keep every row, but fill the gap with an estimate.

  • Constant fill (fillna(0), fillna(value)): replace with a fixed number -- appropriate when "missing" genuinely means a specific value (e.g. a test not attempted counted as 0 marks).
  • Mean/median fill: replace with the column's average -- median is safer when the column has outliers.
  • Mode fill: replace with the most frequent value -- the standard choice for categorical columns.
  • Forward/backward fill (fillna(method='pad') / fillna(method='bfill')): carry the previous (or next) valid value forward -- natural for ordered data, such as a student's test-to-test performance.
  • Interpolation (interpolate()): estimate a value that fits smoothly between the surrounding data points.
  • Model-based (KNN, regression): predict the missing value from the OTHER columns -- more accurate, but heavier to set up, and not covered in this chapter.

Seeing the Difference a Strategy Makes, With the Chapter's Own Data

This chapter's case study (Section 3.8) gives a perfect illustration: Raman was absent for Maths, Science and English in Unit Test 4 (S.St and Hindi ARE present that test). Applying three different, equally legitimate strategies to essentially the same problem -- "what percentage did Raman score?" -- gives three different, honest answers:

Strategy 1 -- delete the incomplete row (dropna()):

dfRaman = df[df['Name']=='Raman']
dfRaman.dropna(inplace=True, how='any')   # UT4's row is removed entirely

With UT4 gone, only 3 Unit Tests remain, so Maths' percentage is computed out of 75 (25x3) instead of 100 -- and comes out to 76.0%.

Strategy 2 -- fill the gap with zero (fillna(0)):

dfRaman = df.loc[df['Name']=='Raman']
(row, col) = dfRaman.shape            # row = 4 (all 4 Unit Tests, NaN included)
dfScience = dfRaman.loc[:, 'Science']
dfFillZeroScience = dfScience.fillna(0)
print(dfFillZeroScience.sum() * 100 / (25 * row), "%")

Output:

60.0 %
``` …

Unlock everything free for 14 days

  • Full step-by-step solutions
  • Concept-first explanations
  • Methods, shortcuts & mistakes
  • PYQ mapping + timed mock tests

Full access for 14 days. No credit card required.