Skip to content
Exercises · Q12

Q.Why estimation is an important concept in data analysis?

Uttarakhand UbseTextbookSubjective· 2mImportance★★★★★est
60% · 32/53 Questions
🔒 Locked · start free trial →

You're viewing a preview — the full solution, concept, methods & PYQ mapping are locked.

Start your 14-day free trial to unlock the full solution →

"Estimation" here refers to Section 3.8.3 of the chapter — filling in a missing value with a reasonable approximation (the previous test's marks, the column average, zero, etc.) using fillna(), instead of discarding the entire row with dropna(). It matters because it lets you keep using the good data in a row that has only one or two missing values.

What "estimation" means in this chapter

This question sits right after Section 3.8 ("Handling Missing Values") in the textbook, which explains that missing values can be handled in two ways:

  1. Drop the row/object that has the missing value (dropna()), or
  2. Fill or estimate the missing value with a plausible substitute (fillna())

"Estimation" is the second strategy — approximating what a missing value probably was, using context such as a nearby value, an average, or a fixed placeholder, rather than throwing away the whole record.

The problem estimation solves

Take the chapter's own case study. Table 3.2 records four Unit Tests for four students. Raman was absent for Unit Test 4 in Maths, Science and English — but he did sit S.St and Hindi that day (S.St = 19, Hindi = 18). If you simply dropna() on Raman's data, pandas removes the entire UT4 row, discarding his perfectly valid S.St and Hindi marks along with the three missing ones. For a student with many missing fields across many rows, this can shrink the dataset drastically and bias any summary statistic computed from what remains.

Estimation avoids that loss: you keep every row, and only the missing cell is replaced.

How estimation is done — fillna()

import pandas as pd
import numpy as np

dfRaman = df[df['Name'] == 'Raman']

# Marks scored by Raman in Science across the 4 Unit Tests
dfScience = dfRaman.loc[:, 'Science']
print(dfScience)

Output:

0    21.0
1    20.0
2    19.0
3     NaN
Name: Science, dtype: float64

Estimate by a fixed value (zero):

dfFillZeroScience = dfScience.fillna(0)
print(dfFillZeroScience)

percent_science = dfFillZeroScience.sum() * 100 / (25 * 4)
print("Percentage of Marks Scored by Raman in Science:", percent_science, "%")

Output:

0    21.0
1    20.0
2    19.0
3     0.0
Name: Science, dtype: float64

Percentage of Marks Scored by Raman in Science: 60.0 %

fillna(0) treats the missing test as a zero score — a conservative estimate that assumes the worst.

Estimate by the value before it (pad/ffill):

dfEng = dfRaman.loc[:, 'Eng']
dfFillPadEng = dfEng.fillna(method='pad')   # carries UT3's mark forward into UT4
print(dfFillPadEng)

percent_eng = dfFillPadEng.sum() * 100 / (25 * 4)
print("Percentage of Marks Scored by Raman in English:", percent_eng, "%")

Output:

0    21.0
1    24.0
2    23.0
3    23.0
Name: Eng, dtype: float64

Percentage of Marks Scored by Raman in English: 91.0 %

Here the missing UT4 English mark is estimated as UT3's mark (23) — a more optimistic, "assume consistency" estimate. fillna(method='bfill') would instead look forward for the next available value.

Why estimation matters, concretely …

Unlock everything free for 14 days

  • Full step-by-step solutions
  • Concept-first explanations
  • Methods, shortcuts & mistakes
  • PYQ mapping + timed mock tests

Full access for 14 days. No credit card required.