Skip to content
Case Study · Q4

Q.Find the attributes which have missing values. Handle the missing values using following two ways:

(i) Replace the missing values by a value before that.
(ii) Remove the rows having missing values from the original dataset. (Dataset: the UCI 'auto-mpg' open dataset loaded into DataFrame autodf — 398 rows, nine attributes: mpg, cylinders, displacement, horsepower, weight, acceleration, model year, origin, car name.)
Uttar Pradesh UpmspTextbookSubjective· 3mImportance★★★★★est
75% · 40/53 Questions
🔒 Locked · start free trial →

You're viewing a preview — the full solution, concept, methods & PYQ mapping are locked.

Start your 14-day free trial to unlock the full solution →

Concept understanding — Missing Data Imputation

Missing Data Imputation: A First Look

Imagine you are a researcher studying consumer behaviour in rural markets. You have a questionnaire with 500 responses, but when you flip through the stack, you notice that some people left certain questions blank. Maybe they skipped the income column because they felt uncomfortable, or they forgot to tick their age group. Now you have a problem: your analysis needs complete information to draw reliable conclusions, but your data has holes.

This is the everyday reality of working with real-world data. Missing data is not an exception — it is the norm. And how you handle those gaps can make or break your entire study.

What Is Missing Data Imputation?

Missing data imputation is the process of filling in the missing values in a dataset with reasonable substitutes, so that the dataset becomes complete and usable for analysis. The word "imputation" comes from the Latin imputare — to bring into the account. You are essentially making an educated guess about what the missing information would have been, based on the information you already have.

Think of it like a jigsaw puzzle where a few pieces have fallen off the table. You do not throw away the entire puzzle. Instead, you look at the surrounding pieces — the colours, the shapes, the patterns — and you reconstruct what the missing piece probably looked like. That reconstruction is imputation.

Why Does It Matter?

If you simply ignore the missing data — a practice called "listwise deletion" — you risk losing a large portion of your sample. In a survey of 500 people, if even 50 have one missing answer, you might end up discarding all 50 responses entirely. That shrinks your sample size, reduces the statistical power of your analysis, and can introduce bias. The people who skipped a question may be systematically different from those who answered it. For example, if higher-income respondents are more likely to skip the income question, deleting them would make your sample look poorer than it really is.

Imputation helps you preserve your sample size and, done carefully, reduces the bias that missing data would otherwise introduce.

Common Approaches (Conceptually, Not Mathematically)

The NCERT textbook for Class 11 Statistics (Chapter 2, on Collection of Data) touches on the problem of non-response and missing information. It emphasises that the researcher must decide how to handle such gaps, and that the method chosen should be appropriate to the nature of the data. Here are the main conceptual strategies:

  • Mean/Median Imputation: You replace the missing value with the average (mean) or middle value (median) of the available responses for that question. This is simple and quick, but it artificially reduces the variability in your data. If you fill every missing income with the average income, your dataset will look more uniform than reality.

  • Mode Imputation: For categorical data — like "urban/rural" or "yes/no" — you fill the gap with the most common category. This works well when one category dominates, but it can distort the distribution if the missing values are not random.

  • Last Observation Carried Forward: In time-series data (e.g., monthly sales figures), if a value is missing, you assume it stayed the same as the previous month. This is intuitive but can hide trends or sudden changes.

  • Hot-Deck Imputation: You find a "donor" record that is similar to the incomplete record on other variables, and copy the missing value from that donor. For instance, if a 35-year-old urban male skipped his income, you look at another 35-year-old urban male who did report income, and use that value. This preserves the natural variability better than mean imputation. …

Unlock everything free for 14 days

  • Full step-by-step solutions
  • Concept-first explanations
  • Methods, shortcuts & mistakes
  • PYQ mapping + timed mock tests

Full access for 14 days. No credit card required.