Q.Consider the given Data-Frame 'health'. Disease name Agent 0 Common cold Virus 1 Chickenpox Virus 2 Cholera Bacteria 3 Tuberculosis Bacteria Write suitable Python statements for the following:
You're viewing a preview — the full solution, concept, methods & PYQ mapping are locked.
Start your 14-day free trial to unlock the full solution →This question tests your ability to manipulate a Pandas DataFrame in Python — specifically, removing a row by condition, adding a new row, and displaying a slice of rows.
Let's start by understanding the DataFrame we're working with. The 'health' DataFrame has two columns: 'Disease name' and 'Agent'. It contains four rows of data, each representing a disease and its causative agent. The diseases listed are Common cold (Virus), Chickenpox (Virus), Cholera (Bacteria), and Tuberculosis (Bacteria). This is a small, clean dataset, perfect for practicing basic DataFrame operations.
Now, the three tasks you need to perform are common data manipulation steps in Pandas. Let's go through each one.
Task (i): Remove the row containing details of disease named Tuberculosis.
To remove a row based on a condition, you use the drop method or boolean indexing. The most straightforward approach here is to use boolean indexing: you select all rows where the 'Disease name' column is NOT equal to 'Tuberculosis'. In Pandas, this is written as:
health = health[health['Disease name'] != 'Tuberculosis']
Alternatively, you could use the drop method with the index label of that row. Since the DataFrame has default integer indices (0,1,2,3), the row for Tuberculosis is at index 3. So you could also write:
health = health.drop(3)
Both methods achieve the same result. The first method is more robust because it works even if the index is not numeric or if the row order changes.
When using drop, remember to assign the result back to the variable (or use inplace=True) — otherwise the original DataFrame remains unchanged.
Task (ii): Add a new disease named 'Malaria' caused by 'Protozoa'.
Adding a new row to a DataFrame is done using the append method or the newer pd.concat function. Since we're adding a single row, append is simpler. You create a new DataFrame (or a dictionary) for the new row and append it to the existing DataFrame.
new_row = {'Disease name': 'Malaria', 'Agent': 'Protozoa'}
health = health.append(new_row, ignore_index=True)
The ignore_index=True parameter is important — it resets the index so the new row gets a fresh index number (in this case, 4, since we removed one row earlier). Without it, the index might become messy.
In newer versions of Pandas (1.4.0+), append has been deprecated. The recommended approach is to use pd.concat. But for CBSE Class 12 exams, append is still accepted and commonly taught. If you want to be future-proof, use:
health = pd.concat([health, pd.DataFrame([new_row])], ignore_index=True)
Task (iii): Display the last 2 rows.
To display the last n rows of a DataFrame, Pandas provides the tail() method. For the last 2 rows, simply write:
print(health.tail(2))
This will show the last two rows of the current DataFrame. After removing Tuberculosis and adding Malaria, the last two rows would be the row for Cholera (if it's still there) and the newly added Malaria row. The exact output depends on the order of operations. …
Unlock everything free for 14 days
- Full step-by-step solutions
- Concept-first explanations
- Methods, shortcuts & mistakes
- PYQ mapping + timed mock tests
Full access for 14 days. No credit card required.