Q.Find the median of the values of the rows of the DataFrame.
[Table: The chapter's case study (Program 3.1) stores unit test marks of 4 students (maximum marks 25 in each subject) in a DataFrame df created from the dictionary marksUT = {'Name':['Raman','Raman','Raman','Zuhaire','Zuhaire','Zuhaire','Ashravy','Ashravy','Ashravy','Mishti','Mishti','Mishti'], 'UT':[1,2,3,1,2,3,1,2,3,1,2,3], 'Maths':[22,21,14,20,23,22,23,24,12,15,18,17], 'Science':[21,20,19,17,15,18,19,22,25,22,21,18], 'S.St':[18,17,15,22,21,19,20,24,19,25,25,20], 'Hindi':[20,22,24,24,25,23,15,17,21,22,24,25], 'Eng':[21,24,23,19,15,13,22,21,23,22,23,20]}.]
You're viewing a preview — the full solution, concept, methods & PYQ mapping are locked.
Start your 14-day free trial to unlock the full solution →Calculate the median of marks across all subjects for each row (each unit test) using df.median(axis=1).
The question asks for the median of each row — that is, for each student's unit test, we want the middle value among the five subject marks. The median is a measure of central tendency that is robust to outliers; if a student scored unusually high or low in one subject, the median gives a better "typical" score than the mean.
The key decision is the axis parameter. In pandas:
axis=0(the default) operates down columns — it would give the median of all students' marks in each subject.axis=1operates across columns (along each row) — exactly what we need here.
We build the DataFrame from the given dictionary, then call .median(axis=1) to compute the median of the five subject columns for each of the twelve rows.
import pandas as pd
marksUT = {
'Name': ['Raman', 'Raman', 'Raman', 'Zuhaire', 'Zuhaire', 'Zuhaire',
'Ashravy', 'Ashravy', 'Ashravy', 'Mishti', 'Mishti', 'Mishti'],
'UT': [1, 2, 3, 1, 2, 3, 1, 2, 3, 1, 2, 3],
'Maths': [22, 21, 14, 20, 23, 22, 23, 24, 12, 15, 18, 17],
'Science': [21, 20, 19, 17, 15, 18, 19, 22, 25, 22, 21, 18],
'S.St': [18, 17, 15, 22, 21, 19, 20, 24, 19, 25, 25, 20],
'Hindi': [20, 22, 24, 24, 25, 23, 15, 17, 21, 22, 24, 25],
'Eng': [21, 24, 23, 19, 15, 13, 22, 21, 23, 22, 23, 20]
}
df = pd.DataFrame(marksUT)
# Calculate median across the five subject columns for each row
row_medians = df.median(axis=1)
print(row_medians)
Output:
0 21.0
1 21.0
2 19.0
3 20.0
4 21.0
5 19.0
6 20.0
7 22.0
8 21.0
9 22.0
10 23.0
11 20.0
dtype: float64
Why this works:
When we call df.median(axis=1), pandas looks at each row and computes the median of all numeric columns. The DataFrame has seven columns: Name, UT, Maths, Science, S.St, Hindi, Eng. The first two are metadata; the last five are the subject marks we care about.
Fortunately, Name is a string (non-numeric) and is automatically excluded. However, UT is numeric. If we want the median of only the five subject columns, we should be explicit:
subject_cols = ['Maths', 'Science', 'S.St', 'Hindi', 'Eng']
row_medians = df[subject_cols].median(axis=1)
print(row_medians)
``` …
Unlock everything free for 14 days
- Full step-by-step solutions
- Concept-first explanations
- Methods, shortcuts & mistakes
- PYQ mapping + timed mock tests
Full access for 14 days. No credit card required.