Skip to content
Activities · Activity 3.2

Q.Find the median of the values of the rows of the DataFrame.
[Table: The chapter's case study (Program 3.1) stores unit test marks of 4 students (maximum marks 25 in each subject) in a DataFrame df created from the dictionary marksUT = {'Name':['Raman','Raman','Raman','Zuhaire','Zuhaire','Zuhaire','Ashravy','Ashravy','Ashravy','Mishti','Mishti','Mishti'], 'UT':[1,2,3,1,2,3,1,2,3,1,2,3], 'Maths':[22,21,14,20,23,22,23,24,12,15,18,17], 'Science':[21,20,19,17,15,18,19,22,25,22,21,18], 'S.St':[18,17,15,22,21,19,20,24,19,25,25,20], 'Hindi':[20,22,24,24,25,23,15,17,21,22,24,25], 'Eng':[21,24,23,19,15,13,22,21,23,22,23,20]}.]

Dnh Dd CbseNCERTSubjective· 2mImportance★★★★★est
9% · 5/53 Questions
🔒 Locked · start free trial →

You're viewing a preview — the full solution, concept, methods & PYQ mapping are locked.

Start your 14-day free trial to unlock the full solution →

Calculate the median of marks across all subjects for each row (each unit test) using df.median(axis=1).

The question asks for the median of each row — that is, for each student's unit test, we want the middle value among the five subject marks. The median is a measure of central tendency that is robust to outliers; if a student scored unusually high or low in one subject, the median gives a better "typical" score than the mean.

The key decision is the axis parameter. In pandas:

  • axis=0 (the default) operates down columns — it would give the median of all students' marks in each subject.
  • axis=1 operates across columns (along each row) — exactly what we need here.

We build the DataFrame from the given dictionary, then call .median(axis=1) to compute the median of the five subject columns for each of the twelve rows.

import pandas as pd

marksUT = {
    'Name': ['Raman', 'Raman', 'Raman', 'Zuhaire', 'Zuhaire', 'Zuhaire',
             'Ashravy', 'Ashravy', 'Ashravy', 'Mishti', 'Mishti', 'Mishti'],
    'UT': [1, 2, 3, 1, 2, 3, 1, 2, 3, 1, 2, 3],
    'Maths': [22, 21, 14, 20, 23, 22, 23, 24, 12, 15, 18, 17],
    'Science': [21, 20, 19, 17, 15, 18, 19, 22, 25, 22, 21, 18],
    'S.St': [18, 17, 15, 22, 21, 19, 20, 24, 19, 25, 25, 20],
    'Hindi': [20, 22, 24, 24, 25, 23, 15, 17, 21, 22, 24, 25],
    'Eng': [21, 24, 23, 19, 15, 13, 22, 21, 23, 22, 23, 20]
}

df = pd.DataFrame(marksUT)

# Calculate median across the five subject columns for each row
row_medians = df.median(axis=1)

print(row_medians)

Output:

0     21.0
1     21.0
2     19.0
3     20.0
4     21.0
5     19.0
6     20.0
7     22.0
8     21.0
9     22.0
10    23.0
11    20.0
dtype: float64

Why this works:

When we call df.median(axis=1), pandas looks at each row and computes the median of all numeric columns. The DataFrame has seven columns: Name, UT, Maths, Science, S.St, Hindi, Eng. The first two are metadata; the last five are the subject marks we care about.

Fortunately, Name is a string (non-numeric) and is automatically excluded. However, UT is numeric. If we want the median of only the five subject columns, we should be explicit:

subject_cols = ['Maths', 'Science', 'S.St', 'Hindi', 'Eng']
row_medians = df[subject_cols].median(axis=1)
print(row_medians)
``` …

Unlock everything free for 14 days

  • Full step-by-step solutions
  • Concept-first explanations
  • Methods, shortcuts & mistakes
  • PYQ mapping + timed mock tests

Full access for 14 days. No credit card required.