Skip to content

Informatics Practices · Ch 3 — Data Handling using Pandas – II

Calculating Mean

3.2.5

Calculating Mean

The mean() method in Pandas calculates the average of numeric values. When called on a DataFrame without any arguments, it returns the mean of every numeric column. This is a column-wise operation by default — it adds up all the values in a column and divides by the number of entries.

For example, if you have a DataFrame df containing marks in subjects like Maths, Science, and English, calling df.mean() will give you a Series where each subject name is paired with its average score. The output will show the mean for each column, and the data type of the result will be float64. The method only works on numeric columns; any non-numeric columns are simply ignored.

Watch out

mean() will silently skip non-numeric columns. If you try to calculate the mean of a column containing text or dates, Pandas will not raise an error — it will just omit that column from the output.

The textbook demonstrates this with a DataFrame df containing marks for five subjects (Maths, Science, S.St, Hindi, Eng) and a column called UT. The output shows:

  • UT: 2.5000
  • Maths: 18.6000
  • Science: 19.8000
  • S.St: 20.0000
  • Hindi: 21.3125
  • Eng: 19.8000

All values are of type float64.

Calculating row-wise averages using axis=1

The real power of mean() appears when you need the average across columns for each row — for instance, finding the average marks a student scored in all subjects in a single unit test. To do this, you pass the argument axis=1 to mean().

The textbook walks through a concrete example using a student named Zuhaire. The DataFrame dfZuhaire contains marks for multiple unit tests. First, you slice the DataFrame to keep only the subject columns (Maths through Eng), excluding any non-subject columns like roll number or name. This is done using .loc[:, 'Maths':'Eng'] and stored in a new variable dfZuhaireMarks.

dfZuhaireMarks = dfZuhaire.loc[:, 'Maths':'Eng']

Printing this sliced DataFrame shows the marks for three unit tests (rows 3, 4, and 5) across five subjects:

MathsScienceS.StHindiEng
32017222419
42315212515
52218192313

Now, to get the average marks obtained by Zuhaire in each unit test, you call mean(axis=1) on this sliced DataFrame:

dfZuhaireMarks.mean(axis=1)

The output is a Series where each row index (the unit test number) is paired with the average of that row's marks:

  • Row 3 (Unit Test 1): 20.4
  • Row 4 (Unit Test 2): 19.8 …
Table 3.11Output of df.mean() -- mean of each column
ColumnMean
UT2.5000
Maths18.6000
Science19.8000
S.St20.0000
Hindi21.3125
DefinitionProgram 3-6

Write the statements to get an average of marks obtained by Zuhaire in all the Unit Tests. This worked example slices out only the subject-mark columns for one student, then takes the row-wise mean (axis …