Informatics Practices · Ch 3 — Data Handling using Pandas – II
Calculating Mean
Calculating Mean
The mean() method in Pandas calculates the average of numeric values. When called on a DataFrame without any arguments, it returns the mean of every numeric column. This is a column-wise operation by default — it adds up all the values in a column and divides by the number of entries.
For example, if you have a DataFrame df containing marks in subjects like Maths, Science, and English, calling df.mean() will give you a Series where each subject name is paired with its average score. The output will show the mean for each column, and the data type of the result will be float64. The method only works on numeric columns; any non-numeric columns are simply ignored.
mean() will silently skip non-numeric columns. If you try to calculate the mean of a column containing text or dates, Pandas will not raise an error — it will just omit that column from the output.
The textbook demonstrates this with a DataFrame df containing marks for five subjects (Maths, Science, S.St, Hindi, Eng) and a column called UT. The output shows:
- UT: 2.5000
- Maths: 18.6000
- Science: 19.8000
- S.St: 20.0000
- Hindi: 21.3125
- Eng: 19.8000
All values are of type float64.
Calculating row-wise averages using axis=1
The real power of mean() appears when you need the average across columns for each row — for instance, finding the average marks a student scored in all subjects in a single unit test. To do this, you pass the argument axis=1 to mean().
The textbook walks through a concrete example using a student named Zuhaire. The DataFrame dfZuhaire contains marks for multiple unit tests. First, you slice the DataFrame to keep only the subject columns (Maths through Eng), excluding any non-subject columns like roll number or name. This is done using .loc[:, 'Maths':'Eng'] and stored in a new variable dfZuhaireMarks.
dfZuhaireMarks = dfZuhaire.loc[:, 'Maths':'Eng']
Printing this sliced DataFrame shows the marks for three unit tests (rows 3, 4, and 5) across five subjects:
| Maths | Science | S.St | Hindi | Eng | |
|---|---|---|---|---|---|
| 3 | 20 | 17 | 22 | 24 | 19 |
| 4 | 23 | 15 | 21 | 25 | 15 |
| 5 | 22 | 18 | 19 | 23 | 13 |
Now, to get the average marks obtained by Zuhaire in each unit test, you call mean(axis=1) on this sliced DataFrame:
dfZuhaireMarks.mean(axis=1)
The output is a Series where each row index (the unit test number) is paired with the average of that row's marks:
- Row 3 (Unit Test 1): 20.4
- Row 4 (Unit Test 2): 19.8 …
| Column | Mean |
|---|---|
| UT | 2.5000 |
| Maths | 18.6000 |
| Science | 19.8000 |
| S.St | 20.0000 |
| Hindi | 21.3125 |
Write the statements to get an average of marks obtained by Zuhaire in all the Unit Tests. This worked example slices out only the subject-mark columns for one student, then takes the row-wise mean (axis …