Informatics Practices · Ch 3 — Data Handling using Pandas – II
Data Aggregations
Data Aggregations
What Aggregation Means
Aggregation is the process of taking a dataset — a column, a set of columns, or the whole DataFrame — and reducing it to a single numeric value. Think of it as summarising: instead of looking at all 12 marks in Maths, you ask for the maximum Maths mark, and you get one number. That's aggregation.
Pandas provides a set of built-in aggregate functions that work on Series and DataFrames. The ones covered in this section are:
max()— the largest valuemin()— the smallest valuesum()— the totalcount()— the number of non-null entriesstd()— standard deviation (a measure of spread)var()— variance (another measure of spread)
These functions can be applied to a single column, multiple columns, or the entire DataFrame at once.
Applying a Single Aggregate Function
The simplest way is to use the .aggregate() method (or its shorthand .agg()) with the name of the function as a string.
df.aggregate('max')
When you run this on the student DataFrame, it returns the maximum value from every column — including non-numeric ones. For example, the output shows:
- Name: Zuhaire (the alphabetically last name)
- UT: 3
- Maths: 24
- Science: 25
- S.St: 25
- Hindi: 25
- Eng: 24
The data type of the result is object because the Name column is a string.
Using Multiple Aggregate Functions Together
You can pass a list of function names to .aggregate() to get several summaries at once.
df.aggregate(['max', 'count'])
This produces a table where each row is an aggregate function and each column is a column from the original DataFrame. For the student data, the output is shown below.
The count row tells you that every column has 12 non-null entries.
Aggregating a Single Column
You can also apply aggregation to just one column. For example:
df['Maths'].aggregate(['max', 'min'])
This returns the max and min shown below.
The axis Parameter
The .aggregate() method accepts an axis parameter. By default, axis=0, which means the aggregation is performed column-wise (down each column). You can explicitly write it:
df['Maths'].aggregate(['max', 'min'], axis=0)
This gives exactly the same result as the previous example — because axis=0 is the default.
Row-wise Aggregation with axis=1
When you set axis=1, the aggregation happens row-wise — across columns for each row. This is useful when you want, for example, the total marks of each student in two subjects.
df[['Maths', 'Science']].aggregate('sum', axis=1)
``` …
| Column | Max Value |
|---|---|
| Name | Zuhaire |
| UT | 3 |
| Maths | 24 |
| Science | 25 |
| S.St | 25 |
| Hindi | 25 |
| Name | UT | Maths | Science | S.St | Hindi | Eng | |
|---|---|---|---|---|---|---|---|
| max | Zuhaire | 3 | 24 | 25 | 25 | 25 | 24 |
| Statistic | Value |
|---|---|
| max | 24 |
| min | 12 |
| Index | Sum |
|---|---|
| 0 | 43 |
| 1 | 41 |
| 2 | 33 |
| 3 | 37 |
| 4 | 38 |
| 5 | 40 |
| 6 | 42 |
| 7 | 46 |
| 8 | 37 |
| 9 | 37 |
| 10 | 39 |