Skip to content

Informatics Practices · Ch 3 — Data Handling using Pandas – II

Data Aggregations

3.3

Data Aggregations

What Aggregation Means

Aggregation is the process of taking a dataset — a column, a set of columns, or the whole DataFrame — and reducing it to a single numeric value. Think of it as summarising: instead of looking at all 12 marks in Maths, you ask for the maximum Maths mark, and you get one number. That's aggregation.

Pandas provides a set of built-in aggregate functions that work on Series and DataFrames. The ones covered in this section are:

  • max() — the largest value
  • min() — the smallest value
  • sum() — the total
  • count() — the number of non-null entries
  • std() — standard deviation (a measure of spread)
  • var() — variance (another measure of spread)

These functions can be applied to a single column, multiple columns, or the entire DataFrame at once.

Applying a Single Aggregate Function

The simplest way is to use the .aggregate() method (or its shorthand .agg()) with the name of the function as a string.

df.aggregate('max')

When you run this on the student DataFrame, it returns the maximum value from every column — including non-numeric ones. For example, the output shows:

  • Name: Zuhaire (the alphabetically last name)
  • UT: 3
  • Maths: 24
  • Science: 25
  • S.St: 25
  • Hindi: 25
  • Eng: 24

The data type of the result is object because the Name column is a string.

Using Multiple Aggregate Functions Together

You can pass a list of function names to .aggregate() to get several summaries at once.

df.aggregate(['max', 'count'])

This produces a table where each row is an aggregate function and each column is a column from the original DataFrame. For the student data, the output is shown below.

The count row tells you that every column has 12 non-null entries.

Aggregating a Single Column

You can also apply aggregation to just one column. For example:

df['Maths'].aggregate(['max', 'min'])

This returns the max and min shown below.

The axis Parameter

The .aggregate() method accepts an axis parameter. By default, axis=0, which means the aggregation is performed column-wise (down each column). You can explicitly write it:

df['Maths'].aggregate(['max', 'min'], axis=0)

This gives exactly the same result as the previous example — because axis=0 is the default.

Row-wise Aggregation with axis=1

When you set axis=1, the aggregation happens row-wise — across columns for each row. This is useful when you want, for example, the total marks of each student in two subjects.

df[['Maths', 'Science']].aggregate('sum', axis=1)
``` …
Table 3.20Output of df.aggregate('max')
ColumnMax Value
NameZuhaire
UT3
Maths24
Science25
S.St25
Hindi25
Table 3.21Output of df.aggregate(['max','count'])
NameUTMathsScienceS.StHindiEng
maxZuhaire32425252524
Table 3.22Output of df['Maths'].aggregate(['max','min'])
StatisticValue
max24
min12
Table 3.23Output of df[['Maths','Science']].aggregate('sum', axis=1)
IndexSum
043
141
233
337
438
540
642
746
837
937
1039