Q.Write the purpose of Data aggregation.
You're viewing a preview — the full solution, concept, methods & PYQ mapping are locked.
Start your 14-day free trial to unlock the full solution →Data aggregation combines many rows into one summary value (or a few, per group) -- the same idea whether you use SQL's GROUP BY, spreadsheet subtotals, or pandas' aggregate(). Its purpose is to make large datasets comprehensible: fewer numbers, more meaning.
Why Aggregate?
Raw data is rarely useful on its own -- a table of a thousand individual sales, sensor readings, or (as in this chapter's case study) unit-test marks tells you very little at a glance. Aggregation transforms that raw detail into one meaningful number (or a small handful) per column, per group, or per category, which is what actually supports:
- Summarization -- turning many rows into a few interpretable numbers (a class's average marks, rather than every individual mark).
- Trend/pattern identification -- spotting the best/worst performer, the busiest month, the weakest subject.
- Decision-making -- teachers, managers, or analysts act on summaries, not raw rows.
- Reporting and dashboards -- charts and report cards show aggregated figures (totals, averages, maxima), not raw transaction logs.
- Reduced data volume -- a summary is far cheaper to store, transmit, and query than the full underlying dataset.
Seeing It in Code
This chapter's own case-study df demonstrates several aggregate functions directly (Section 3.3), all built on the pattern df.aggregate(func):
print(df.aggregate('max'))
Output:
Name Zuhaire
UT 3
Maths 24
Science 25
S.St 25
Hindi 25
Eng 24
dtype: object
A single call summarizes all seven columns down to their best value -- this is aggregation in its simplest form: many rows collapsed into one row of maxima. Note that max() even works on the text Name column (alphabetically latest).
print(df.aggregate(['max', 'count']))
Output:
Name UT Maths Science S.St Hindi Eng
max Zuhaire 3 24 25 25 25 24
count 12 12 12 12 12 12 12
Passing a LIST of function names runs several aggregations in one call -- here both the maximum AND the count of each column, side by side. This is exactly the "one command, multiple summaries" convenience that makes aggregation efficient for reporting.
print(df['Maths'].aggregate(['max', 'min']))
Output:
max 24
min 12
Name: Maths, dtype: int64
Aggregation isn't restricted to the whole DataFrame -- here it's scoped to just the Maths column, showing the class's highest and lowest Maths score in one call.
print(df[['Maths', 'Science']].aggregate('sum', axis=1))
Output:
0 43
1 41
2 33
3 37
4 38
5 40
6 42
7 46
8 37
9 37
10 39
11 35 …
Unlock everything free for 14 days
- Full step-by-step solutions
- Concept-first explanations
- Methods, shortcuts & mistakes
- PYQ mapping + timed mock tests
Full access for 14 days. No credit card required.