Skip to content

Informatics Practices · Ch 3 — Data Handling using Pandas – II

Introduction

3.1

Introduction

As discussed in the previous chapter, Pandas is a well-established Python library used for manipulating, processing, and analysing data. You have already learned the basic operations on Series and DataFrame — creating them and accessing data from them. Pandas offers much more powerful and useful functions for data analysis, and this chapter builds on that foundation.

In this chapter, you will work with more advanced features of DataFrame: sorting data, answering analytical questions using the data, cleaning data, and applying various useful functions to it. To make these techniques concrete, the chapter uses one running example dataset throughout, introduced next as a case study, so that every new method you learn is demonstrated on the same familiar data.

Table 3.1Case Study Result
Name/SubjectsUnit TestMathsScienceS.St.HindiEng
Raman12221182021
Raman22120172224
Raman31419152423
Zuhaire12017222419
Zuhaire22315212515
Zuhaire32218192313
Aashravy12319201522
Aashravy22422241721
Aashravy31225192123
Mishti11522252222
Mishti21821252423
Mishti31718202520

Case Study

This chapter's running example is a case study built around marks scored in unit tests held at a school (Table 3.1). For each unit test, the marks scored by every student in the class are recorded, out of a maximum of 25 marks per subject. The subjects tracked are Maths, Science, Social Studies (S.St.), Hindi, and English.

For simplicity, the textbook assumes a class of just four students — Raman, Zuhaire, Ashravy (the table heading spells this name 'Aashravy', while the code below spells it 'Ashravy' — a minor inconsistency in the source; both refer to the same student), and Mishti — with marks recorded across three unit tests.

The data is stored in a DataFrame named marksUT, built from a Python dictionary where each key becomes a column name and each value is a list of that column's entries:

import pandas as pd

marksUT = {
    'Name': ['Raman','Raman','Raman','Zuhaire','Zuhaire','Zuhaire',
             'Ashravy','Ashravy','Ashravy','Mishti','Mishti','Mishti'],
    'UT': [1,2,3,1,2,3,1,2,3,1,2,3],
    'Maths': [22,21,14,20,23,22,23,24,12,15,18,17],
    'Science': [21,20,19,17,15,18,19,22,25,22,21,18],
    'S.St': [18,17,15,22,21,19,20,24,19,25,25,20],
    'Hindi': [20,22,24,24,25,23,15,17,21,22,24,25],
    'Eng': [21,24,23,19,15,13,22,21,23,22,23,20]
}

df = pd.DataFrame(marksUT)
print(df)

Running this produces a DataFrame with 12 rows (indexed 0 to 11) and 7 columns — one row per student per unit test. Notice that each student appears three times, once per unit test: this is a "long format" layout, where every row is a single observation of one student's performance in one test.

Every technique introduced in the rest of this chapter — sorting, aggregating, grouping, reshaping, and handling missing values — is demonstrated on this same df DataFrame, so each new method builds directly on data you already understand.

Note

This DataFrame-creation walkthrough is Program 3-1 in the textbook.

A second, larger case study — built on the real-world UCI "auto-mpg" dataset — appears later in the chapter as a solved worked example, applying everything learned here to an external, open dataset.