Skip to content
← Informatics Practices

Informatics Practices · Class 12 Optional

Ch 3Data Handling using Pandas – II — Class 12 Informatics Practices, concept-first.

As discussed in the previous chapter, Pandas is a well-established Python library used for manipulating, processing, and analysing data. You have already learned the basic operations on Series and DataFrame — creating them and accessing data from them.

53

Q&A

24

Concepts

Not available

Exam weightage

Start learning — read this chapter →

Key concepts

Hover a concept to preview it and jump to its most relevant Q&A.

In previous exams

How often this chapter’s concepts have been examined — real appearance data, never estimated.

Chapter contents

The NCERT structure, section by section. Open a section to see its questions, then read the concept-first solution.

3.1

Introduction

As discussed in the previous chapter, Pandas is a well-established Python library used for manipulating, processing, and analysing data.

Case Study

This chapter's running example is a case study built around marks scored in unit tests held at a school (Table 3.1).

3.2

Descriptive Statistics

Descriptive statistics are the methods we use to get a basic understanding of a dataset — to summarise it rather than look at every single value.

3.2.1

Calculating Maximum Values

The max() method in Pandas returns the largest value from a DataFrame. It works on all data types — for numbers it gives the highest numeric value, and for strings it gives the alphabetically last nam…

3.2.2

Calculating Minimum Values

The min() method on a DataFrame returns the smallest value from each column (or each row, if you specify the axis).

3.2.3

Calculating Sum of Values

3 Q

The sum() method in Pandas is used to calculate totals from a DataFrame. When called on the entire DataFrame without any column filter, it returns the sum of every column — including text columns.

3.2.4

Calculating Number of Values

DataFrame.count() is the method used to find out how many non-null values exist in a DataFrame. It does not count missing (NaN) entries — only actual data.

3.2.5

Calculating Mean

2 Q

The mean() method in Pandas calculates the average of numeric values. When called on a DataFrame without any arguments, it returns the mean of every numeric column.

3.2.6

Calculating Median

2 Q

The median() function in Pandas returns the middle value of a dataset. For a DataFrame, calling .median() gives you the median of each numeric column individually.

3.2.7

Calculating Mode

The mode is the value that appears most frequently in a dataset. While the mean and median give you a central tendency based on calculation or position, the mode tells you which value is most common —…

3.2.8

Calculating Quartile

The quantile() method on a DataFrame is the tool for calculating quartiles. A quartile divides a sorted dataset into four equal parts.

3.2.9

Calculating Variance

Variance measures how spread out a set of numbers is. In Pandas, the .var() method on a DataFrame or Series calculates this spread.

3.2.10

Calculating Standard Deviation

The standard deviation tells you how spread out the numbers in a dataset are from their average (mean).

3.3

Data Aggregations

Aggregation is the process of taking a dataset — a column, a set of columns, or the whole DataFrame — and reducing it to a single numeric value.

3.4

Sorting a DataFrame

Sorting means arranging data elements in a specified order — either ascending (smallest to largest) or descending (largest to smallest).

3.5

GROUP BY Functions

2 Q

The groupby() function in Pandas is used to split a DataFrame into groups based on some criterion — typically the values in one or more columns.

3.6

Altering the Index

The index of a DataFrame is the set of row labels used to access and retrieve data quickly. By default, Pandas assigns a numeric index starting from 0, as shown in the sample DataFrame of student mark…

3.7

Other DataFrame Operations

In this section, we learn more techniques and functions that can be used to manipulate and analyse data in a DataFrame.

3.7.1

Reshaping Data

The way a dataset is arranged into rows and columns is referred to as the shape of the data. Reshaping data means changing that shape — rearranging the dataset — to make it suitable for particular ana…

(A)

Pivot

2 Q

The pivot function is used to reshape the data and create a new DataFrame from the original one. The textbook builds the idea through Example 3.1 — sales and profit data of four stores (S1, S2, S3 and…

(B)

Pivoting by Multiple Columns

To pivot on more than one value column at a time, pass a list of column names to the values parameter of the pivot() function.

(C)

Pivot Table

The pivottable() function works like the pivot() function, but with one crucial difference: when several rows share the same values for the specified index/column combination, it aggregates those rows…

3.8

Handling Missing Values

A DataFrame can contain many rows, where each row holds values for different columns (attributes). When a value for a particular column is absent, it is called a missing value.

3.8.1

Checking Missing Values

2 Q

Real-world data is often incomplete. A student might be absent for a test, a sensor might fail to record a reading, or a survey question might be left blank.

3.8.2

Dropping Missing Values

When a dataset contains missing values (shown as NaN in Pandas), you have two broad choices: either drop the rows that contain those missing values, or replace them with some other value.

3.8.3

Estimating Missing Values

Missing values in a dataset are a loss of information. Instead of simply dropping those rows, you can fill them with an estimated or approximated value.

3.9

Import and Export of Data between Pandas and MySQL

To work with real-world data, you will rarely type every value into a DataFrame by hand. Instead, data usually lives in files (like CSV or text files) or inside a database.

3.9.1

Importing Data from MySQL to Pandas

The core idea is simple: you take a table that lives inside a MySQL database and pull it into your Python environment as a pandas DataFrame.

3.9.2

Exporting Data from Pandas to MySQL

Exporting data from Pandas to MySQL means writing the contents of a pandas DataFrame into a table inside a MySQL database.

Summary

- Descriptive statistics are used to quantitatively summarise a dataset. - Pandas provides many statistical functions for analysing data, including max(), min(), mean(), median(), mode(), std(), and v…

Exercises

CBSE Sample Papers

Questions from official CBSE sample papers.

More questions