Informatics Practices · Class 12 Optional
Ch 3Data Handling using Pandas – II — Class 12 Informatics Practices, concept-first.
As discussed in the previous chapter, Pandas is a well-established Python library used for manipulating, processing, and analysing data. You have already learned the basic operations on Series and DataFrame — creating them and accessing data from them.
Key concepts
Hover a concept to preview it and jump to its most relevant Q&A.
DataFrame Creation
Think of a spreadsheet you might use to track your monthly expenses. You have rows — one for each day or each purchase — and columns with labels like "Date", "Item", "Category", "Amount", and "Payment Method".
Most relevant Q&A
- Store the Result data in a DataFrame called marksUT. (Create the DataFrame from the case-study marks of the four students — Raman, Zuhaire,…Preview
- Write the statement which will sort the marks in English in the DataFrame df based on Unit Test 3, in descending order.Preview
- Write the steps required to read data from a MySQL database to a DataFrame.Preview
- Assuming the given table: Product. Write the python code for the following: | Item | Company | Rupees | USD | |------|---------|--------|---…Preview
- State whether the following statement is True or False: In Python, we cannot create an empty DataFrame.Preview
In previous exams
How often this chapter’s concepts have been examined — real appearance data, never estimated.
Chapter contents
The NCERT structure, section by section. Open a section to see its questions, then read the concept-first solution.
Introduction
As discussed in the previous chapter, Pandas is a well-established Python library used for manipulating, processing, and analysing data.
Case Study
This chapter's running example is a case study built around marks scored in unit tests held at a school (Table 3.1).
Descriptive Statistics
Descriptive statistics are the methods we use to get a basic understanding of a dataset — to summarise it rather than look at every single value.
Calculating Maximum Values
The max() method in Pandas returns the largest value from a DataFrame. It works on all data types — for numbers it gives the highest numeric value, and for strings it gives the alphabetically last nam…
Calculating Minimum Values
The min() method on a DataFrame returns the smallest value from each column (or each row, if you specify the axis).
Calculating Sum of Values
3 QThe sum() method in Pandas is used to calculate totals from a DataFrame. When called on the entire DataFrame without any column filter, it returns the sum of every column — including text columns.
+−Programs1 question
+−Think & Reflect1 question
Calculating Number of Values
DataFrame.count() is the method used to find out how many non-null values exist in a DataFrame. It does not count missing (NaN) entries — only actual data.
+−Programs1 question
Calculating Mean
2 QThe mean() method in Pandas calculates the average of numeric values. When called on a DataFrame without any arguments, it returns the mean of every numeric column.
+−Programs1 question
Calculating Median
2 QThe median() function in Pandas returns the middle value of a dataset. For a DataFrame, calling .median() gives you the median of each numeric column individually.
Calculating Mode
The mode is the value that appears most frequently in a dataset. While the mean and median give you a central tendency based on calculation or position, the mode tells you which value is most common —…
Calculating Quartile
The quantile() method on a DataFrame is the tool for calculating quartiles. A quartile divides a sorted dataset into four equal parts.
Calculating Variance
Variance measures how spread out a set of numbers is. In Pandas, the .var() method on a DataFrame or Series calculates this spread.
Calculating Standard Deviation
The standard deviation tells you how spread out the numbers in a dataset are from their average (mean).
Data Aggregations
Aggregation is the process of taking a dataset — a column, a set of columns, or the whole DataFrame — and reducing it to a single numeric value.
Sorting a DataFrame
Sorting means arranging data elements in a specified order — either ascending (smallest to largest) or descending (largest to smallest).
GROUP BY Functions
2 QThe groupby() function in Pandas is used to split a DataFrame into groups based on some criterion — typically the values in one or more columns.
+−Programs1 question
Altering the Index
The index of a DataFrame is the set of row labels used to access and retrieve data quickly. By default, Pandas assigns a numeric index starting from 0, as shown in the sample DataFrame of student mark…
Other DataFrame Operations
In this section, we learn more techniques and functions that can be used to manipulate and analyse data in a DataFrame.
Reshaping Data
The way a dataset is arranged into rows and columns is referred to as the shape of the data. Reshaping data means changing that shape — rearranging the dataset — to make it suitable for particular ana…
Pivot
2 QThe pivot function is used to reshape the data and create a new DataFrame from the original one. The textbook builds the idea through Example 3.1 — sales and profit data of four stores (S1, S2, S3 and…
+−Worked Examples1 question
Pivoting by Multiple Columns
To pivot on more than one value column at a time, pass a list of column names to the values parameter of the pivot() function.
Pivot Table
The pivottable() function works like the pivot() function, but with one crucial difference: when several rows share the same values for the specified index/column combination, it aggregates those rows…
Handling Missing Values
A DataFrame can contain many rows, where each row holds values for different columns (attributes). When a value for a particular column is absent, it is called a missing value.
Checking Missing Values
2 QReal-world data is often incomplete. A student might be absent for a test, a sensor might fail to record a reading, or a survey question might be left blank.
Dropping Missing Values
When a dataset contains missing values (shown as NaN in Pandas), you have two broad choices: either drop the rows that contain those missing values, or replace them with some other value.
Estimating Missing Values
Missing values in a dataset are a loss of information. Instead of simply dropping those rows, you can fill them with an estimated or approximated value.
Import and Export of Data between Pandas and MySQL
To work with real-world data, you will rarely type every value into a DataFrame by hand. Instead, data usually lives in files (like CSV or text files) or inside a database.
Importing Data from MySQL to Pandas
The core idea is simple: you take a table that lives inside a MySQL database and pull it into your Python environment as a pandas DataFrame.
Exporting Data from Pandas to MySQL
Exporting data from Pandas to MySQL means writing the contents of a pandas DataFrame into a table inside a MySQL database.
Summary
- Descriptive statistics are used to quantitatively summarise a dataset. - Pandas provides many statistical functions for analysing data, including max(), min(), mean(), median(), mode(), std(), and v…
Exercises
+−Show 14 questionsHide questions14 questions
- Q1Write the statement to install the python connector to connect MySQL i.e. pymysql.Free
- Q2Explain the difference between pivot() and pivot_table() function?Free
- Q3What is sqlalchemy?Free
- Q4Can you sort a DataFrame with respect to multiple columns?Preview
- Q5What are missing values? What are the strategies to handle them?Preview
- Q6Define the following terms: Median, Standard Deviation and variance.Preview
- Q7What do you understand by the term MODE? Name the function which is used to calculate it.Preview
- Q8Write the purpose of Data aggregation.Preview
- Q9Explain the concept of GROUP BY with help on an example.Preview
- Q10Write the steps required to read data from a MySQL database to a DataFrame.Preview
- Q11Explain the importance of reshaping of data with an example.Preview
- Q12Why estimation is an important concept in data analysis?Preview
- Q13Assuming the given table: Product. Write the python code for the following: | Item | Company | Rupees | USD | |------|---------|--------|---…Preview
- Q14Write the python statement for the following question on the basis of given dataset: [Table: the printed dataset — | | Name | Degree | Score…Preview
CBSE Sample Papers
Questions from official CBSE sample papers.
+−Show 9 questionsHide questions9 questions
- Q1Which of the following is NOT true with respect to CSV files ? (A) Values are separated by commas. (B) to_csv() can be used to save a datafr…Preview
- Q2We can add a new row to a DataFrame DF using the ______ method. (A) DF.add() (B) DF.loc[] (C) DF.loc() (D) DF.addloc[]Preview
- Q3CSV stands for ______ Separated Values. (A) Colon (B) CTRL (C) Comma (D) CaretPreview
- Q4In a Pandas DataFrame, which value for axis will be used to delete a column ? (A) axis = -1 (B) axis = 0 (C) axis = 1 (D) axis = 2Preview
- Q5(a) List any two characteristics of Pandas library in Python. OR (b) Briefly explain the purpose of Boolean Indexing with respect to DataFra…Preview
- Q6State whether the following statement is True or False: In Python, we cannot create an empty DataFrame.Preview
- Q7Which of the following Python statements is used to change a column label in a DataFrame, df? (A) df = df.rename({old_name: new_name}, axis=…Preview
- Q8In Python Pandas, DataFrame.______[] is used for label indexing with DataFrames. (A) label (B) index (C) labindex (D) locPreview
- Q9Assertion (A): The drop() method in Pandas can be used to delete rows and columns from a DataFrame. Reason (R): The axis parameter in the dr…Preview
More questions
+−Show 8 questionsHide questions8 questions
- Q1Solved Case Study based on Open Datasets: UCI dataset is a collection of open datasets, available to the public for experimentation and rese…Free
- Q2Give description of the generated DataFrame autodf. (Dataset: the UCI 'auto-mpg' open dataset — 398 rows, nine attributes: mpg, cylinders, d…Free
- Q3Display the first 10 rows of the DataFrame autodf. (Dataset: the UCI 'auto-mpg' open dataset — 398 rows, nine attributes: mpg, cylinders, di…Free
- Q4Find the attributes which have missing values. Handle the missing values using following two ways: (i) Replace the missing values by a value…Preview
- Q5Print the details of the car which gave the maximum mileage. (Dataset: the UCI 'auto-mpg' open dataset loaded into DataFrame autodf — 398 ro…Preview
- Q6Find the average displacement of the car given the number of cylinders. (Dataset: the UCI 'auto-mpg' open dataset loaded into DataFrame auto…Preview
- Q7What is the average number of cylinders in a car? (Dataset: the UCI 'auto-mpg' open dataset loaded into DataFrame autodf — 398 rows, nine at…Preview
- Q8Determine the no. of cars with weight greater than the average weight. (Dataset: the UCI 'auto-mpg' open dataset loaded into DataFrame autod…Preview