Skip to content

Informatics Practices · Ch 2 — Data Handling using Pandas – I

Accessing DataFrames Element through Slicing

2.3.4

Accessing DataFrames Element through Slicing

Slicing in DataFrames

Slicing is a powerful way to select a subset of rows and/or columns from a DataFrame. Unlike Python lists, slicing in DataFrames is inclusive of the end value — meaning both the start and the stop labels are included in the result.

Slicing Rows by Label

You can retrieve a set of consecutive rows by using their row labels with the .loc[] accessor. For example, if you have a DataFrame called ResultDF with row labels like 'Maths', 'Science', 'Hindi', etc., the following statement will display rows from 'Maths' through 'Science':

>>> ResultDF.loc['Maths': 'Science']

This will show all columns for the rows labelled 'Maths' and 'Science' — both are included.

Important

In DataFrames, slicing with .loc[] is inclusive of the end label. This is different from Python list slicing, which excludes the end index.

Slicing Rows and Selecting a Single Column

You can combine a row slice with a specific column name to get values from only that column for the selected rows. The syntax is:

>>> ResultDF.loc['Maths': 'Science', 'Arnab']

This returns the marks of Arnab in Maths and Science — two values, one per row.

Slicing Both Rows and Columns

To access a rectangular block of data, use a slice of row labels and a slice of column labels together:

>>> ResultDF.loc['Maths': 'Science', 'Arnab':'Samridhi']

This gives you a smaller DataFrame containing rows Maths and Science, and columns Arnab, Ramit, and Samridhi (all inclusive).

Slicing Rows with a List of Columns

Instead of a column slice, you can provide a list of specific column names. This is useful when the columns you want are not consecutive:

>>> ResultDF.loc['Maths': 'Science', ['Arnab', 'Samridhi']]

This returns only the columns Arnab and Samridhi for the rows Maths and Science.

Tip

Use a list of column names when you need non-adjacent columns. Use a slice when the columns are in a continuous range.

Filtering Rows Using Boolean Lists

You can also filter rows by passing a list of Boolean values (True or False) to .loc[]. Each value in the list corresponds to a row in the DataFrame — True means the row is included, False means it is omitted.

For example, if ResultDF has three rows (Maths, Science, Hindi), the following statement will show only the first and third rows, skipping the second row (Science):

>>> ResultDF.loc[[True, False, True]]

The output will display all columns for Maths and Hindi, but not for Science. …

Table 2.8ResultDF.loc['Maths':'Science']
ArnabRamitSamridhiRiyaMallika
Maths9092898194
Table 2.9ResultDF.loc['Maths':'Science', 'Arnab':'Samridhi']
ArnabRamitSamridhi
Maths909289
Table 2.10ResultDF.loc['Maths':'Science', ['Arnab', 'Samridhi']]
ArnabSamridhi
Maths9089

Filtering Rows in DataFrames

Filtering rows in a DataFrame using .loc[] can also be done with a Boolean list — a list of True/False values, one per row, in the same order as the DataFrame's rows. True keeps that row, False omits it.

For example, ResultDF has three rows: Maths, Science, Hindi. Passing [True, False, True] keeps the first and third rows (Maths and Hindi) and skips the second (Science):

>>> ResultDF.loc[[True, False, True]]
ArnabRamitSamridhiRiyaMallika
Maths9092898194
Hindi9796886799