Skip to content

Informatics Practices · Ch 2 — Data Handling using Pandas – I

Accessing DataFrames Element through Indexing

2.3.3

Accessing DataFrames Element through Indexing

Individual data elements in a DataFrame can be accessed using indexing. Pandas offers two ways of indexing DataFrames: label based indexing, where we pick out data by the names of rows and columns, and Boolean indexing, where we pick out data by applying True/False conditions to the values themselves. …

(A)

Label Based Indexing

Pandas provides several methods for label based indexing, and the most important of them for DataFrames is DataFrame.loc[ ]. With loc[ ] we address data by the names of rows and columns rather than by their positions.

The textbook continues with the ResultDF DataFrame created earlier (at this point it holds the Maths, Science and Hindi rows):

>>> ResultDF
         Arnab  Ramit  Samridhi  Riya  Mallika
Maths       90     92        89    81       94
Science     91     81        91    71       95
Hindi       97     96        88    67       99

A single row label returns the row as a Series. Passing one row label inside loc[ ] gives back that entire row, presented as a Series whose index is the column labels:

>>> ResultDF.loc['Science']
Arnab       91
Ramit       81
Samridhi    91
Riya        71
Mallika     95
Name: Science, dtype: int64

An integer inside loc[ ] is a label, not a position. When the row label passed is an integer value, loc[ ] interprets it as a label of the index, not as an integer position along the index. The book shows this with a DataFrame of multiples of 10, whose default index labels happen to be the integers 0, 1, 2, …:

>>> dFrame10Multiples = pd.DataFrame([10,20,30,40,50])
>>> dFrame10Multiples.loc[2]
0    30
Name: 2, dtype: int64

Here loc[2] returns the row whose label is 2 (the value 30) — it just happens that label and position coincide for a default index.

A single column label returns the column as a Series. To select a column with loc[ ], we give : for the rows (meaning "all rows") followed by the column label:

>>> ResultDF.loc[:, 'Arnab']
Maths      90
Science    91
Hindi      97
Name: Arnab, dtype: int64

The textbook points out that the same result — the marks of Arnab in all the subjects — can also be obtained with the plain square-bracket column access we met earlier:

>>> print(df['Arnab'])

Both forms return Arnab's column as a Series.

A list of row labels returns a DataFrame. To read more than one row at once, pass a list of row labels — note the double square brackets [[ ]]. Where a single label gave a Series, a list of labels gives back a DataFrame:

(B)

Boolean Indexing

Boolean means a binary variable that can represent either of two states — True (indicated by 1) or False (indicated by 0). In Boolean indexing, we select subsets of data based on the actual values in the DataFrame rather than their row or column labels. In practice this means we can write conditions on column (or row) names and use them to filter data values.

Consider the ResultDF DataFrame again. The following statement asks, for the Maths row, which students scored more than 90 — and displays True or False for each student depending on whether the data value satisfies the condition or not:

>>> ResultDF.loc['Maths'] > 90
Arnab       False
Ramit        True
Samridhi    False
Riya        False
Mallika      True
Name: Maths, dtype: bool

Only Ramit (92) and Mallika (94) scored above 90 in Maths, so only their entries are True. Notice the result is itself a Series — of dtype: bool.

The same idea works down a column. To check in which subjects Arnab has scored more than 90, we apply the condition to Arnab's column:

>>> ResultDF.loc[:, 'Arnab'] > 90
Maths      False
Science     True
Hindi       True
Name: Arnab, dtype: bool

Arnab's 91 in Science and 97 in Hindi satisfy the condition, while his 90 in Maths does not (the condition is strictly greater than 90). …