Skip to content

Informatics Practices · Ch 2 — Data Handling using Pandas – I

Accessing Elements of a Series

2.2.2

Accessing Elements of a Series

Once a Series has been created, we need ways to get at its contents. There are two common ways of accessing the elements of a Series: indexing and slicing. …

(A)

Indexing

Indexing is the mechanism for accessing individual elements of a Series, and it works very much like indexing on NumPy arrays. Pandas supports two kinds of indexes:

  • Positional index — an integer that corresponds to the element's position in the Series, counting from 0.
  • Labelled index — any user-defined label that has been assigned as the index.

Accessing a value by positional index. When a Series has the default index, we simply put the position number in square brackets:

>>> seriesNum = pd.Series([10,20,30])
>>> seriesNum[2]
30

The value 30 is displayed because it sits at positional index 2 (positions are 0, 1, 2 for the three values).

Accessing a value by labelled index. When explicit labels have been given to a Series, those labels can themselves be used inside the square brackets. Here the value 3 is displayed for the labelled index Mar:

>>> seriesMnths = pd.Series([2,3,4], index=["Feb","Mar","Apr"])
>>> seriesMnths["Mar"]
3

The same idea works with any labels. In the next example, the value NewDelhi is displayed for the labelled index India:

>>> seriesCapCntry = pd.Series(['NewDelhi', 'WashingtonDC', 'London', 'Paris'],
                               index=['India', 'USA', 'UK', 'France'])
>>> seriesCapCntry['India']
'NewDelhi'

Importantly, giving a Series labels does not take away positional access — we can still reach an element through its position:

>>> seriesCapCntry[1]
'WashingtonDC'

Accessing more than one element at a time. A list of positional integers, or a list of index labels, selects several elements in one statement. The elements come back in the order you asked for them:

>>> seriesCapCntry[[3,2]]
France     Paris
UK        London
dtype: object
>>> seriesCapCntry[['UK','USA']]
UK           London
USA    WashingtonDC
dtype: object

Note the double square brackets — the outer pair does the indexing, the inner pair builds the list of positions or labels.

Changing the index values. The index associated with a Series is not fixed forever; assigning a new sequence to the index attribute replaces the existing labels:

>>> seriesCapCntry.index = [10,20,30,40]
>>> seriesCapCntry …
(B)

Slicing

Sometimes we need only a part of a Series rather than a single element or the whole thing. Extracting a contiguous portion is called slicing, and it works like slicing on NumPy arrays: we specify start and end parameters as [start:end] after the Series name.

Slicing with positional indices — the end is excluded. When positions are used, the value at the end position is left out, so exactly end - start data values are extracted. Consider the familiar Series:

>>> seriesCapCntry = pd.Series(['NewDelhi', 'WashingtonDC', 'London', 'Paris'],
                               index=['India', 'USA', 'UK', 'France'])
>>> seriesCapCntry[1:3]   # excludes the value at index position 3
USA    WashingtonDC
UK           London
dtype: object

Only the values at positions 1 and 2 appear in the output — position 3 (Paris) is excluded.

Slicing with labelled indices — the end is included. If labels are used for slicing, the value at the end label is part of the output. This is the key difference from positional slicing:

>>> seriesCapCntry['USA':'France']
USA       WashingtonDC
UK              London
France           Paris
dtype: object

Here the slice runs from label USA up to and including label France.

Reversing a Series with a slice. A step of -1 returns the Series in reverse order:

>>> seriesCapCntry[::-1]
France           Paris
UK              London
USA       WashingtonDC
India         NewDelhi
dtype: object

Modifying values through a slice. Slicing is not only for reading — assigning to a slice updates the selected elements. First build a Series of the numbers 10 to 15 with letter labels:

>>> import numpy as np
>>> seriesAlph = pd.Series(np.arange(10,16,1),
                           index = ['a', 'b', 'c', 'd', 'e', 'f'])
>>> seriesAlph
a    10
b    11
c    12
d    13
e    14
f    15
dtype: int32

Assigning through a positional slice again excludes the end position — only positions 1 and 2 (labels b and c) are changed here:

>>> seriesAlph[1:3] = 50
>>> seriesAlph
a    10
b    50
c    50
d    13
e    14
f    15
dtype: int32
``` …