Skip to content

Informatics Practices · Ch 2 — Data Handling using Pandas – I

Creation of Series

2.2.1

Creation of Series

There are different ways in which a Series can be created in Pandas. Whichever way we choose, the first step is always the same: to create or use a Series, we must first import the Pandas library. …

(A)

Creation of Series from Scalar Values

The simplest way to build a Series is directly from scalar values — a plain Python list of data items passed to pd.Series():

>>> import pandas as pd   #import Pandas with alias pd
>>> series1 = pd.Series([10,20,30])  #create a Series
>>> print(series1)  #Display the series

Output:

0    10
1    20
2    30
dtype: int64

Observe that the output is shown in two columns — the index on the left and the data value on the right. If we do not explicitly specify an index while creating the Series, Pandas assigns default indices ranging from 0 through N − 1, where N is the number of data elements (here 3, so the indices run 0, 1, 2). The dtype: int64 line reports the data type of the values.

We are not restricted to the default index: we can assign our own user-defined labels to the index and later use them to access elements of the Series. The labels can even be numbers in any (random) order:

>>> series2 = pd.Series(["Kavi","Shyam","Ravi"], index=[3,5,1])
>>> print(series2)  #Display the series

Output:

3     Kavi
5    Shyam
1     Ravi
dtype: object

Here the data values Kavi, Shyam and Ravi have the index values 3, 5 and 1, respectively — the index is whatever we declare it to be, not necessarily 0-based or ordered. Note also that dtype is now object, since the values are strings.

Letters or strings work as indices just as well:

>>> series2 = pd.Series([2,3,4], index=["Feb","Mar","Apr"])
>>> print(series2) #Display the series

Output:

Feb    2
Mar    3 …
(B)

Creation of Series from NumPy Arrays

A Series can also be created from a one-dimensional (1D) NumPy array. NumPy is imported alongside Pandas (conventionally with the alias np), the array is built with np.array(), and the array object is passed to pd.Series():

>>> import numpy as np  # import NumPy with alias np
>>> import pandas as pd
>>> array1 = np.array([1,2,3,4])
>>> series3 = pd.Series(array1)
>>> print(series3)

Output:

0    1
1    2
2    3
3    4
dtype: int32

As with scalar values, no explicit index means the default 0 to N − 1 index is generated (0 to 3 here for the 4 array elements). Letters or strings can again be supplied as indices through the index= parameter:

>>> series4 = pd.Series(array1, index = ["Jan", "Feb", "Mar", "Apr"])
>>> print(series4)

Output:

Jan    1
Feb    2
Mar    3
Apr    4
dtype: int32

One rule must be respected when index labels are passed with an array: the index and the array must be of the same size, otherwise Pandas raises a ValueError. In the example below, array1 contains 4 values but only 3 indices are supplied, so the error is displayed:

>>> series5 = pd.Series(array1, index = ["Jan", "Feb", "Mar"])

Output:

ValueError: Length of passed values is 4, index implies 3

(This size-matching constraint did not arise with the default index, because Pandas generates exactly as many default labels as there are data values.) …

(C)

Creation of Series from Dictionary

Recall that a Python dictionary stores key: value pairs, and a value can be quickly retrieved when its key is known. This structure maps naturally onto a Series: the dictionary keys can be used to construct the index, while the dictionary values become the data. When a dictionary is passed to pd.Series(), each key becomes an index label for its corresponding value:

>>> dict1 = {'India': 'NewDelhi', 'UK': 'London', 'Japan': 'Tokyo'}
>>> print(dict1)  #Display the dictionary

Output:

{'India': 'NewDelhi', 'UK': 'London', 'Japan': 'Tokyo'}
>>> series8 = pd.Series(dict1)
>>> print(series8)  #Display the series

Output:

India    NewDelhi
UK         London
Japan       Tokyo
dtype: object
``` …