Informatics Practices · Ch 2 — Data Handling using Pandas – I
Creation of Series
Creation of Series
There are different ways in which a Series can be created in Pandas. Whichever way we choose, the first step is always the same: to create or use a Series, we must first import the Pandas library. …
Creation of Series from Scalar Values
The simplest way to build a Series is directly from scalar values — a plain Python list of data items passed to pd.Series():
>>> import pandas as pd #import Pandas with alias pd
>>> series1 = pd.Series([10,20,30]) #create a Series
>>> print(series1) #Display the series
Output:
0 10
1 20
2 30
dtype: int64
Observe that the output is shown in two columns — the index on the left and the data value on the right. If we do not explicitly specify an index while creating the Series, Pandas assigns default indices ranging from 0 through N − 1, where N is the number of data elements (here 3, so the indices run 0, 1, 2). The dtype: int64 line reports the data type of the values.
We are not restricted to the default index: we can assign our own user-defined labels to the index and later use them to access elements of the Series. The labels can even be numbers in any (random) order:
>>> series2 = pd.Series(["Kavi","Shyam","Ravi"], index=[3,5,1])
>>> print(series2) #Display the series
Output:
3 Kavi
5 Shyam
1 Ravi
dtype: object
Here the data values Kavi, Shyam and Ravi have the index values 3, 5 and 1, respectively — the index is whatever we declare it to be, not necessarily 0-based or ordered. Note also that dtype is now object, since the values are strings.
Letters or strings work as indices just as well:
>>> series2 = pd.Series([2,3,4], index=["Feb","Mar","Apr"])
>>> print(series2) #Display the series
Output:
Feb 2
Mar 3 …
Creation of Series from NumPy Arrays
A Series can also be created from a one-dimensional (1D) NumPy array. NumPy is imported alongside Pandas (conventionally with the alias np), the array is built with np.array(), and the array object is passed to pd.Series():
>>> import numpy as np # import NumPy with alias np
>>> import pandas as pd
>>> array1 = np.array([1,2,3,4])
>>> series3 = pd.Series(array1)
>>> print(series3)
Output:
0 1
1 2
2 3
3 4
dtype: int32
As with scalar values, no explicit index means the default 0 to N − 1 index is generated (0 to 3 here for the 4 array elements). Letters or strings can again be supplied as indices through the index= parameter:
>>> series4 = pd.Series(array1, index = ["Jan", "Feb", "Mar", "Apr"])
>>> print(series4)
Output:
Jan 1
Feb 2
Mar 3
Apr 4
dtype: int32
One rule must be respected when index labels are passed with an array: the index and the array must be of the same size, otherwise Pandas raises a ValueError. In the example below, array1 contains 4 values but only 3 indices are supplied, so the error is displayed:
>>> series5 = pd.Series(array1, index = ["Jan", "Feb", "Mar"])
Output:
ValueError: Length of passed values is 4, index implies 3
(This size-matching constraint did not arise with the default index, because Pandas generates exactly as many default labels as there are data values.) …
Creation of Series from Dictionary
Recall that a Python dictionary stores key: value pairs, and a value can be quickly retrieved when its key is known. This structure maps naturally onto a Series: the dictionary keys can be used to construct the index, while the dictionary values become the data. When a dictionary is passed to pd.Series(), each key becomes an index label for its corresponding value:
>>> dict1 = {'India': 'NewDelhi', 'UK': 'London', 'Japan': 'Tokyo'}
>>> print(dict1) #Display the dictionary
Output:
{'India': 'NewDelhi', 'UK': 'London', 'Japan': 'Tokyo'}
>>> series8 = pd.Series(dict1)
>>> print(series8) #Display the series
Output:
India NewDelhi
UK London
Japan Tokyo
dtype: object
``` …