Informatics Practices · Ch 2 — Data Handling using Pandas – I
Creation of DataFrame
Creation of DataFrame
There are a number of ways to create a DataFrame, and this section lists some of them. Each of the lettered sub-sections that follow demonstrates one creation method with working code and its output. …
Creation of an empty DataFrame
The simplest DataFrame of all is an empty one — no columns, no rows. It is created by calling the DataFrame() constructor with no arguments:
>>> import pandas as pd
>>> dFrameEmt = pd.DataFrame()
>>> dFrameEmt
Empty DataFrame
Columns: []
Index: []
``` …
Creation of DataFrame from NumPy ndarrays
Pandas can build a DataFrame directly from NumPy ndarrays. To try this, first import NumPy and create three arrays:
>>> import numpy as np
>>> array1 = np.array([10,20,30])
>>> array2 = np.array([100,200,300])
>>> array3 = np.array([-10,-20,-30, -40])
From a single ndarray. Passing one ndarray to pd.DataFrame() creates a simple DataFrame without any column labels:
>>> dFrame4 = pd.DataFrame(array1)
>>> dFrame4
0
0 10
1 20
2 30
Since we supplied no labels, pandas falls back to its defaults: the single column is labelled 0, and the rows get the default index 0, 1, 2 — one row per element of the array.
From more than one ndarray. We can also pass a list of ndarrays. In that case each ndarray becomes one row of the DataFrame, and we can name the columns through the columns parameter:
>>> dFrame5 = pd.DataFrame([array1, array3, array2], columns=['A', 'B', 'C', 'D'])
>>> dFrame5
A B C D
0 10 20 30 NaN
1 -10 -20 -30 -40.0
2 100 200 300 NaN
Two things are worth observing in this output:
- The rows appear in the order the arrays were listed:
array1becomes row 0,array3becomes row 1, andarray2becomes row 2. …
Creation of DataFrame from List of Dictionaries
A DataFrame can also be created from a list of dictionaries. Each dictionary in the list becomes one row of the DataFrame:
# Create list of dictionaries
>>> listDict = [{'a':10, 'b':20}, {'a':5, 'b':10, 'c':20}]
>>> dFrameListDict = pd.DataFrame(listDict)
>>> dFrameListDict
a b c
0 10 20 NaN
1 5 10 20.0
The rules pandas follows here are:
- Dictionary keys become the column labels, and the values corresponding to each key are taken as the row's data.
- The number of rows equals the number of dictionaries in the list. In this example the list holds two dictionaries, so the DataFrame has two rows.
- The number of columns equals the maximum number of keys in any dictionary of the list. The first dictionary has two keys (
a,b) but the second has three (a,b,c), so the DataFrame gets three columns. NaN(Not a Number) is inserted wherever a corresponding value for a column is missing. The first dictionary has no keyc, so row 0 showsNaNin columnc. …
Creation of DataFrame from Dictionary of Lists
DataFrames can also be created from a dictionary of lists. Here each key of the dictionary names a column, and the list stored against that key supplies the values going down that column.
Consider a dictionary with the keys 'State', 'GArea' (geographical area) and 'VDF' (very dense forest), where the value against each key is a list:
>>> dictForest = {'State': ['Assam', 'Delhi', 'Kerala'],
'GArea': [78438, 1483, 38852],
'VDF' : [2797, 6.72, 1663]}
>>> dFrameForest = pd.DataFrame(dictForest)
>>> dFrameForest
State GArea VDF
0 Assam 78438 2797.00
1 Delhi 1483 6.72
2 Kerala 38852 1663.00
Note that dictionary keys become column labels by default, and the lists supply the row data. This is why the textbook says a DataFrame can be thought of as a dictionary of lists or a dictionary of series — each column is one named sequence of values.
Changing the sequence of columns. We can control the order in which the columns appear by passing a particular sequence of the dictionary keys as the columns parameter:
>>> dFrameForest1 = pd.DataFrame(dictForest, columns=['State', 'VDF', 'GArea'])
>>> dFrameForest1
State VDF GArea …
Creation of DataFrame from Series
A DataFrame can be created from one or more Series. Consider the following three Series (note that seriesC deliberately uses a different set of index labels):
seriesA = pd.Series([1,2,3,4,5],
index = ['a', 'b', 'c', 'd', 'e'])
seriesB = pd.Series([1000,2000,-1000,-5000,1000],
index = ['a', 'b', 'c', 'd', 'e'])
seriesC = pd.Series([10,20,-10,-50,100],
index = ['z', 'y', 'a', 'c', 'e'])
From a single Series. Passing one Series gives a DataFrame with as many rows as the Series has elements, but only one column. The Series' index labels become the row labels:
>>> dFrame6 = pd.DataFrame(seriesA)
>>> dFrame6
0
a 1
b 2
c 3
d 4
e 5
From more than one Series. To use several Series, pass them in a list. Now the picture flips: the labels of the Series become the column names, and each Series becomes one row of the DataFrame:
>>> dFrame7 = pd.DataFrame([seriesA, seriesB])
>>> dFrame7
a b c d e
0 1 2 3 4 5
1 1000 2000 -1000 -5000 1000
When the Series have different labels. Now look at what happens when the two Series do not share the same set of labels:
>>> dFrame8 = pd.DataFrame([seriesA, seriesC])
>>> dFrame8
a b c d e z y
0 1.0 2.0 3.0 4.0 5.0 NaN NaN
1 -10.0 NaN -50.0 NaN 100.0 10.0 20.0
``` …
Creation of DataFrame from Dictionary of Series
A dictionary of Series can also be used to create a DataFrame. The textbook's example is ResultSheet, a dictionary of Series holding the marks of 5 students in three subjects. The students' names are the keys of the dictionary, and the index values of each Series are the subject names:
>>> ResultSheet = {
'Arnab': pd.Series([90, 91, 97],
index=['Maths','Science','Hindi']),
'Ramit': pd.Series([92, 81, 96],
index=['Maths','Science','Hindi']),
'Samridhi': pd.Series([89, 91, 88],
index=['Maths','Science','Hindi']),
'Riya': pd.Series([81, 71, 67],
index=['Maths','Science','Hindi']),
'Mallika': pd.Series([94, 95, 99],
index=['Maths','Science','Hindi'])}
>>> ResultDF = pd.DataFrame(ResultSheet)
>>> ResultDF
Arnab Ramit Samridhi Riya Mallika
Maths 90 92 89 81 94
Science 91 81 91 71 95
Hindi 97 96 88 67 99
Here the dictionary keys (student names) have become the column labels, and the Series index values (subject names) have become the row labels. Every column of a DataFrame is itself a Series, which we can confirm with type():
>>> type(ResultDF.Arnab)
<class 'pandas.core.series.Series'>
Row labels are the union of all the indexes. When a DataFrame is created from a dictionary of Series, the resulting index (the row labels) is a union of all the Series indexes used to create it. The following example makes this visible — Series1 is indexed by a…e while Series2 and Series3 are indexed by z, y, a, c, e:
dictForUnion = {'Series1': pd.Series([1,2,3,4,5],
index = ['a', 'b', 'c', 'd', 'e']),
'Series2': pd.Series([10,20,-10,-50,100],
index = ['z', 'y', 'a', 'c', 'e']),
'Series3': pd.Series([10,20,-10,-50,100],
index = ['z', 'y', 'a', 'c', 'e'])}
>>> dFrameUnion = pd.DataFrame(dictForUnion)
>>> dFrameUnion
Series1 Series2 Series3
a 1.0 -10.0 -10.0 …