Skip to content

Informatics Practices · Ch 2 — Data Handling using Pandas – I

Creation of DataFrame

2.3.1

Creation of DataFrame

There are a number of ways to create a DataFrame, and this section lists some of them. Each of the lettered sub-sections that follow demonstrates one creation method with working code and its output. …

(A)

Creation of an empty DataFrame

The simplest DataFrame of all is an empty one — no columns, no rows. It is created by calling the DataFrame() constructor with no arguments:

>>> import pandas as pd
>>> dFrameEmt = pd.DataFrame()
>>> dFrameEmt
Empty DataFrame
Columns: []
Index: []
``` …
(B)

Creation of DataFrame from NumPy ndarrays

Pandas can build a DataFrame directly from NumPy ndarrays. To try this, first import NumPy and create three arrays:

>>> import numpy as np
>>> array1 = np.array([10,20,30])
>>> array2 = np.array([100,200,300])
>>> array3 = np.array([-10,-20,-30, -40])

From a single ndarray. Passing one ndarray to pd.DataFrame() creates a simple DataFrame without any column labels:

>>> dFrame4 = pd.DataFrame(array1)
>>> dFrame4
    0
0  10
1  20
2  30

Since we supplied no labels, pandas falls back to its defaults: the single column is labelled 0, and the rows get the default index 0, 1, 2 — one row per element of the array.

From more than one ndarray. We can also pass a list of ndarrays. In that case each ndarray becomes one row of the DataFrame, and we can name the columns through the columns parameter:

>>> dFrame5 = pd.DataFrame([array1, array3, array2], columns=['A', 'B', 'C', 'D'])
>>> dFrame5
     A    B    C     D
0   10   20   30   NaN
1  -10  -20  -30 -40.0
2  100  200  300   NaN

Two things are worth observing in this output:

  • The rows appear in the order the arrays were listed: array1 becomes row 0, array3 becomes row 1, and array2 becomes row 2. …
(C)

Creation of DataFrame from List of Dictionaries

A DataFrame can also be created from a list of dictionaries. Each dictionary in the list becomes one row of the DataFrame:

# Create list of dictionaries
>>> listDict = [{'a':10, 'b':20}, {'a':5, 'b':10, 'c':20}]
>>> dFrameListDict = pd.DataFrame(listDict)
>>> dFrameListDict
    a   b     c
0  10  20   NaN
1   5  10  20.0

The rules pandas follows here are:

  • Dictionary keys become the column labels, and the values corresponding to each key are taken as the row's data.
  • The number of rows equals the number of dictionaries in the list. In this example the list holds two dictionaries, so the DataFrame has two rows.
  • The number of columns equals the maximum number of keys in any dictionary of the list. The first dictionary has two keys (a, b) but the second has three (a, b, c), so the DataFrame gets three columns.
  • NaN (Not a Number) is inserted wherever a corresponding value for a column is missing. The first dictionary has no key c, so row 0 shows NaN in column c. …
(D)

Creation of DataFrame from Dictionary of Lists

DataFrames can also be created from a dictionary of lists. Here each key of the dictionary names a column, and the list stored against that key supplies the values going down that column.

Consider a dictionary with the keys 'State', 'GArea' (geographical area) and 'VDF' (very dense forest), where the value against each key is a list:

>>> dictForest = {'State': ['Assam', 'Delhi', 'Kerala'],
                  'GArea': [78438, 1483, 38852],
                  'VDF'  : [2797, 6.72, 1663]}
>>> dFrameForest = pd.DataFrame(dictForest)
>>> dFrameForest
    State   GArea    VDF
0   Assam   78438  2797.00
1   Delhi   1483     6.72
2   Kerala  38852  1663.00

Note that dictionary keys become column labels by default, and the lists supply the row data. This is why the textbook says a DataFrame can be thought of as a dictionary of lists or a dictionary of series — each column is one named sequence of values.

Changing the sequence of columns. We can control the order in which the columns appear by passing a particular sequence of the dictionary keys as the columns parameter:

>>> dFrameForest1 = pd.DataFrame(dictForest, columns=['State', 'VDF', 'GArea'])
>>> dFrameForest1
    State      VDF   GArea …
(E)

Creation of DataFrame from Series

A DataFrame can be created from one or more Series. Consider the following three Series (note that seriesC deliberately uses a different set of index labels):

seriesA = pd.Series([1,2,3,4,5],
                    index = ['a', 'b', 'c', 'd', 'e'])
seriesB = pd.Series([1000,2000,-1000,-5000,1000],
                    index = ['a', 'b', 'c', 'd', 'e'])
seriesC = pd.Series([10,20,-10,-50,100],
                    index = ['z', 'y', 'a', 'c', 'e'])

From a single Series. Passing one Series gives a DataFrame with as many rows as the Series has elements, but only one column. The Series' index labels become the row labels:

>>> dFrame6 = pd.DataFrame(seriesA)
>>> dFrame6
   0
a  1
b  2
c  3
d  4
e  5

From more than one Series. To use several Series, pass them in a list. Now the picture flips: the labels of the Series become the column names, and each Series becomes one row of the DataFrame:

>>> dFrame7 = pd.DataFrame([seriesA, seriesB])
>>> dFrame7
      a     b     c     d     e
0     1     2     3     4     5
1  1000  2000 -1000 -5000  1000

When the Series have different labels. Now look at what happens when the two Series do not share the same set of labels:

>>> dFrame8 = pd.DataFrame([seriesA, seriesC])
>>> dFrame8
      a    b     c    d      e     z     y
0   1.0  2.0   3.0  4.0    5.0   NaN   NaN
1 -10.0  NaN -50.0  NaN  100.0  10.0  20.0
``` …
(F)

Creation of DataFrame from Dictionary of Series

A dictionary of Series can also be used to create a DataFrame. The textbook's example is ResultSheet, a dictionary of Series holding the marks of 5 students in three subjects. The students' names are the keys of the dictionary, and the index values of each Series are the subject names:

>>> ResultSheet = {
'Arnab':    pd.Series([90, 91, 97],
                      index=['Maths','Science','Hindi']),
'Ramit':    pd.Series([92, 81, 96],
                      index=['Maths','Science','Hindi']),
'Samridhi': pd.Series([89, 91, 88],
                      index=['Maths','Science','Hindi']),
'Riya':     pd.Series([81, 71, 67],
                      index=['Maths','Science','Hindi']),
'Mallika':  pd.Series([94, 95, 99],
                      index=['Maths','Science','Hindi'])}
>>> ResultDF = pd.DataFrame(ResultSheet)
>>> ResultDF
         Arnab  Ramit  Samridhi  Riya  Mallika
Maths       90     92        89    81       94
Science     91     81        91    71       95
Hindi       97     96        88    67       99

Here the dictionary keys (student names) have become the column labels, and the Series index values (subject names) have become the row labels. Every column of a DataFrame is itself a Series, which we can confirm with type():

>>> type(ResultDF.Arnab)
<class 'pandas.core.series.Series'>

Row labels are the union of all the indexes. When a DataFrame is created from a dictionary of Series, the resulting index (the row labels) is a union of all the Series indexes used to create it. The following example makes this visible — Series1 is indexed by a…e while Series2 and Series3 are indexed by z, y, a, c, e:

dictForUnion = {'Series1': pd.Series([1,2,3,4,5],
                            index = ['a', 'b', 'c', 'd', 'e']),
                'Series2': pd.Series([10,20,-10,-50,100],
                            index = ['z', 'y', 'a', 'c', 'e']),
                'Series3': pd.Series([10,20,-10,-50,100],
                            index = ['z', 'y', 'a', 'c', 'e'])}
>>> dFrameUnion = pd.DataFrame(dictForUnion)
>>> dFrameUnion
   Series1  Series2  Series3
a      1.0    -10.0    -10.0 …