Skip to content

Informatics Practices · Ch 2 — Data Handling using Pandas – I

Attributes of DataFrames

2.3.6

Attributes of DataFrames

>>> ForestArea = {
        'Assam' : pd.Series([78438, 2797, 10192, 15116],
                  index = ['GeoArea', 'VeryDense', 'ModeratelyDense', 'OpenForest']),
        'Kerala' : pd.Series([38852, 1663, 9407, 9251],
                  index = ['GeoArea', 'VeryDense', 'ModeratelyDense', 'OpenForest']),
        'Delhi' : pd.Series([1483, 6.72, 56.24, 129.45],
                  index = ['GeoArea', 'VeryDense', 'ModeratelyDense', 'OpenForest'])}

>>> ForestAreaDF = pd.DataFrame(ForestArea)
>>> ForestAreaDF
                 Assam  Kerala    Delhi
GeoArea          78438   38852  1483.00
VeryDense         2797    1663     6.72
ModeratelyDense  10192    9407    56.24
OpenForest       15116    9251   129.45

Attributes of DataFrames

Just like Series, a DataFrame has certain built-in properties called attributes that give you quick access to information about the DataFrame — its structure, labels, data types, and values. You access an attribute by writing the DataFrame name, a dot, and then the attribute name (e.g., ForestAreaDF.index).

The textbook uses a real-world example: a DataFrame called ForestAreaDF built from the State of Forest Report 2017 (Forest Survey of India). It contains data for three states — Assam, Kerala, and Delhi — across four row categories: geographical area, area under very dense forest, area under moderately dense forest, and area under open forest (all in square kilometres).

Here is how the DataFrame looks when printed:

             Assam    Kerala    Delhi
GeoArea       78438     38852   1483.00
VeryDense      2797      1663      6.72
ModeratelyDense 10192     9407     56.24
OpenForest    15116      9251    129.45

Notice that Delhi’s values are floats (with decimals) while Assam and Kerala are integers — this matters when you check data types.


DataFrame.index

Displays the row labels (the index) of the DataFrame.

>>> ForestAreaDF.index
Index(['GeoArea', 'VeryDense', 'ModeratelyDense', 'OpenForest'], dtype='object')

This tells you the four row names. The index is an Index object, and its dtype is 'object' because the labels are strings.


DataFrame.columns

Displays the column labels.

>>> ForestAreaDF.columns
Index(['Assam', 'Kerala', 'Delhi'], dtype='object')

Here, the columns are the three state names.


DataFrame.dtypes

Shows the data type of each column in the DataFrame. This is extremely useful when you want to check whether a column contains integers, floats, or strings.

>>> ForestAreaDF.dtypes
Assam      int64
Kerala     int64
Delhi     float64
dtype: object

Assam and Kerala are int64 (whole numbers), while Delhi is float64 (decimal numbers). The overall output is itself a Series with dtype: object.


DataFrame.values

Returns a NumPy ndarray containing all the values in the DataFrame — without any row or column labels. This is a pure numerical array.

>>> ForestAreaDF.values
array([[7.8438e+04, 3.8852e+04, 1.4830e+03],
       [2.7970e+03, 1.6630e+03, 6.7200e+00],
       [1.0192e+04, 9.4070e+03, 5.6240e+01],
       [1.5116e+04, 9.2510e+03, 1.2945e+02]])

The numbers are shown in scientific notation (e.g., 7.8438e+04 means 78,438). The array is two-dimensional, matching the shape of the DataFrame.


DataFrame.shape

Returns a tuple (number_of_rows, number_of_columns).

>>> ForestAreaDF.shape
(4, 3)

This tells you the DataFrame has 4 rows and 3 columns.


DataFrame.size

Returns the total number of elements (values) in the DataFrame — that is, rows × columns.

>>> ForestAreaDF.size
12

Since 4 rows × 3 columns = 12 values.

Watch out

Do not confuse size with shape. shape gives the dimensions as a tuple; size gives the total count of elements.


DataFrame.T

Transposes the DataFrame — rows become columns and columns become rows.

>>> ForestAreaDF.T
              GeoArea  VeryDense  ModeratelyDense  OpenForest
Assam         78438.0    2797.00         10192.00     15116.00
Kerala        38852.0    1663.00          9407.00      9251.00
Delhi          1483.0       6.72            56.24       129.45

After transposing, the original row labels (GeoArea, VeryDense, etc.) become column headers, and the original column labels (Assam, Kerala, Delhi) become row indices. Notice that all values are now shown as floats because the DataFrame had mixed integer and float columns — transposing converts everything to a common type.


DataFrame.head(n)

Displays the first n rows of the DataFrame. If you don't pass a number, it defaults to showing the first 5 rows.

>>> ForestAreaDF.head(2)
             Assam  Kerala    Delhi
GeoArea      78438   38852  1483.00
VeryDense     2797    1663     6.72

Here, head(2) shows only the first two rows. This is very handy when you have a large DataFrame and just want a quick peek at the top.


DataFrame.tail(n)

Displays the last n rows of the DataFrame. Default is also 5 rows.

>>> ForestAreaDF.tail(2)
                 Assam  Kerala    Delhi
ModeratelyDense  10192    9407    56.24 …
Table 2.4Some Attributes of Pandas DataFrame
Attribute NamePurposeExample
DataFrame.indexto display row labels>>> ForestAreaDF.index
Index(['GeoArea', 'VeryDense', 'ModeratelyDense', 'OpenForest'], dtype ='object')
DataFrame.columnsto display column labels>>> ForestAreaDF.columns
Index(['Assam', 'Kerala', 'Delhi'], dtype='object')
DataFrame.dtypesto display data type of each column in the DataFrame>>> ForestAreaDF.dtypes
Assam int64
Kerala int64
Delhi float64
dtype: object
DataFrame.valuesto display a NumPy ndarray having all the values in the DataFrame, without the axes labels>>> ForestAreaDF.values
array([[7.8438e+04, 3.8852e+04, 1.4830e+03],
       [2.7970e+03, 1.6630e+03, 6.7200e+00],
       [1.0192e+04, 9.4070e+03, 5.6240e+01],
       [1.5116e+04, 9.2510e+03, 1.2945e+02]])
DataFrame.shapeto display a tuple representing the dimensionality of the DataFrame>>> ForestAreaDF.shape
(4, 3)
It means ForestAreaDF has 4 rows and 3 columns.
DataFrame.sizeto display a tuple representing the dimensionality of the DataFrame>>> ForestAreaDF.size
12
This means the ForestAreaDF has 12 values in it.
DataFrame.Tto transpose the DataFrame. Means, row indices and column labels of the DataFrame replace each other's position>>> ForestAreaDF.T
      GeoArea VeryDense ModeratelyDense OpenForest
Assam 78438.0 2797.00 10192.00 15116.00
Kerala38852.0 1663.00 9407.00 9251.00
Delhi 1483.0 6.72 56.24 129.45
DataFrame.head(n)to display the first n rows in the DataFrame>>> ForestAreaDF.head(2)
              Assam Kerala Delhi
GeoArea 78438 38852 1483.00
VeryDense 2797 1663 6.72
displays the first 2 rows of the DataFrame ForestAreaDF.If the parameter n is not specified by default it gives the first 5 rows of the DataFrame.