Informatics Practices · Ch 2 — Data Handling using Pandas – I
DataFrame
DataFrame
A DataFrame is what you reach for when a single Series — one column of data — is not enough. Real-world data is tabular: a mark sheet with student names, subjects, and marks; a restaurant menu with item names, prices, and availability; a train reservation chart with seat numbers, passenger names, and berth types. Pandas handles all of this with a DataFrame.
A DataFrame is a two-dimensional, labelled data structure. Think of it as a table in a MySQL database or a spreadsheet in Excel. It has both rows and columns, and crucially, it has a row index (the labels for each row) and a column index (the labels for each column). This double-indexing is what makes it a labelled structure.
One important feature: unlike a NumPy array where every element must be of the same data type, a DataFrame can hold columns of different types. One column might contain integers (like roll numbers), another strings (like student names), and a third floats (like percentages) or even boolean values (like pass/fail status). This flexibility mirrors the way data is stored in a database table.
The textbook illustrates this with a small table showing the geographical area and area under very dense forests for three Indian states. The row indexes are 1, 2, and 3 (for Assam, Delhi, and Kerala), and the column indexes are "State", "Geographical Area (sq Km)", and "Area under Very Dense Forests (sq Km)". This is a classic DataFrame: rows are observations (states), columns are variables (area, forest cover). …
| State | Geographical Area (sq Km) | Area under Very Dense Forests (sq Km) | |
|---|---|---|---|
| 1 | Assam | 78438 | 2797 |