Informatics Practices · Ch 2 — Data Handling using Pandas – I
Importing a CSV file to a DataFrame
Importing a CSV file to a DataFrame
A CSV (Comma Separated Values) file stores tabular data as plain text, where each line is a row and values are separated by commas. Pandas provides the read_csv() function to load such a file directly into a DataFrame. This is the most common way to bring external data into your Python environment for analysis.
Consider a file named ResultData.csv stored at the path C:/NCERT. The file contains the following data:
| RollNo | Name | Eco | Maths |
|---|---|---|---|
| 1 | Arnab | 18 | 57 |
| 2 | Kritika | 23 | 45 |
| 3 | Divyam | 51 | 37 |
| 4 | Vivaan | 40 | 60 |
| 5 | Aaroosh | 18 | 27 |
To load this into a DataFrame named marks, you write:
>>> marks = pd.read_csv("C:/NCERT/ResultData.csv", sep=",", header=0)
>>> marks
RollNo Name Eco Maths
0 1 Arnab 18 57
1 2 Kritika 23 45
2 3 Divyam 51 37
3 4 Vivaan 40 60
4 5 Aaroosh 18 27
Notice that the DataFrame automatically assigns integer row indices (0, 1, 2, …) and uses the first row of the CSV as column names.
Parameters of read_csv()
-
First parameter (file path): The string
"C:/NCERT/ResultData.csv"gives the full location of the CSV file. You must include the path so Python can find the file on your system. -
sep: This parameter tells pandas what character separates the values in the file. Common separators are comma (,), semicolon (;), or tab (\t). The default value forsepis a space. In the example,sep=","is used because the file is comma-separated. -
header: This parameter specifies which row in the file should be used as the column names for the DataFrame. It also marks the row from which the actual data begins.header=0means the first line (row 0) of the file is taken as column names, and the data starts from the second line onward.- By default,
header=0is assumed, so you can omit it if the first row of your CSV contains the column headers.
Using the names parameter to supply custom column labels
If your CSV file does not have a header row, or if you want to assign different column names than those in the file, you can use the names parameter. When names is provided, pandas ignores any header in the file and uses the list you supply.
For example, suppose you have a file ResultData1.csv that contains the same data but without column names. You can create a DataFrame with your own labels:
>>> marks1 = pd.read_csv("C:/NCERT/ResultData1.csv", sep=",", names=['RNo', 'StudentName', 'Sub1', 'Sub2'])
>>> marks1
RNo StudentName Sub1 Sub2
0 1 Arnab 18 57
1 2 Kritika 23 45
2 3 Divyam 51 37
3 4 Vivaan 40 60 …
| RollNo | Name | Eco | Maths |
|---|---|---|---|
| 1 | Arnab | 18 | 57 |
| 2 | Kritika | 23 | 45 |
| 3 | Divyam | 51 | 37 |
| RNo | StudentName | Sub1 | Sub2 |
|---|---|---|---|
| 1 | Arnab | 18 | 57 |
| 2 | Kritika | 23 | 45 |
| 3 | Divyam | 51 | 37 |