Informatics Practices · Ch 2 — Data Handling using Pandas – I
Importing and Exporting Data between CSV Files and DataFrames
Importing and Exporting Data between CSV Files and DataFrames
We often have data stored in CSV (Comma Separated Values) files — a very common format for spreadsheets and databases. A CSV file stores tabular data as plain text, where each line is a row and values within a row are separated by commas. Pandas makes it straightforward to bring this data into a DataFrame for analysis, and also to save a DataFrame back into a CSV file for sharing or storage.
Importing Data from a CSV File
To create a DataFrame from a CSV file, you use the read_csv() function. The basic syntax is:
import pandas as pd
df = pd.read_csv('filename.csv')
Here, 'filename.csv' is the path to your CSV file. The function reads the file, automatically treats the first row as column headers (unless told otherwise), and returns a DataFrame. This is the most common way to load external data into a Pandas program.
The name "CSV" stands for Comma Separated Values. The file's extension is .csv.
Exporting Data to a CSV File
Once you have processed or created a DataFrame, you can save it as a CSV file using the to_csv() method. The syntax is:
df.to_csv('output_filename.csv')
This writes the DataFrame df into a new file named output_filename.csv. By default, it includes the row index (the leftmost column of the DataFrame) and the column headers. If you want to exclude the index, you can add the parameter index=False:
df.to_csv('output_filename.csv', index=False)
This is useful when you want a clean CSV file without the extra index column.