Informatics Practices · Ch 2 — Data Handling using Pandas – I
Summary
Summary
This chapter covered the two core Pandas data structures and how to move data between them and CSV files.
- NumPy, Pandas, and Matplotlib are the key Python libraries for scientific and analytical work; Pandas is installed with
pip install pandas. - A Series is a one-dimensional array of values, each with an associated data label (its index).
- A Series can be accessed by indexing (a single element) or slicing (a range of elements) — either by positional index (integer position, starting at 0) or by labelled index (a user-defined label).
- With positional slicing, the value at the end position is excluded; with labelled slicing, the value at the end label is included.
- Basic arithmetic operations on Series can be written using either operators (
+,-,*,/) or the equivalent methods (.add(),.sub(),.mul(),.div()). - When two Series are combined, Pandas aligns values by index label first; a label missing from one side produces
NaNin the result rather than raising an error. - A DataFrame is a two-dimensional labelled data structure, like a spreadsheet — it has both a row index and a column index.
- Creating a DataFrame from a dictionary makes each key a column label; a DataFrame can be thought of as a dictionary of Series that all share the same row index.
pandas.read_csv()loads data from a file on disk into a DataFrame;DataFrame.to_csv()writes a DataFrame back out to a text/CSV file.DataFrame.Treturns the transpose (rows and columns swapped).DataFrame.loc[]is used for label-based indexing and slicing of rows (and optionally columns) — every label used must actually exist in the index, or Pandas raises aKeyError.DataFrame.append()merges two DataFrames by stacking the rows of the second below the first, taking the union of their columns (any column missing from one side is filled withNaN). …