Skip to content

Informatics Practices · Ch 2 — Data Handling using Pandas – I

Summary

Summary

This chapter covered the two core Pandas data structures and how to move data between them and CSV files.

  • NumPy, Pandas, and Matplotlib are the key Python libraries for scientific and analytical work; Pandas is installed with pip install pandas.
  • A Series is a one-dimensional array of values, each with an associated data label (its index).
  • A Series can be accessed by indexing (a single element) or slicing (a range of elements) — either by positional index (integer position, starting at 0) or by labelled index (a user-defined label).
  • With positional slicing, the value at the end position is excluded; with labelled slicing, the value at the end label is included.
  • Basic arithmetic operations on Series can be written using either operators (+, -, *, /) or the equivalent methods (.add(), .sub(), .mul(), .div()).
  • When two Series are combined, Pandas aligns values by index label first; a label missing from one side produces NaN in the result rather than raising an error.
  • A DataFrame is a two-dimensional labelled data structure, like a spreadsheet — it has both a row index and a column index.
  • Creating a DataFrame from a dictionary makes each key a column label; a DataFrame can be thought of as a dictionary of Series that all share the same row index.
  • pandas.read_csv() loads data from a file on disk into a DataFrame; DataFrame.to_csv() writes a DataFrame back out to a text/CSV file.
  • DataFrame.T returns the transpose (rows and columns swapped).
  • DataFrame.loc[] is used for label-based indexing and slicing of rows (and optionally columns) — every label used must actually exist in the index, or Pandas raises a KeyError.
  • DataFrame.append() merges two DataFrames by stacking the rows of the second below the first, taking the union of their columns (any column missing from one side is filled with NaN). …