Skip to content

Informatics Practices · Ch 3 — Data Handling using Pandas – II

Summary

Summary

  • Descriptive statistics are used to quantitatively summarise a dataset.
  • Pandas provides many statistical functions for analysing data, including max(), min(), mean(), median(), mode(), std(), and var().
  • Sorting arranges data in a specified order — ascending or descending.
  • The row/column labels of a DataFrame can be changed, a process called altering the index; reset_index() and set_index() are used for this.
  • Missing values are a hindrance in data analysis and must be handled properly.
  • There are two main strategies for handling missing data: remove the row (or column) with the missing value entirely, or replace it with an appropriate value (zero, the average, etc.).
  • Changing the structure of a DataFrame is called reshaping; Pandas provides pivot() and pivot_table() for this. …