Skip to content

Informatics Practices · Ch 3 — Data Handling using Pandas – II

Descriptive Statistics

3.2

Descriptive Statistics

Descriptive statistics are the methods we use to get a basic understanding of a dataset — to summarise it rather than look at every single value. In the context of a DataFrame, these methods help us quickly grasp the central tendency, spread, and shape of the data. The textbook introduces the following descriptive statistical measures that can be applied to a DataFrame: max, min, count, sum, mean, median, mode, quartiles, and variance. Each of these is demonstrated using the DataFrame df that was created earlier in the chapter.

  • max and min return the largest and smallest values in each column of the DataFrame, respectively. They give you the range of the data at a glance.
  • count returns the number of non-null (non-missing) values in each column. This is useful for checking how much data is actually present.
  • sum adds up all the values in each column. For numeric columns, this gives a total; for non-numeric columns, it may concatenate strings (depending on the data type).
  • mean calculates the arithmetic average of each numeric column. It is the sum divided by the count.
  • median is the middle value when the data in each column is sorted. If there is an even number of values, it is the average of the two middle values. The median is less affected by extreme values than the mean.
  • mode returns the most frequently occurring value(s) in each column. A column can have multiple modes (if several values appear equally often).
  • quartiles divide the sorted data into four equal parts. The first quartile (Q1) is the 25th percentile, the second quartile (Q2) is the median (50th percentile), and the third quartile (Q3) is the 75th percentile. These help understand the spread of the data.
  • variance measures how far the values in a column are spread out from the mean. A higher variance indicates greater dispersion. …