Skip to content

Informatics Practices · Ch 3 — Data Handling using Pandas – II

Calculating Standard Deviation

3.2.10

Calculating Standard Deviation

The standard deviation tells you how spread out the numbers in a dataset are from their average (mean). A small standard deviation means the values are clustered close to the mean; a large one means they are widely scattered. In Pandas, you can calculate it for any numeric column or DataFrame using the .std() method.

Using DataFrame.std()

The method DataFrame.std() returns the standard deviation of the values in each column. Internally, it first calculates the variance (the average of the squared differences from the mean) and then takes its square root.

For example, if you have a DataFrame df with columns for student marks in different subjects, you can get the standard deviation for a selection of those columns:

df[['Maths','Science','S. St','Hindi','Eng']].std()

The output will be a Series showing the standard deviation for each subject. For the textbook's example data, the result is shown below.

This tells you, for instance, that Maths marks have a standard deviation of about 3.98, while Science marks are less spread out, with a standard deviation of about 2.67.

The describe() Function — A One-Stop Summary

Instead of calculating each statistic separately, you can use DataFrame.describe() to get a whole set of descriptive statistical values in a single command. This function is extremely useful for getting a quick overview of your data.

When you call df.describe(), it automatically includes:

  • count — the number of non-null values in each column
  • mean — the average value
  • std — the standard deviation
  • min — the smallest value
  • 25% — the first quartile (25th percentile)
  • 50% — the median (50th percentile)
  • 75% — the third quartile (75th percentile)
  • max — the largest value

For the textbook's example DataFrame, df.describe() produces the following table. …

Table 3.18Output of df[[...]].std() -- standard deviation of the five subject columns
ColumnStd. Deviation
Maths3.980064
Science2.667140
S.St3.146667
Hindi3.157483
Table 3.19Output of df.describe()
UTMathsScienceS.StHindiEng
count12.00000012.00000012.0000012.00000012.00000012.000000
mean2.00000019.25000019.7500020.41666721.83333320.500000
std0.8528033.9800642.667143.1466673.1574833.370999
min1.00000012.00000015.0000015.00000015.00000013.000000
25%1.00000016.50000018.0000018.75000020.75000019.750000
50%2.00000020.50000019.5000020.00000022.50000021.500000