Skip to content

Informatics Practices · Ch 3 — Data Handling using Pandas – II

Calculating Variance

3.2.9

Calculating Variance

Variance measures how spread out a set of numbers is. In Pandas, the .var() method on a DataFrame or Series calculates this spread. It tells you, on average, how far each value deviates from the mean, squared so that positive and negative differences don't cancel each other out.

The textbook defines variance as "the average of squared differences from the mean." This means you take every score, subtract the mean of all scores, square that difference, and then average all those squared values. A higher variance indicates that the data points are more spread out from the mean; a lower variance means they are clustered closer to it.

To compute variance for multiple columns at once, you pass a list of column names inside double square brackets to the DataFrame, then call .var(). The textbook shows this exact syntax:

df[['Maths','Science','S. St','Hindi','Eng']].var()

This returns a Series where each subject name is the index and the corresponding value is its variance. The output from the textbook is:

  • Maths: 15.84
  • Science: 7.11
  • S. St: 9.90
  • Hindi: 9.97
  • English: 11.36

The data type of the entire result is float64. Notice that the variance for Maths is the highest among these subjects, meaning the students' Maths scores varied the most from the average. Science has the lowest variance, indicating the scores were more consistent. …

Table 3.17Output of df[[...]].var() -- variance of the five subject columns
ColumnVariance
Maths15.840909
Science7.113636
S.St9.901515
Hindi9.969697