Skip to content

Business Mathematics and Basic Statistics · Ch 16 — Bivariate Statistics — Correlation and Regression

Idea of Bivariate Data and Covariance

1

Idea of Bivariate Data and Covariance

So far, every measure this course has looked at — mean, mode, median, dispersion — has described a single variable at a time: marks in one subject, wages of one factory's workers, ages of one clinic's patients. Many real business questions, though, are really about the relationship between two variables measured on the same set of items: does higher advertising spending go with higher sales? Does a higher price go with lower quantity demanded? Data of this kind, where every unit in the sample carries a pair of values (x,y)(x, y) instead of one, is called bivariate data.

The first natural question about a bivariate data set is whether the two variables move together, and if so, in which direction. Covariance answers this by averaging the product of each variable's deviation from its own mean.

Note

Covariance

For nn paired observations (xi,yi)(x_i, y_i),

Cov(x,y)=1n∑(xi−xˉ)(yi−yˉ)=1n∑xiyi−xˉyˉ\text{Cov}(x,y) = \dfrac{1}{n}\sum (x_i-\bar{x})(y_i-\bar{y}) = \dfrac{1}{n}\sum x_iy_i - \bar{x}\bar{y}

Both forms always give the same value — the second (the 'direct' form, using raw products xiyix_iy_i) is usually faster by hand, while the first (the 'deviation' form) is more useful for understanding why covariance works: when xx and yy tend to be above their means together (or below their means together), the products (xi−xˉ)(yi−yˉ)(x_i-\bar{x})(y_i-\bar{y}) are mostly positive, driving the covariance positive; when one tends to be above its mean while the other is below, the products are mostly negative.

Sign, but not size, is meaningful on its own. A positive covariance means xx and yy tend to move in the same direction; a negative covariance means they tend to move in opposite directions; a covariance near zero suggests little linear association. However, covariance's numerical size depends on the units the two variables are measured in (₹ vs kg vs years), so it cannot by itself say how strong the association is — that is exactly the gap the correlation coefficient of the next section fills, by dividing covariance down to a fixed, unit-free scale.

Understanding bivariate data and computing covariance is the foundation on which every other idea in this chapter — correlation, rank correlation, and regression — is built, and this two-variable way of thinking is central to the West Bengal HS commerce mathematics and statistics syllabus's treatment of Business Mathematics and Basic Statistics.

Definition 1Bivariate Data

Data in which each unit of observation carries a pair of values (x,y)(x, y) for two variables measured together, e.g. advertising expenditure and sales for each of several months.

Definition 2Covariance

A measure of how two variables vary together: Cov(x,y)=1n∑(xi−xˉ)(yi−yˉ)\text{Cov}(x,y) = \dfrac{1}{n}\sum(x_i-\bar{x})(y_i-\bar{y}). Positive covariance indicates the variables tend to move in the same direction; negative covariance indicates they tend to move in opposite directions.