Skip to content

Mathematics and Statistics · Ch 13 — Bivariate Frequency Distribution and Chi-Square Statistic

Bivariate Data and the Bivariate Frequency Distribution

1

Bivariate Data and the Bivariate Frequency Distribution

Every measure met so far in this course — mean, median, quartiles, standard deviation, skewness — describes just one variable at a time: the marks of a class in a single subject, or the daily wages of one factory's workers. Very often, though, a business question is really about two variables measured together on the same units — the marks a student scores in Mathematics and in Statistics, or the age and the monthly income of each customer. Data in which every unit of observation carries a pair of values (x,y)(x, y) is called bivariate data ('bi' = two, 'variate' = variable).

When only a handful of pairs are recorded, they can simply be listed. But when there are many pairs, a raw list becomes unreadable, so — exactly as a single variable's many values are compressed into a frequency distribution — a bivariate data set is compressed into a bivariate frequency distribution, displayed as a two-way table.

Note

Two-Way (Bivariate) Frequency Table

The distinct values or class-intervals of one variable, xx, label the rows; the distinct values or class-intervals of the other variable, yy, label the columns. Each cell holds the frequency fijf_{ij} — the number of pairs whose xx-value falls in row ii and whose yy-value falls in column jj. The sum of all cell frequencies is the total number of pairs, N=∑i∑jfijN = \sum_i \sum_j f_{ij}.

To build the table, each observed pair is read off and a tally mark placed in the single cell where its row-class and column-class cross; the tallies are then counted into frequencies. A small worked layout with xx taking values 1,2,31, 2, 3 and yy taking values 1,2,31, 2, 3 looks like this:

x\yx \backslash yy=1y=1y=2y=2y=3y=3Row total
x=1x=12204
x=2x=22215
x=3x=30123
Column total45312

Here the cell in row x=2x=2, column y=1y=1 holds the frequency 2, meaning two of the twelve pairs were (2,1)(2, 1). The bottom-right corner, 12, is the grand total NN and must equal both the sum of the row totals (4+5+34+5+3) and the sum of the column totals (4+5+34+5+3) — a quick built-in arithmetic check on the whole table.

Organising paired data into a bivariate frequency distribution is the starting point of the Maharashtra Board Std XI commerce mathematics and statistics treatment of two-variable data, and the two-way frequency table is the same standard device used to summarise bivariate data in statistics courses across every board and stream.

Definition 1Bivariate Data

Data in which each unit of observation carries a pair of values (x,y)(x, y) for two variables measured together, e.g. the Mathematics mark and the Statistics mark of each student in a class.

Definition 2Bivariate Frequency Distribution

A summary of many (x,y)(x, y) pairs in a two-way table whose rows are the values/classes of xx, whose columns are the values/classes of yy, and whose cell fijf_{ij} counts the pairs falling in row ii and column jj; the total of all cells is NN, the number of pairs.