Skip to content

Mathematics · Ch 14 — Probability Distributions

Random variables

14.1

Random variables

A random variable formalises an idea we already use informally: instead of listing every raw outcome of a random experiment, we are usually interested in a single number derived from that outcome — the number of heads when two coins are tossed, the sum shown by two dice, or the number of defective items in a sample. Formally, a random variable is a real-valued function defined on the sample space SS of a random experiment: X:S→RX : S \to R. Its domain is the sample space and its co-domain is the set of real numbers. We abbreviate 'random variable' as r.v., and usually denote it by a capital letter such as XX, YY or ZZ; a specific value it takes is written with the corresponding small letter xx, and the set of every outcome that produces that value is the event [X=x][X = x].

Three quick illustrations show the idea. Throwing two dice has 36 raw outcomes, but if only the sum of the two numbers matters, there are just 11 distinct values, from 2 to 12. Tossing a coin 10 times has 210=10242^{10} = 1024 raw outcomes, but the number of heads among the 10 tosses takes only 11 distinct values, from 0 to 10. Choosing 4 items at random from a lot of 20 that contains 6 defectives has many raw outcomes, but the number of defective items among the four chosen takes only 5 distinct values, from 0 to 4. In every case there is a definite rule assigning a unique number to each outcome, and because that number changes from outcome to outcome it is genuinely a variable — a random variable, since it is derived from the outcomes of a random experiment.

A fully worked example makes the mechanics explicit. Suppose three seeds are sown and we record, for each, whether it germinates (Y) or not (N). The sample space has 23=82^3 = 8 outcomes: S={YYY,YYN,YNY,NYY,YNN,NYN,NNY,NNN}S = \{YYY, YYN, YNY, NYY, YNN, NYN, NNY, NNN\}. Let XX count how many times Y appears in an outcome. Then X(YYY)=3X(YYY) = 3; X(YYN)=X(YNY)=X(NYY)=2X(YYN) = X(YNY) = X(NYY) = 2; X(YNN)=X(NYN)=X(NNY)=1X(YNN) = X(NYN) = X(NNY) = 1; and X(NNN)=0X(NNN) = 0. So XX takes exactly four possible values, {0,1,2,3}\{0, 1, 2, 3\} — this set is called the range of XX. The four events defined by these values are [X=0]={NNN}[X=0] = \{NNN\}, [X=1]={YNN,NYN,NNY}[X=1] = \{YNN, NYN, NNY\}, [X=2]={YYN,YNY,NYY}[X=2] = \{YYN, YNY, NYY\} and [X=3]={YYY}[X=3] = \{YYY\}.

A sample space need not be finite for a random variable to be well defined on it. Consider tossing a coin repeatedly until a head appears for the first time. The sample space is the unending list S={H,TH,TTH,TTTH,…}S = \{H, TH, TTH, TTTH, \ldots\}, which is countably infinite. If XX is the number of tosses needed to get the first head, then XX can equal any positive integer, so its range {1,2,3,…}\{1, 2, 3, \ldots\} is also countably infinite. This example is picked up again once probability mass functions are introduced (section 7.3.1).