Applied Mathematics · Ch 7 — Inferential Statistics
Population and Sample
Population and Sample
A population is the complete group of individuals or objects you want to draw conclusions about — every student enrolled in a school, every chip a factory produces in a week. The population size is simply the count of members in that group. Studying an entire population is rarely practical, so we instead work with a sample — a small, manageable subset of the population — and use it to make inferences about the whole. This is sampling: choosing a subset without having to examine every individual, then extending what we learn to the full population. A vaccine maker testing side-effects on all of India before rollout is impossible; testing on a carefully chosen group from each demographic and extrapolating is not.
Sampling methods split into two families. Probability sampling uses randomization so every member of the population has an equal chance of selection — this keeps the process unbiased. Non-probability sampling skips randomization, which risks a biased result where some elements are systematically more or less likely to appear.
Two probability methods matter here. In simple random sampling, every individual is chosen purely by chance — from a population of size , a sample of size is one of equally likely subsets, and each has the same probability of being picked. In systematic random sampling, you fix a starting point and then pick every -th member thereafter — for instance, starting at position 2 and selecting every 3rd student from a numbered class list.
A representative sample genuinely mirrors the population's characteristics — it needs a random-selection component and a size large enough to capture the population's variability. When a sample fails to reflect the population's true parameter, it's called an unrepresentative (or biased) sample, and the resulting distortion is selection bias — dialling only listed telephone numbers for a survey, for example, silently excludes every unlisted household.
Sampling is called unbiased when every member has an equal chance of selection (as in simple random or systematic sampling) and biased when the process systematically favours certain outcomes — such as convenience sampling, where only easily reachable individuals are included. Other recognized forms of bias include voluntary response bias (only those who choose to participate respond), undercoverage (parts of the population are underrepresented), and response bias (survey conditions distort answers).
Even a properly random sample won't exactly match the population — the gap between a population parameter and the corresponding sample statistic is the sampling error:
…