Sampling and Variables in Economics
You already do sampling without knowing it. When you taste one spoonful of dal to judge whether the whole pot needs salt, you are sampling. You assume that spoonful represents the entire pot. If you stirred well, it probably does. If you didn't, you might oversalt everything.
Economics faces the same problem — but with millions of households, firms, or prices. We cannot ask every single person in India how much they earn, or measure every single price in every single market. So we take a sample, and from that sample we try to learn about the whole population.
What is a Population? What is a Sample?
Population is the entire group you want to study. If you want to know the average monthly expenditure of households in Delhi, the population is all households in Delhi. If you want the average yield of wheat in Punjab, the population is every wheat field in Punjab.
Sample is a subset of that population — a few hundred households, or a few dozen fields — that you actually collect data from.
The goal of sampling is to estimate a characteristic of the population (called a parameter) using a statistic calculated from the sample. The sample must be representative — it should mirror the population in the ways that matter.
If your sample of Delhi households only includes people from South Delhi, it will overestimate average expenditure. That is sampling bias — the sample does not look like the population.
Why Sampling Matters in Economics
Governments and businesses make decisions worth crores based on sample data. The National Sample Survey Office (NSSO) does not interview every Indian household — it surveys a carefully chosen sample of about 1.2 lakh households. From that, it estimates national unemployment, consumption, and poverty rates.
A badly chosen sample gives wrong estimates. Wrong estimates lead to bad policy. That is why sampling design — how you pick your sample — is a serious subject.
Types of Sampling
1. Random Sampling (Simple Random Sampling)
Every member of the population has an equal chance of being selected. Like drawing names from a hat.
Advantage: Eliminates bias. The sample is likely to be representative.
Disadvantage: Hard to do in practice — you need a complete list of the population (a sampling frame), which often does not exist.
2. Stratified Sampling
Divide the population into groups (strata) that are internally similar but different from each other — for example, urban and rural households, or income brackets. Then take a random sample from each group.
Why use it? If you know that urban and rural spending patterns are very different, a simple random sample might accidentally get too many urban households. Stratified sampling ensures each group is represented in proportion to its size in the population.
3. Cluster Sampling
Divide the population into clusters (like villages or city blocks), randomly pick a few clusters, and then survey everyone in those clusters.
Why use it? Cheaper and easier — you only travel to a few locations. But less precise than stratified sampling because people within a cluster tend to be similar.
4. Convenience Sampling
You survey whoever is easiest to reach — people in a mall, students in your school.
This is almost always biased. The people easiest to reach are rarely representative of the whole population. Never trust conclusions from convenience samples in serious economics.
What is a Variable?
A variable is any characteristic that can take different values for different members of the population. Age, income, gender, education level, price of wheat — all are variables.
In any study, you have:
- Dependent variable: The outcome you are trying to explain or predict.
- Independent variable: The factor you think influences the dependent variable.
Example: You want to study how education affects income.
- Dependent variable: Income (what you want to explain)
- Independent variable: Education level (what you think causes the change)
Types of Variables
Quantitative Variables (Numeric)
Can be measured in numbers.
- Discrete: Can only take whole number values. Number of children in a family (0, 1, 2, ...). You cannot have 2.5 children.
- Continuous: Can take any value within a range. Height, weight, temperature, income in rupees (₹25,000.50 is possible).
Qualitative Variables (Categorical)
Cannot be measured numerically — they describe categories. …