Sampling and Variables in Economics
You already do sampling without knowing it. When you taste one spoonful of dal to judge whether the whole pot needs salt, you are sampling. You assume that spoonful represents the entire pot. If you stirred well, it probably does. If you didn't, you might oversalt everything.
Economics faces the same problem — but with millions of households, firms, or prices. We cannot ask every single person in India how much they earn, or measure every single price in every single market. So we take a sample, and from that sample we try to learn about the whole population.
What is a Population? What is a Sample?
Population is the entire group you want to study. If you want to know the average monthly expenditure of households in Delhi, the population is all households in Delhi. If you want the average yield of wheat in Punjab, the population is every wheat field in Punjab.
Sample is a subset of that population — a few hundred households, or a few dozen fields — that you actually collect data from.
The goal of sampling is to estimate a characteristic of the population (called a parameter) using a statistic calculated from the sample. The sample must be representative — it should mirror the population in the ways that matter.
If your sample of Delhi households only includes people from South Delhi, it will overestimate average expenditure. That is sampling bias — the sample does not look like the population.
Why Sampling Matters in Economics
Governments and businesses make decisions worth crores based on sample data. The National Sample Survey Office (NSSO) does not interview every Indian household — it surveys a carefully chosen sample of about 1.2 lakh households. From that, it estimates national unemployment, consumption, and poverty rates.
A badly chosen sample gives wrong estimates. Wrong estimates lead to bad policy. That is why sampling design — how you pick your sample — is a serious subject.
Types of Sampling
1. Random Sampling (Simple Random Sampling)
Every member of the population has an equal chance of being selected. Like drawing names from a hat.
Advantage: Eliminates bias. The sample is likely to be representative.
Disadvantage: Hard to do in practice — you need a complete list of the population (a sampling frame), which often does not exist.
2. Stratified Sampling
Divide the population into groups (strata) that are internally similar but different from each other — for example, urban and rural households, or income brackets. Then take a random sample from each group.
Why use it? If you know that urban and rural spending patterns are very different, a simple random sample might accidentally get too many urban households. Stratified sampling ensures each group is represented in proportion to its size in the population.
3. Cluster Sampling
Divide the population into clusters (like villages or city blocks), randomly pick a few clusters, and then survey everyone in those clusters.
Why use it? Cheaper and easier — you only travel to a few locations. But less precise than stratified sampling because people within a cluster tend to be similar.
4. Convenience Sampling
You survey whoever is easiest to reach — people in a mall, students in your school.
This is almost always biased. The people easiest to reach are rarely representative of the whole population. Never trust conclusions from convenience samples in serious economics.
What is a Variable?
A variable is any characteristic that can take different values for different members of the population. Age, income, gender, education level, price of wheat — all are variables.
In any study, you have:
- Dependent variable: The outcome you are trying to explain or predict.
- Independent variable: The factor you think influences the dependent variable.
Example: You want to study how education affects income.
- Dependent variable: Income (what you want to explain)
- Independent variable: Education level (what you think causes the change)
Types of Variables
Quantitative Variables (Numeric)
Can be measured in numbers.
- Discrete: Can only take whole number values. Number of children in a family (0, 1, 2, ...). You cannot have 2.5 children.
- Continuous: Can take any value within a range. Height, weight, temperature, income in rupees (₹25,000.50 is possible).
Qualitative Variables (Categorical)
Cannot be measured numerically — they describe categories.
- Nominal: Categories with no natural order. Gender (male/female), religion, type of crop grown.
- Ordinal: Categories that can be ordered. Education level (primary, secondary, graduate, postgraduate). You know graduate is higher than secondary, but you cannot say by how much.
In economics, many variables that seem quantitative are treated as categorical for analysis. Income brackets (₹0–₹2.5 lakh, ₹2.5–₹5 lakh, etc.) are ordinal, not continuous, because you lose the exact value.
How Sampling and Variables Connect
When you design a sample, you must decide:
- Which variables to measure — what questions to ask.
- How to measure them — is income self-reported? Is it annual or monthly?
- Which sampling method will give you reliable data on those variables.
If your variable is "monthly household expenditure" and you only sample during Diwali week, your estimates will be too high — that is a measurement error caused by bad timing, not bad sampling.
A Simple Example
Suppose you want to estimate the average monthly pocket money of Class 12 students in your city.
- Population: All Class 12 students in the city.
- Sample: 200 students chosen randomly from 10 schools (stratified by school type: government, private, and Kendriya Vidyalaya).
- Variable: Monthly pocket money in rupees (continuous, quantitative).
- Dependent variable: Pocket money.
- Independent variable: Type of school (categorical, nominal).
You collect the data, calculate the sample average — say ₹1,200 — and use that as your estimate for the population average. The accuracy of that estimate depends entirely on how well your sample represents the population.
The key idea: A sample is only useful if it is representative. A representative sample requires careful design — random selection, proper stratification, and unbiased measurement of variables. Without that, your numbers are meaningless, no matter how many decimal places you report.