Example 2. Using the step deviation method, calculate the coefficient of correlation between price index (X) and money supply (Y).
| Price index (X) | Money supply in Rs crores (Y) |
|---|---|
| 120 | 1800 |
| 150 | 2000 |
| 190 | 2500 |
| 220 | 2700 |
| 230 | 3000 |
(Take A = 100; h = 10; B = 1700; k = 100.)
Concept understanding — Pearson Correlation
Pearson Correlation: From Intuition to Precision
Imagine you're tracking two things over time — say, the number of ice creams sold at your school canteen and the outside temperature. On hot days, both go up; on cool days, both drop. They seem to move together. That's the basic idea of correlation: a measure of how two variables move in relation to each other.
But not all co-movement is equal. Sometimes one variable goes up while the other goes down — like the price of a good and the quantity demanded (law of demand). Sometimes they seem to have no connection at all — like the number of students in your class and the price of tea in China. Pearson correlation gives us a single number that captures the strength and direction of this linear relationship.
The Precise Meaning
Pearson correlation coefficient, usually denoted by r, measures the linear relationship between two variables X and Y. It answers: If I know how far X is from its average, how far (and in which direction) is Y from its average, on average?
The formula is:
r=∑(Xi−Xˉ)2⋅∑(Yi−Yˉ)2∑(Xi−Xˉ)(Yi−Yˉ)
Where:
- Xi, Yi are individual observations
- Xˉ, Yˉ are the means (averages) of X and Y
- ∑ means "sum over all observations"
The numerator is the covariance — it tells you whether deviations from the mean tend to be in the same direction (positive product) or opposite directions (negative product). The denominator is the product of the standard deviations of X and Y, which scales the result so that r always lies between −1 and +1.
r is unitless and always between −1 and +1:
- r=+1: perfect positive linear relationship (all points lie on an upward-sloping line)
- r=−1: perfect negative linear relationship (all points lie on a downward-sloping line)
- r=0: no linear relationship (but there could still be a non-linear one!)
Why It Matters in Economics
Economics is full of pairs of variables that we suspect move together. Pearson correlation gives us a first, clean check on whether that suspicion holds water.
Example 1: Consumption and Income. Keynes said consumption depends on income. If you plot household consumption against household income for a sample of families, you'd expect a positive r — higher income families tend to consume more. A value close to +0.8 or +0.9 would be strong evidence for that relationship.
Example 2: Price and Quantity Demanded. The law of demand says price and quantity demanded move in opposite directions. A negative r between price and quantity (holding other factors constant) would confirm this. But here's the catch — in real market data, price and quantity are determined simultaneously by supply and demand, so a simple correlation might not show the expected negative sign. That's why economists use more advanced tools (like regression) to isolate the relationship.
Example 3: Investment and Interest Rates. You'd expect a negative correlation — when interest rates are high, borrowing is expensive, so investment falls. But the relationship might be weak (r close to 0) because investment also depends on expectations, technology, and government policy.
Correlation does NOT imply causation. Just because ice cream sales and drowning incidents are positively correlated (both peak in summer) does NOT mean ice cream causes drowning. The common cause is hot weather, which makes people both buy ice cream and go swimming. In economics, this is a constant trap — GDP and money supply are correlated, but which causes which? The answer requires theory, not just correlation.
Visualising It
Draw a scatter plot with X on the horizontal axis and Y on the vertical axis. If the points cluster around a straight line sloping upward, r is positive and strong. If they cluster around a line sloping downward, r is negative and strong. If they form a shapeless cloud, r is near zero.
But here's the nuance: a perfect circle of points has r=0 even though X and Y are clearly related (non-linearly). Pearson correlation only captures linear relationships. Two variables could be perfectly related by a U-shaped curve and still have r=0.
Always plot your data first. A single outlier can dramatically change r, and a non-linear relationship can hide behind a near-zero r. The number alone is never enough — you need to see the shape.
A Quick Worked Example
Suppose you have data on five families:
| Family | Income (₹'000) | Consumption (₹'000) |
|---|---|---|
| A | 10 | 8 |
| B | 20 | 15 |
| C | 30 | 22 |
| D | 40 | 30 |
| E | 50 | 38 |
Mean income Xˉ=30, mean consumption Yˉ=22.6.
Compute deviations and products:
- For family A: (10−30)(8−22.6)=(−20)(−14.6)=292
- For family B: (20−30)(15−22.6)=(−10)(−7.6)=76
- For family C: (30−30)(22−22.6)=(0)(−0.6)=0
- For family D: (40−30)(30−22.6)=(10)(7.4)=74
- For family E: (50−30)(38−22.6)=(20)(15.4)=308
Sum of products = 292+76+0+74+308=750
Sum of squared deviations for income: (−20)2+(−10)2+02+102+202=400+100+0+100+400=1000
Sum of squared deviations for consumption: (−14.6)2+(−7.6)2+(−0.6)2+7.42+15.42=213.16+57.76+0.36+54.76+237.16=563.2
r=1000×563.2750=563200750≈750.47750≈0.999
That's almost perfect positive correlation — consumption rises almost exactly linearly with income in this small sample.
The Bottom Line
Pearson correlation is your first tool for spotting linear relationships in economic data. It's simple, intuitive, and powerful — but limited. Use it to check whether two variables move together, but never to claim that one causes the other. In economics, where everything is connected to everything else, that caution is worth its weight in gold.
The step-deviation method rescales the large price-index and money-supply figures using an assumed origin and common factor before computing Karl Pearson's coefficient, since correlation is unaffected by such a change of origin and scale.
With step deviations U=10X−100 and V=100Y−1700:
r=[N∑U2−(∑U)2][N∑V2−(∑V)2]N∑UV−∑U∑V=434×490455≈0.98
A very high positive correlation between the price index and money supply.
Rescaling with A=100,h=10 and B=1700,k=100 gives step deviations U and V. Then ∑U=41, ∑V=35, ∑UV=378, ∑U2=423, ∑V2=343, so r≈+0.98 — an almost perfect positive relationship between price index and money supply.
Concept first
The step-deviation method keeps large figures manageable. Correlation is unaffected by a change of origin (subtracting A, B) or scale (dividing by h, k), so we may compute r on the coded values U,V and it equals rXY:
U=hX−A,V=kY−B
r=[N∑U2−(∑U)2][N∑V2−(∑V)2]N∑UV−∑U∑V
The working table
A=100, h=10, B=1700, k=100, N=5.
| X | Y | U=10X−100 | V=100Y−1700 | UV | U2 | V2 |
|---|---|---|---|---|---|---|
| 120 | 1800 | 2 | 1 | 2 | 4 | 1 |
| 150 | 2000 | 5 | 3 | 15 | 25 | 9 |
| 190 | 2500 | 9 | 8 | 72 | 81 | 64 |
| 220 | 2700 | 12 | 10 | 120 | 144 | 100 |
| 230 | 3000 | 13 | 13 | 169 | 169 | 169 |
| Σ | 41 | 35 | 378 | 423 | 343 |
Substituting
Numerator=N∑UV−∑U∑V=5(378)−(41)(35)=1890−1435=455
N∑U2−(∑U)2=5(423)−412=2115−1681=434
N∑V2−(∑V)2=5(343)−352=1715−1225=490
r=434×490455=212660455=461.15455=0.9867
Interpretation
r≈+0.98 is very close to +1: as the price index rises, money supply rises almost in lock-step — a strong positive association.
r=434×490455≈+0.98
A very high positive correlation between the price index and money supply.
- JKBOSE Class 11 (Commerce) 2026Set ANNUAL6 marksQ.Explain Karl Pearson's coefficient of correlation and calculate it for the following data:
X 2 4 6 8 10 Y 3 5 7 9 11 (OR)Discuss the uses and limitations of index numbers in economic analysis.›Reveal solutionSolution
MAIN: Karl Pearson's coefficient of correlation (r) measures the degree and direction of linear relationship between two variables; for the given data it works out to r = +1, a perfect positive correlation. OR: index numbers are valuable tools for comparing variables over time and guiding policy, but have specific limitations arising from how they are constructed.
Main Question — Karl Pearson's Coefficient of Correlation
Meaning: Karl Pearson's coefficient of correlation is a mathematical measure of the degree of linear relationship between two quantitative variables X and Y. It is given by:
r = Σ(dx·dy) / √(Σdx² × Σdy²)
where dx = X − X̄ (deviation of X from its mean) and dy = Y − Ȳ (deviation of Y from its mean). The value of r always lies between −1 and +1: +1 indicates perfect positive correlation, −1 indicates perfect negative correlation, and 0 indicates no linear correlation.
Calculation for the given data:
X Y dx = X−X̄ dy = Y−Ȳ dx·dy dx² dy² 2 3 −4 −4 16 16 16 4 5 −2 −2 4 4 4 6 7 0 0 0 0 0 8 9 2 2 4 4 4 10 11 4 4 16 16 16 Total 40 40 40 N = 5
X̄ = ΣX/N = (2+4+6+8+10)/5 = 30/5 = 6
Ȳ = ΣY/N = (3+5+7+9+11)/5 = 35/5 = 7
r = Σdx·dy / √(Σdx² × Σdy²) = 40 / √(40 × 40) = 40 / 40 = 1
Since Y = X + 1 for every pair, X and Y increase together in perfect step, giving a perfect positive correlation.
Or — Uses and Limitations of Index Numbers
Uses of Index Numbers:
- Measuring relative change — index numbers show the percentage change in a variable (prices, production, cost of living) between a base period and a current period.
- Framing economic policy — government uses price indices (WPI, CPI) to frame monetary, fiscal and wage policy.
- Measuring cost of living / inflation — the Consumer Price Index is used to determine changes in the cost of living, and to grant dearness allowance (DA) adjustments to employees.
- Deflating/comparing real values — index numbers help convert nominal values (e.g., nominal GDP, nominal wages) into real values by removing the effect of price changes.
- Studying trends — index numbers help compare and study trends in variables like industrial production or agricultural output over many years.
Limitations of Index Numbers:
-
Based on a sample, not all items — index numbers cannot cover every single commodity/item in the economy, so results are approximate, based only on a representative selection of items.
-
Choice of base year — if the base year is not 'normal' (e.g., affected by a war, famine or abnormal event), comparisons drawn from it may be misleading.
-
Choice and accuracy of weights — different commodities are given different weights based on importance, and improper or outdated weights can distort the index.
-
Limited applicability — a general price index may not reflect the true price experience of a specific individual or region, since consumption patterns differ.
-
Formula-dependent differences — different formulae (Laspeyres, Paasche, Fisher) can give somewhat different index values for the same data, so results are formula-sensitive, not a single absolute truth.
✓Final answerMain — Karl Pearson's r = +1 (perfect positive correlation) for the given data. Or — Index numbers are useful for measuring relative/temporal change, guiding economic policy, measuring the cost of living and deflating nominal values, but are limited by their sample-based coverage, sensitivity to the choice of base year and weights, and formula sensitivity.
- JKBOSE Class 11 (Commerce) 2023Set ANNUAL6 marksQ.What is correlation ? What are its various types ? Discuss.(OR)Calculate mean, median and mode from the following data :
Marks F 0-10 4 10-20 9 20-30 21 30-40 37 40-50 10 50-60 10 60-70 9 ›Reveal solutionSolution
Correlation studies how two variables move together (positively, negatively, or not at all), and is classified by direction, number of variables, and the nature of the relationship; separately (OR), computing mean, median and mode of a grouped frequency distribution gives three distinct but related measures of central tendency — here 35.6, ≈34.32, and ≈33.72 respectively.
Part 1 — Correlation: meaning and types
Correlation refers to the statistical technique used to study and measure the degree and direction of the relationship between two (or more) variables, such that a change in one variable is associated with a change in the other. It does not establish causation, only association.
Types of correlation:
- Based on direction:
- Positive correlation: both variables move in the same direction (e.g., income and expenditure).
- Negative correlation: the variables move in opposite directions (e.g., price and quantity demanded).
- Based on number of variables:
- Simple correlation: relationship between only two variables.
- Partial correlation: relationship between two variables, keeping the effect of other variables constant.
- Multiple correlation: the joint relationship of one variable with two or more other variables simultaneously.
- Based on nature/proportion of change:
- Linear correlation: the ratio of change between the variables is constant throughout (points would lie on a straight line if plotted).
- Non-linear (curvilinear) correlation: the ratio of change between the variables is not constant.
Correlation is commonly measured by Karl Pearson's coefficient of correlation (for linear relationships) or Spearman's Rank Correlation (for ranked/ordinal data).
Part 2 (OR) — Mean, Median, Mode of the given data
Data: Marks (class) and Frequency (F):
Marks F Mid-value (m) d=(m-35)/10 f×d 0-10 4 5 -3 -12 10-20 9 15 -2 -18 20-30 21 25 -1 -21 30-40 37 35 0 0 40-50 10 45 1 10 50-60 10 55 2 20 60-70 9 65 3 27 Total N=100 Σfd=6 Mean (step-deviation / assumed-mean method, A=35, h=10):
Mean = A + (Σfd/N) × h = 35 + (6/100) × 10 = 35 + 0.6 = 35.6
Median:
Cumulative frequencies: 4, 13, 34, 71, 81, 91, 100. N/2 = 50, which first exceeds at the 30-40 class (cf=71), so the median class is 30-40, with cf (of preceding class) = 34, f = 37, L = 30, h = 10.
Median = L + [(N/2 − cf)/f] × h = 30 + [(50−34)/37] × 10 = 30 + (16/37)×10 = 30 + 4.32 = 34.32 (approx.)
Mode:
The highest frequency (37) is in the class 30-40, so this is the modal class: L=30, f1=37, f0=21 (preceding class 20-30), f2=10 (following class 40-50), h=10.
Mode = L + [(f1−f0)/(2f1−f0−f2)] × h = 30 + [(37−21)/(74−21−10)] × 10 = 30 + (16/43)×10 = 30 + 3.72 = 33.72 (approx.)
✓Final answerCorrelation is classified as Positive/Negative, Simple/Partial/Multiple, and Linear/Non-linear. (OR) Mean = 35.6, Median ≈ 34.32, Mode ≈ 33.72.
- Based on direction:
🎓Unlock everything free for 14 days
- ✓Full step-by-step solutions
- ✓Concept-first explanations
- ✓Methods, shortcuts & mistakes
- ✓PYQ mapping + timed mock tests
Full access for 14 days. No credit card required.