Imagine you're tracking two things over time — say, the number of ice creams sold at your school canteen and the outside temperature. On hot days, both go up; on cool days, both drop. They seem to move together. That's the basic idea of correlation: a measure of how two variables move in relation to each other.
But not all co-movement is equal. Sometimes one variable goes up while the other goes down — like the price of a good and the quantity demanded (law of demand). Sometimes they seem to have no connection at all — like the number of students in your class and the price of tea in China. Pearson correlation gives us a single number that captures the strength and direction of this linear relationship.
The Precise Meaning
Pearson correlation coefficient, usually denoted by r, measures the linear relationship between two variables X and Y. It answers: If I know how far X is from its average, how far (and in which direction) is Y from its average, on average?
The formula is:
r=∑(Xi−Xˉ)2⋅∑(Yi−Yˉ)2∑(Xi−Xˉ)(Yi−Yˉ)
Where:
Xi, Yi are individual observations
Xˉ, Yˉ are the means (averages) of X and Y
∑ means "sum over all observations"
The numerator is the covariance — it tells you whether deviations from the mean tend to be in the same direction (positive product) or opposite directions (negative product). The denominator is the product of the standard deviations of X and Y, which scales the result so that r always lies between −1 and +1.
Important
r is unitless and always between −1 and +1:
r=+1: perfect positive linear relationship (all points lie on an upward-sloping line)
r=−1: perfect negative linear relationship (all points lie on a downward-sloping line)
r=0: no linear relationship (but there could still be a non-linear one!)
Why It Matters in Economics
Economics is full of pairs of variables that we suspect move together. Pearson correlation gives us a first, clean check on whether that suspicion holds water.
Example 1: Consumption and Income. Keynes said consumption depends on income. If you plot household consumption against household income for a sample of families, you'd expect a positive r — higher income families tend to consume more. A value close to +0.8 or +0.9 would be strong evidence for that relationship.
Example 2: Price and Quantity Demanded. The law of demand says price and quantity demanded move in opposite directions. A negative r between price and quantity (holding other factors constant) would confirm this. But here's the catch — in real market data, price and quantity are determined simultaneously by supply and demand, so a simple correlation might not show the expected negative sign. That's why economists use more advanced tools (like regression) to isolate the relationship.
Example 3: Investment and Interest Rates. You'd expect a negative correlation — when interest rates are high, borrowing is expensive, so investment falls. But the relationship might be weak (r close to 0) because investment also depends on expectations, technology, and government policy.
Watch out
Correlation does NOT imply causation. Just because ice cream sales and drowning incidents are positively correlated (both peak in summer) does NOT mean ice cream causes drowning. The common cause is hot weather, which makes people both buy ice cream and go swimming. In economics, this is a constant trap — GDP and money supply are correlated, but which causes which? The answer requires theory, not just correlation.
Visualising It
Draw a scatter plot with X on the horizontal axis and Y on the vertical axis. If the points cluster around a straight line sloping upward, r is positive and strong. If they cluster around a line sloping downward, r is negative and strong. If they form a shapeless cloud, r is near zero.
But here's the nuance: a perfect circle of points has r=0 even though X and Y are clearly related (non-linearly). Pearson correlation only captures linear relationships. Two variables could be perfectly related by a U-shaped curve and still have r=0.
Tip
Always plot your data first. A single outlier can dramatically change r, and a non-linear relationship can hide behind a near-zero r. The number alone is never enough — you need to see the shape.
A Quick Worked Example
Suppose you have data on five families:
Family
Income (₹'000)
Consumption (₹'000)
A
10
8
B
20
15
C
30
22
D
40
30
E
50
38
Mean income Xˉ=30, mean consumption Yˉ=22.6.
Compute deviations and products:
For family A: (10−30)(8−22.6)=(−20)(−14.6)=292
For family B: (20−30)(15−22.6)=(−10)(−7.6)=76
For family C: (30−30)(22−22.6)=(0)(−0.6)=0
For family D: (40−30)(30−22.6)=(10)(7.4)=74
For family E: (50−30)(38−22.6)=(20)(15.4)=308
Sum of products = 292+76+0+74+308=750
Sum of squared deviations for income: (−20)2+(−10)2+02+102+202=400+100+0+100+400=1000
Sum of squared deviations for consumption: (−14.6)2+(−7.6)2+(−0.6)2+7.42+15.42=213.16+57.76+0.36+54.76+237.16=563.2
r=1000×563.2750=563200750≈750.47750≈0.999
That's almost perfect positive correlation — consumption rises almost exactly linearly with income in this small sample.
The Bottom Line
Pearson correlation is your first tool for spotting linear relationships in economic data. It's simple, intuitive, and powerful — but limited. Use it to check whether two variables move together, but never to claim that one causes the other. In economics, where everything is connected to everything else, that caution is worth its weight in gold.
Karl Pearson's coefficient of correlation is worked out here using deviations of each observation from its own arithmetic mean, which keeps the arithmetic simpler than working with the raw values directly.
✓Final answer
Karl Pearson's coefficient of correlation is
r=∑x2⋅∑y2∑xy=112×3842=65.2442≈0.644
A moderately strong positive correlation — more years of schooling tend to go with a higher yield per acre.
Using deviations from the means (Xˉ=6, Yˉ=7), ∑xy=42, ∑x2=112, ∑y2=38, giving r≈+0.644 — a moderately strong positive relationship between a farmer's years of schooling and the annual yield per acre.
Concept first
Karl Pearson's coefficient of correlation (r) measures the direction and strength of the linear association between two quantitative variables. It is a pure number lying between −1 and +1. Taking deviations from the arithmetic means simplifies the arithmetic:
r=∑x2∑y2∑xy,x=X−Xˉ,y=Y−Yˉ
The working table
Xˉ=742=6,Yˉ=749=7.
X
Y
x=X−6
y=Y−7
xy
x2
y2
0
4
-6
-3
18
36
9
2
4
-4
-3
12
16
9
4
6
-2
-1
2
4
1
6
10
0
3
0
0
9
8
10
2
3
6
4
9
10
8
4
1
4
16
1
12
7
6
0
0
36
0
Σ
0
0
42
112
38
Substituting
r=112×3842=425642=65.2442=0.6438
Interpretation
The value +0.644 is positive and moderately high: schooling and yield move together, though not perfectly — other factors (soil, irrigation, inputs) also affect yield.
✓Final answer
r=∑x2∑y2∑xy=112×3842≈+0.644
A moderately strong positive correlation between years of schooling and annual yield per acre.
Same / Similar Concept — real previous-year questions on the same or a closely similar concept, not this exact question.
BSEH Haryana Senior Secondary Class 11 (Commerce) 2021Set ANNUAL1 markMCQ
Q.Correctly match the following — Correlation Coefficient :
(a) Measure of Poverty
(b) Economic Activity
(c) Liberalisation
(d) Quartile
(e) 1948
(f) Arithmetic Mean
(g) Karl Pearson
(h) Land Reforms
(i) Regional Rural Banks
(j) Diagramatic Presentation of Data
›Reveal solutionSolution
Correlation Coefficient = Karl Pearson.
The mathematical coefficient of correlation (the product-moment coefficient) was developed by Karl Pearson. So 'Correlation Coefficient' matches 'Karl Pearson'. (Spearman gave the rank correlation.)