Informatics Practices · Ch 2 — Data Handling using Pandas – I
Mathematical Operations on Series
Mathematical Operations on Series
In Class XI we saw that basic mathematical operations — addition, subtraction, multiplication, division, and so on — applied to two NumPy arrays act on each corresponding pair of elements. Pandas lets us do the same with two Series, with one important twist: the operation is carried out by index matching. Pandas lines up elements that share the same index label, operates on those pairs, and fills every position where a match is missing with NaN by default.
To see this behaviour concretely, the textbook sets up two Series whose index labels only partly overlap:
>>> seriesA = pd.Series([1,2,3,4,5], index = ['a', 'b', 'c', 'd', 'e'])
>>> seriesA
a 1
b 2
c 3
d 4
e 5
dtype: int64
>>> seriesB = pd.Series([10,20,-10,-50,100], index = ['z', 'y', 'a', 'c', 'e'])
>>> seriesB
z 10
y 20
a -10
c -50
e 100
dtype: int64
``` …
Addition of two Series
Two Series can be added in two ways.
Method 1 — the + operator. The simplest way is to add the two Series directly. Pandas matches elements by index label; wherever one of the two elements (or both) has no value, the result of the addition is NaN:
>>> seriesA + seriesB
a -9.0
b NaN
c -47.0
d NaN
e 105.0
y NaN
z NaN
dtype: float64
Table 2.2 sets out exactly which values were matched at each index while performing this addition. Reading it: at index a, 1 (from seriesA) + (-10) (from seriesB) gives -9.0; at c, 3 + (-50) gives -47.0; and at e, 5 + 100 gives 105.0. At b and d only seriesA has a value (2 and 4), and at y and z only seriesB has one (20 and 10) — so all four of those positions come out as NaN.
Method 2 — the add() method with fill_value. When we do not want NaN values in the output, we use the Series method add() together with the parameter fill_value, which replaces any missing value with the value we specify. Calling seriesA.add(seriesB) is equivalent to seriesA + seriesB, but add() lets us explicitly state what to substitute for any element missing from seriesA or seriesB:
>>> seriesA.add(seriesB, fill_value=0)
a -9.0
b 2.0
c -47.0
d 4.0
e 105.0
y 20.0
z 10.0
dtype: float64
Table 2.3 shows the matching for this version: the missing partners at b, d, y and z are treated as 0, so those positions now yield 2.0, 4.0, 20.0 and 10.0 instead of NaN, while the fully matched indices a, c and e are unchanged.
The contrast between the two tables is the whole point: Table 2.2 shows the output without replacing missing values, Table 2.3 the output after replacing them with 0. …
| index | value from seriesA | value from seriesB | seriesA + seriesB |
|---|---|---|---|
| a | 1 | -10 | -9.0 |
| b | 2 | NaN | |
| c | 3 | -50 | -47.0 |
| d | 4 | NaN |
| index | value from seriesA | value from seriesB | seriesA + seriesB |
|---|---|---|---|
| a | 1 | -10 | -9.0 |
| b | 2 | 0 | 2.0 |
| c | 3 | -50 | -47.0 |
| d | 4 | 0 | 4.0 |
Subtraction of two Series
Subtraction, like addition, can be done in two different ways.
Method 1 — the - operator. Direct subtraction matches elements by index label and produces NaN wherever either Series is missing a value:
>>> seriesA - seriesB # using subtraction operator
a 11.0
b NaN
c 53.0
d NaN
e -95.0
y NaN
z NaN
dtype: float64
Only the shared labels produce numbers: at a, 1 - (-10) = 11.0; at c, 3 - (-50) = 53.0; at e, 5 - 100 = -95.0. The unmatched labels b, d, y and z all give NaN.
Method 2 — the sub() method with fill_value. To avoid the NaNs, we replace the missing values before subtracting — here with 1000 — using the explicit subtraction method sub():
>>> seriesA.sub(seriesB, fill_value=1000)
# using fill value 1000 while making explicit
# call of the method
a 11.0
b -998.0
c 53.0
d -996.0
e -95.0
y 980.0
z 990.0
dtype: float64
Check the filled positions against the rule: at b, seriesB's missing value becomes 1000, so 2 - 1000 = -998.0; at d, 4 - 1000 = -996.0; at y, seriesA's missing value becomes 1000, so 1000 - 20 = 980.0; and at z, 1000 - 10 = 990.0. The fully matched indices a, c and e are the same as with the plain operator. …
Multiplication of two Series
Multiplication of two Series follows the same two-way pattern.
Method 1 — the * operator. Direct multiplication matches elements by index label; unmatched positions come out as NaN:
>>> seriesA * seriesB # using multiplication operator
a -10.0
b NaN
c -150.0
d NaN
e 500.0
y NaN
z NaN
dtype: float64
At the shared labels: a gives 1 × (-10) = -10.0, c gives 3 × (-50) = -150.0, and e gives 5 × 100 = 500.0. Every label present in only one of the two Series (b, d, y, z) produces NaN.
Method 2 — the mul() method with fill_value. To get a concrete output instead of NaN, we replace the missing values with 0 before multiplying, using the explicit multiplication method mul():
>>> seriesA.mul(seriesB, fill_value=0)
# using fill value 0 while making
# explicit call of the method
a -10.0
b 0.0
c -150.0
d 0.0
e 500.0
y 0.0
z 0.0
dtype: float64
Since anything multiplied by the fill value 0 is 0, all four previously-NaN positions now show 0.0, while the matched indices are unchanged. …
Division of two Series
Division of two Series, once again, can be done in two different ways.
Method 1 — the / operator. Direct division matches elements by index label and yields NaN wherever a partner value is missing:
>>> seriesA/seriesB # using division operator
a -0.10
b NaN
c -0.06
d NaN
e 0.05
y NaN
z NaN
dtype: float64
At the shared labels: a gives 1 ÷ (-10) = -0.10, c gives 3 ÷ (-50) = -0.06, and e gives 5 ÷ 100 = 0.05. The labels found in only one Series (b, d, y, z) give NaN.
Method 2 — the div() method with fill_value. Replacing the missing values with 0 before dividing, via the explicit division method div():
>>> seriesA.div(seriesB, fill_value=0)
# using fill value 0 while making explicit
# call of the method
a -0.10
b inf
c -0.06
d inf
e 0.05
y 0.00
z 0.00
dtype: float64
``` …