Skip to content

Informatics Practices · Ch 2 — Data Handling using Pandas – I

Mathematical Operations on Series

2.2.5

Mathematical Operations on Series

In Class XI we saw that basic mathematical operations — addition, subtraction, multiplication, division, and so on — applied to two NumPy arrays act on each corresponding pair of elements. Pandas lets us do the same with two Series, with one important twist: the operation is carried out by index matching. Pandas lines up elements that share the same index label, operates on those pairs, and fills every position where a match is missing with NaN by default.

To see this behaviour concretely, the textbook sets up two Series whose index labels only partly overlap:

>>> seriesA = pd.Series([1,2,3,4,5], index = ['a', 'b', 'c', 'd', 'e'])
>>> seriesA
a    1
b    2
c    3
d    4
e    5
dtype: int64
>>> seriesB = pd.Series([10,20,-10,-50,100], index = ['z', 'y', 'a', 'c', 'e'])
>>> seriesB
z     10
y     20
a    -10
c    -50
e    100
dtype: int64
``` …
(A)

Addition of two Series

Two Series can be added in two ways.

Method 1 — the + operator. The simplest way is to add the two Series directly. Pandas matches elements by index label; wherever one of the two elements (or both) has no value, the result of the addition is NaN:

>>> seriesA + seriesB
a     -9.0
b      NaN
c    -47.0
d      NaN
e    105.0
y      NaN
z      NaN
dtype: float64

Table 2.2 sets out exactly which values were matched at each index while performing this addition. Reading it: at index a, 1 (from seriesA) + (-10) (from seriesB) gives -9.0; at c, 3 + (-50) gives -47.0; and at e, 5 + 100 gives 105.0. At b and d only seriesA has a value (2 and 4), and at y and z only seriesB has one (20 and 10) — so all four of those positions come out as NaN.

Method 2 — the add() method with fill_value. When we do not want NaN values in the output, we use the Series method add() together with the parameter fill_value, which replaces any missing value with the value we specify. Calling seriesA.add(seriesB) is equivalent to seriesA + seriesB, but add() lets us explicitly state what to substitute for any element missing from seriesA or seriesB:

>>> seriesA.add(seriesB, fill_value=0)
a     -9.0
b      2.0
c    -47.0
d      4.0
e    105.0
y     20.0
z     10.0
dtype: float64

Table 2.3 shows the matching for this version: the missing partners at b, d, y and z are treated as 0, so those positions now yield 2.0, 4.0, 20.0 and 10.0 instead of NaN, while the fully matched indices a, c and e are unchanged.

The contrast between the two tables is the whole point: Table 2.2 shows the output without replacing missing values, Table 2.3 the output after replacing them with 0. …

Table 2.2Details of addition of two series
indexvalue from seriesAvalue from seriesBseriesA + seriesB
a1-10-9.0
b2NaN
c3-50-47.0
d4NaN
Table 2.3Details of addition of two series using add() method
indexvalue from seriesAvalue from seriesBseriesA + seriesB
a1-10-9.0
b202.0
c3-50-47.0
d404.0
(B)

Subtraction of two Series

Subtraction, like addition, can be done in two different ways.

Method 1 — the - operator. Direct subtraction matches elements by index label and produces NaN wherever either Series is missing a value:

>>> seriesA - seriesB   # using subtraction operator
a     11.0
b      NaN
c     53.0
d      NaN
e    -95.0
y      NaN
z      NaN
dtype: float64

Only the shared labels produce numbers: at a, 1 - (-10) = 11.0; at c, 3 - (-50) = 53.0; at e, 5 - 100 = -95.0. The unmatched labels b, d, y and z all give NaN.

Method 2 — the sub() method with fill_value. To avoid the NaNs, we replace the missing values before subtracting — here with 1000 — using the explicit subtraction method sub():

>>> seriesA.sub(seriesB, fill_value=1000)
# using fill value 1000 while making explicit
# call of the method
a      11.0
b    -998.0
c      53.0
d    -996.0
e     -95.0
y     980.0
z     990.0
dtype: float64

Check the filled positions against the rule: at b, seriesB's missing value becomes 1000, so 2 - 1000 = -998.0; at d, 4 - 1000 = -996.0; at y, seriesA's missing value becomes 1000, so 1000 - 20 = 980.0; and at z, 1000 - 10 = 990.0. The fully matched indices a, c and e are the same as with the plain operator. …

(C)

Multiplication of two Series

Multiplication of two Series follows the same two-way pattern.

Method 1 — the * operator. Direct multiplication matches elements by index label; unmatched positions come out as NaN:

>>> seriesA * seriesB   # using multiplication operator
a    -10.0
b      NaN
c   -150.0
d      NaN
e    500.0
y      NaN
z      NaN
dtype: float64

At the shared labels: a gives 1 × (-10) = -10.0, c gives 3 × (-50) = -150.0, and e gives 5 × 100 = 500.0. Every label present in only one of the two Series (b, d, y, z) produces NaN.

Method 2 — the mul() method with fill_value. To get a concrete output instead of NaN, we replace the missing values with 0 before multiplying, using the explicit multiplication method mul():

>>> seriesA.mul(seriesB, fill_value=0)
# using fill value 0 while making
# explicit call of the method
a    -10.0
b      0.0
c   -150.0
d      0.0
e    500.0
y      0.0
z      0.0
dtype: float64

Since anything multiplied by the fill value 0 is 0, all four previously-NaN positions now show 0.0, while the matched indices are unchanged. …

(D)

Division of two Series

Division of two Series, once again, can be done in two different ways.

Method 1 — the / operator. Direct division matches elements by index label and yields NaN wherever a partner value is missing:

>>> seriesA/seriesB   # using division operator
a    -0.10
b      NaN
c    -0.06
d      NaN
e     0.05
y      NaN
z      NaN
dtype: float64

At the shared labels: a gives 1 ÷ (-10) = -0.10, c gives 3 ÷ (-50) = -0.06, and e gives 5 ÷ 100 = 0.05. The labels found in only one Series (b, d, y, z) give NaN.

Method 2 — the div() method with fill_value. Replacing the missing values with 0 before dividing, via the explicit division method div():

>>> seriesA.div(seriesB, fill_value=0)
# using fill value 0 while making explicit
# call of the method
a    -0.10
b      inf
c    -0.06
d      inf
e     0.05
y     0.00
z     0.00
dtype: float64
``` …