Informatics Practices · Ch 2 — Data Handling using Pandas – I
Pandas Series Vs NumPy ndarray
Pandas Series Vs NumPy ndarray
The Core Idea: Label-Driven Alignment
The fundamental difference between a Pandas Series and a NumPy ndarray is not about speed or memory — it is about identity. A NumPy array is a grid of numbers accessed by their integer position. A Series is a collection of values, each with a label (its index). This label changes everything about how operations work.
When you perform an operation between two Series, Pandas does not care about the order of the values. It cares about the labels. It automatically aligns the data based on those labels before doing the calculation. This is called automatic label alignment, and it is the single most important idea in this section.
Duplicate Index Values: Allowed but Risky
Pandas allows non-unique index values — meaning you can have two rows with the same label. This is something NumPy arrays do not support. However, there is a catch. If you attempt an operation that cannot handle duplicate index values, Pandas will raise an exception at that moment. The duplicates are fine until you try something that requires unique labels, like certain alignment operations or the .loc accessor.
Just because Pandas allows duplicate index values does not mean it is always safe. If you try an operation that requires unique labels and duplicates exist, your code will crash with an exception. Always check for duplicates before performing label-sensitive operations.
Automatic Label Alignment in Action
Here is where the power of Series becomes clear. Suppose you have two Series with different labels, or the same labels in a different order. In NumPy, this would fail or produce meaningless results. In Pandas, it works seamlessly.
When you write an operation between two Series that are not aligned — meaning their labels do not match or are not in the same order — Pandas does the following:
- It looks at the labels of both Series.
- It finds the union of all labels from both Series.
- It performs the operation only on labels that exist in both Series.
- For any label that exists in one Series but not the other, the result is marked as NaN (Not a Number).
This means you never have to manually reorder or align your data before performing calculations. You can write code that assumes the labels will take care of themselves. This grants immense freedom and flexibility in interactive data analysis and research.
The result of an operation between unaligned Series always has the union of the indexes involved. Missing labels produce NaN, not an error. This is the opposite of NumPy, where mismatched arrays cause alignment to fail entirely.
The Comparison Table
The textbook provides a clear side-by-side comparison. Here is the essence of it:
| Feature | Pandas Series | NumPy ndarray |
|---|---|---|
| Index type | User-defined labeled index (numbers or letters) | Integer positional index only |
| Index flexibility | Can be indexed in descending order, or with any custom labels | Index starts at 0 and is fixed |
| Handling misaligned data | Produces NaN for missing labels | Alignment fails — no NaN concept |
| Memory usage | Requires more memory | Occupies less memory |
What This Means in Practice
The practical consequence is straightforward. With NumPy arrays, if you try to add two arrays of different lengths or with mismatched positions, you get an error. With Pandas Series, you get a result — one that includes NaN where data is missing. This makes Series far more forgiving and practical for real-world data, where datasets rarely align perfectly.
The textbook emphasises that you can write computations without considering whether all Series involved have the same label or not. This is the key takeaway: label alignment is automatic, not manual.
A Note on Memory …
| Pandas Series | NumPy Arrays |
|---|---|
| In series we can define our own labeled index to access elements of an array. These can be numbers or letters. | NumPy arrays are accessed by their integer position using numbers only. |
| The elements can be indexed in descending order also. | The indexing starts with zero for the first element and the index is fixed. |