Skip to content

Informatics Practices · Ch 3 — Data Handling using Pandas – II

Altering the Index

3.6

Altering the Index

Altering the Index

The index of a DataFrame is the set of row labels used to access and retrieve data quickly. By default, Pandas assigns a numeric index starting from 0, as shown in the sample DataFrame of student marks across different Unit Tests (UT). In that table, the first column (0, 1, 2, …) is the default index.

However, the default index may not always suit your needs. You might want to use an existing column (like student name or UT number) as the index, or you may need to clean up a non‑continuous index after slicing data. Pandas provides three key functions to alter the index: reset_index(), drop(), and set_index().


The Problem of Non‑Continuous Index After Slicing

When you filter a DataFrame (for example, selecting only rows where UT == 1), the resulting slice retains the original index values. These values are no longer consecutive — they are the original row numbers from the parent DataFrame. For instance, after selecting Unit Test 1 marks, the index might be 0, 3, 6, 9 instead of 0, 1, 2, 3. This non‑continuous index can be inconvenient for further operations.


Using reset_index() to Create a New Continuous Index

The reset_index() function resets the index to a default sequence of integers (0, 1, 2, …). By default, it does not modify the original DataFrame — it returns a new DataFrame. To change the original DataFrame itself, use the inplace=True parameter.

When you call reset_index(inplace=True), a new column named 'index' is added to the DataFrame. This column holds the old index values, so you can still refer to them if needed. The new continuous index becomes the row label.

Note

After reset_index(), the old index is preserved as a separate column called 'index'. You can keep it or drop it.


Dropping the Old Index Column with drop()

If you do not need the old index values, you can remove the 'index' column using the drop() function. Specify columns=['index'] and set inplace=True to modify the DataFrame directly.

dfUT1.drop(columns=['index'], inplace=True)

After this, the DataFrame has only the new continuous index (0, 1, 2, …) and no trace of the original index.


Changing the Index to Another Column with set_index()

You can make any existing column the new row index using set_index(). For example, to set the 'Name' column as the index:

dfUT1.set_index('Name', inplace=True)

Now the row labels become the student names (Raman, Zuhaire, Ashravy, Mishti), and the 'Name' column disappears from the data area — it becomes the index. The original numeric index is replaced.

Watch out

set_index() does not keep the old index as a column. If you need the old index later, reset it first or store it separately.


Reverting to the Previous Index

To revert back to the default numeric index after using set_index(), call reset_index() again. This time, you can specify the name of the column that is currently the index (e.g., 'Name') to move it back to a regular column:

dfUT1.reset_index('Name', inplace=True)

This restores the default integer index and brings the 'Name' column back into the DataFrame as a regular column.


Summary of Functions …
Table 3.56Output of print(df) -- the case-study DataFrame with its default index
IndexNameUTMathsScienceS.StHindiEng
0Raman12221182021
1Raman22120172224
2Raman31419152423
3Zuhaire12017222419
4Zuhaire22315212515
5Zuhaire32218192313
6Ashravy12319201522
7Ashravy22422241721
8Ashravy31225192123
9Mishti11522252222
10Mishti21821252423
11Mishti31718202520
Table 3.33Output of dfUT1 = df[df.UT==1] -- non-continuous index after slicing
IndexNameUTMathsScienceS.StHindiEng
0Raman12221182021
3Zuhaire12017222419
6Ashravy12319201522
Table 3.34Output of dfUT1.reset_index(inplace=True) -- new continuous index added
indexNameUTMathsScienceS.StHindiEng
00Raman12221182021
13Zuhaire12017222419
26Ashravy12319201522
39Mishti11522252222
Table 3.35Output of dfUT1.drop(columns=['index'], inplace=True) -- old index column dropped
IndexNameUTMathsScienceS.StHindiEng
0Raman12221182021
1Zuhaire12017222419
2Ashravy12319201522
Table 3.36Output of dfUT1.set_index('Name', inplace=True) -- Name column becomes the index
NameUTMathsScienceS.StHindiEng
Raman12221182021
Zuhaire12017222419
Ashravy12319201522