Skip to content
Tasks

Q.Utilising an open data source: Open Government Data (OGD) Platform India www.data.gov.in is a platform for supporting open data initiative of Government of India. From this platform, let us consider the dataset "Special Tabulation on Adolescent and youth population classified by various parameters for India, States and Union Territories, 2011". The dataset was contributed by the Ministry of Home Affairs, Government of India, and released under National Data Sharing and Accessibility Policy (NDSAP). The dataset was published on portal on 07/09/2015. Statistics of the Data Set:
Number of rows: 12168
Number of columns: 123 Descriptions of some of the columns are given below:
Area Name: Name of the states and union territories
Total/Rural/Urban: Data about the total, rural or urban areas of a state or UT
Adolescent and youth: Data for different age groups
Total Male: Total number of males
Total Female: Total number of females
Literates-M: total number of literate males
Literates-F: total number of literate females Task 1: What is the total population, total male population and total female population aged 10 to 24 in India?

Yanam BieapTextbookSubjective· 3mImportance★★★★★est
100% · 10/10 Questions
🔒 Locked · start free trial →

You're viewing a preview — the full solution, concept, methods & PYQ mapping are locked.

Start your 14-day free trial to unlock the full solution →

A complete pandas data-handling workflow on a real open-government dataset: read the CSV, inspect its shape and columns, filter the India rows, rename unwieldy column names, select the three age groups 10-14 / 15-19 / 20-24, sum them, and plot a bar chart. Answer: 364,659,883 people aged 10-24 (190,985,924 male, 173,673,959 female).

The dataset is the Census-2011 special tabulation on adolescent and youth population from data.gov.in. It has 12,168 rows and 123 columns: one row per (area, class, age-category) combination, where area runs from INDIA down to individual districts, class is Total/Rural/Urban, and age category takes six values — '10-14', '15-19', '20-24', 'Adolescent (10-19)', 'All Ages', 'Youth (15-24)'.

Step 0 — Import the libraries.

import pandas as pd
import matplotlib.pyplot as plt

Step 1 — Read the CSV into a DataFrame.

df = pd.read_csv("PCA_AY_2011_Revised.csv")   # give the path on your computer

Step 2 — Check the shape.

print(df.shape)
(12168, 123)

12,168 rows and 123 columns — far more than we need for this task.

Step 3 — See what the columns are.

print(df.columns.values)

The list starts with 'Table No.' 'State Code' 'District Code' 'Area Name' 'Total/ Rural/ Urban' 'Adolescent and youth categories' 'Total Population - Persons' 'Total Population - Males' 'Total Population - Females' ... and continues through literacy and worker breakdowns. For this task we only need Area Name, the Total/Rural/Urban class, the age category, and the three population columns.

Step 4 — Filter to the rows and columns we need.

The question is about India as a whole, so we keep only rows where Area Name is 'INDIA' (the district-level rows are not needed), and slice the columns from 'Area Name' through 'Total Population - Females':

df1 = df.loc[df['Area Name'] == 'INDIA',
             'Area Name':'Total Population - Females']

Step 5 — Rename the columns. The census column names are long; short names make the rest of the code readable:

df1.columns = ['Area', 'Class', 'Category', 'TotalPop', 'MalePop', 'FemalePop']
print(df1.head())
    Area  Class            Category    TotalPop    MalePop  FemalePop
0  INDIA  Total            All Ages  1210854977  623270258  587584719
1  INDIA  Total               10-14   132709212   69418835   63290377
2  INDIA  Total               15-19   120526449   63982396   56544053
3  INDIA  Total               20-24   111424222   57584693   53839529
4  INDIA  Total  Adolescent (10-19)   253235661  133401231  119834430

There are 18 INDIA rows in all: the six age categories, each repeated for Class = Total, Rural and Urban.

Step 6 — Select exactly the age groups asked for. "Aged 10 to 24" is the union of the three disjoint categories '10-14', '15-19' and '20-24'. The composite categories ('Adolescent (10-19)', 'Youth (15-24)', 'All Ages') overlap these, so they must be excluded. Likewise we keep only Class == 'Total':

ages = df1[(df1['Class'] == 'Total') &
           (df1['Category'].isin(['10-14', '15-19', '20-24']))]
print(ages[['Category', 'TotalPop', 'MalePop', 'FemalePop']])
  Category   TotalPop   MalePop  FemalePop
1    10-14  132709212  69418835   63290377
2    15-19  120526449  63982396   56544053
3    20-24  111424222  57584693   53839529
``` …

Unlock everything free for 14 days

  • Full step-by-step solutions
  • Concept-first explanations
  • Methods, shortcuts & mistakes
  • PYQ mapping + timed mock tests

Full access for 14 days. No credit card required.