Skip to content

Informatics Practices · Ch 7 — Project Based Learning

Project IV: Utilising an open data source to use a national, state or district level Dataset

7.4.4

Project IV: Utilising an open data source to use a national, state or district level Dataset

Understanding the Project

This project is about learning to work with a real, large-scale government dataset. The goal is not just to write code, but to understand how to ask meaningful questions of data and answer them step by step. You will use the Open Government Data (OGD) Platform India, specifically the website www.data.gov.in, which is the official platform for sharing data under the Government of India's open data initiative.

The dataset you will use is called "Special Tabulation on Adolescent and youth population classified by various parameters for India, States and Union Territories, 2011". It was contributed by the Ministry of Home Affairs and released under the National Data Sharing and Accessibility Policy (NDSAP) on 07/09/2015. This is a massive dataset with 12,168 rows and 123 columns, so you will learn how to filter and summarise it efficiently.

Understanding the Dataset Columns

The dataset contains a wealth of information broken down by state, area type (total/rural/urban), and age group. Here are the key columns you will work with:

  • State: Serial numbers assigned to states.
  • Area Name: The name of the state or union territory.
  • Total/Rural/Urban: Indicates whether the data is for the total, rural, or urban population of that area.
  • Adolescent and youth: The specific age group (e.g., 10-14, 15-19, 20-24).
  • Total Male / Total Female: Total number of males and females.
  • SC-M / SC-F: Total number of males and females belonging to Scheduled Castes.
  • ST-M / ST-F: Total number of males and females belonging to Scheduled Tribes.
  • Literates-M / Literates-F: Total number of literate males and females.
  • LiteratesSC-M / LiteratesSC-F: Literate males and females of Scheduled Castes.
  • LiteratesST-M / LiteratesST-F: Literate males and females of Scheduled Tribes.
  • Illiterates-M / Illiterates-F: Total number of illiterate males and females.
  • IlliteratesSC-M / IlliteratesSC-F: Illiterate males and females of Scheduled Castes.
  • IlliteratesST-M / IlliteratesST-F: Illiterate males and females of Scheduled Tribes.
  • MainWorker-M / MainWorker-F: Total number of main worker males and females.
  • MainWorkerSC-M / MainWorkerSC-F: Main worker males and females of Scheduled Castes.
  • MainWorkerST-M / MainWorkerST-F: Main worker males and females of Scheduled Tribes.
  • MarginalWorker-M / MarginalWorker-F: Total number of marginal worker males and females.
  • MarginalWorkerSC-M / MarginalWorkerSC-F: Marginal worker males and females of Scheduled Castes.
  • MarginalWorkerST-M / MarginalWorkerST-F: Marginal worker males and females of Scheduled Tribes.

Types of Questions You Can Answer

With such a rich dataset, you can answer a wide variety of analytical questions. The textbook provides a list of 11 sample queries that demonstrate the range of possibilities. Your project should involve selecting 4-5 of these (or similar ones) and solving them step by step with full documentation.

Here are the sample questions:

  1. What is the total population, total male population and total female population aged 10 to 24 in India?
  2. Which State or Union Territory in India has the maximum number of illiterates in the youth ages?
  3. What is the percentage of people working as a marginal worker?
  4. List the top 5 states or union territories which have the maximum population working as a marginal worker.
  5. Compare the sex ratio of urban areas and rural areas using appropriate graph.
  6. Which state has the highest and the lowest percentage of literate Scheduled Tribes and Scheduled Castes?
  7. For each state, compare the number of female marginal workers with the number of male marginal workers using appropriate graphs.
  8. What percentage of Scheduled Tribes lives in urban areas? Draw a pie chart showing the proportion of literate and illiterate scheduled tribes living in urban areas.
  9. What is the state wise ratio of literates vs. illiterates in all age groups?
  10. Which state is home to the maximum number of ST in India? Which state has the minimum number of ST in India?
  11. For each state, find the number of literate females and literate males. Draw a bar graph for the same. Which state has the highest ratio of literate female vs literate male and which state has the minimum?

Step-by-Step Example: Solving Question 1

The textbook walks through the solution for the first question to give you a clear template for solving the others. The question is: What is the total population, total male population and total female population aged 10 to 24 in India?

The solution follows a structured, step-by-step approach.

Prerequisite: You must first download the CSV file (PCA_AY_2011_Revised.csv) using the QR code provided at the beginning of the chapter.

Step 0: Import Required Libraries

You start by importing the two essential Python libraries: pandas for data manipulation and matplotlib.pyplot for plotting.

import pandas as pd
import matplotlib.pyplot as plt

Step 1: Read the CSV File into a DataFrame

Load the CSV file into a pandas DataFrame.

data = pd.read_csv("PCA_AY_2011_Revised.csv")
df = pd.DataFrame(data)

Step 2: Check the Shape of the DataFrame

This confirms the size of the dataset.

print(df.shape)

The output will show (12168, 123), confirming 12,168 rows and 123 columns.

Step 3: Display the Columns

View all column names to understand what data is available.

print(df.columns.values)

You will see a list of 123 column names, including 'Area Name', 'Total/Rural/Urban', 'Adolescent and youth categories', 'Total Population - Persons', and many more.

Step 4: Filter Data

a. Identify the columns you need. For this question, you only need 'Area Name', 'Total/Rural/Urban', 'Adolescent and youth categories', and 'Total Population - Persons', 'Total Population - Males', 'Total Population - Females'.

b. Identify the rows you need. Check the values in the 'Area Name' column. You will see that the first few rows correspond to 'INDIA', followed by rows for states and districts. For this question, you only need data for 'INDIA'.

Step 5: Create a New DataFrame with Filtered Data

Use .loc[] to select only the rows where 'Area Name' is 'INDIA' and only the columns from 'Area Name' to 'Total Population - Females'.

df1 = df.loc[(df['Area Name'] == 'INDIA'), 'Area Name':'Total Population - Females']

Step 6: Rename Columns for Ease of Use

The original column names are very long. Rename them to shorter, more manageable names.

df1.columns = ['Area', 'Class', 'Category', 'TotalPop', 'MalePop', 'FemalePop']

Step 7: Group Data as per Requirement

When you print df1, you will see that the 'Category' column contains six different categories: '10-14', '15-19', '20-24', 'Adolescent (10-19)', 'All Ages', and 'Youth (15-24)'. To get the total population for each age group across all of India, you need to group by 'Category' and sum the population columns.

d = df1.groupby('Category').sum()

This gives you a new DataFrame d with the total population for each category. Since you are only interested in the specific age groups 10-14, 15-19, and 20-24, you drop the other rows.

d = d.drop(['Adolescent (10-19)', 'All Ages', 'Youth (15-24)'], axis=0)
``` …
DefinitionTask 1

What is the total population, total male population and total female population aged 10 to 24 in India? This worked example solves the first of the project's suggested questions end to end -- reading the Census CSV, checking its shape, filtering to the relevant age group, and aggregating the population totals -- as a model for how a student would …

Table 7.1Output of print(df.columns.values) -- partial column list (123 total)
['Table No.' 'State Code' 'District Code' 'Area Name'
 'Total/ Rural/ Urban' 'Adolescent and youth categories'
 'Total Population - Persons' 'Total Population - Males'
 'Total Population - Females' 'Scheduled Caste - Persons'
 'Scheduled Caste - Males' 'Scheduled Caste - Females'
                     ...              
 'Scheduled Tribe Marginal Worker - Household Industry - Males'
 'Scheduled Tribe Marginal Worker - Household Industry - Females'
 'Scheduled Tribe Marginal Worker - Other Workers - Persons'
 'Scheduled Tribe Marginal Worker - Other Workers - Males' …
Table 7.2Output of print(df['Area Name']) -- Step 4(b)
IndexArea Name
0INDIA
1INDIA
2INDIA
3INDIA
4INDIA
......
12163District - South Andaman (03)
12164District - South Andaman (03)
12165District - South Andaman (03)
12166District - South Andaman (03)
Table 7.3Output of print(df1) -- the filtered, renamed DataFrame (Step 6)
AreaClassCategoryTotalPopMalePopFemalePop
0INDIATotalAll Ages1210854977623270258587584719
1INDIATotal10-141327092126941883563290377
2INDIATotal15-191205264496398239656544053
3INDIATotal20-241114242225758469353839529
4INDIATotalAdolescent (10-19)253235661133401231119834430
5INDIATotalYouth (15-24)231950671121567089110383582
6INDIARuralAll Ages833748852427781058405967794
7INDIARural10-14968044945048815846316336
8INDIARural15-19839024724457055739331915
9INDIARural20-24738350463813866235696384
10INDIARuralAdolescent (10-19)1807069669505871585648251
11INDIARuralYouth (15-24)1577375188270921975028299
12INDIAUrbanAll Ages377106125195489200181616925
13INDIAUrban10-14359047181893067716974041
14INDIAUrban15-19366239771941183917212138
Table 7.4Output of d = df1.GROUP BY('Category').sum() -- Step 7
CategoryTotalPopMalePopFemalePop
10-14265418424138837670126580754
15-19241052898127964792113088106
20-24222848444115169386107679058
Adolescent (10-19)506471322266802462239668860
Table 7.5Output after d.drop(['Adolescent (10-19)','All Ages','Youth (15-24)'], axis=0)
CategoryTotalPopMalePopFemalePop
10-14265418424138837670126580754
15-19241052898127964792113088106
20-24222848444115169386107679058
Figure 7.2Barchart showing population in different categories
Fig. 7.2 — Barchart showing population in different categories

Drawn by us to help you understand the concept clearly, and verified to make sure it's accurate. For exams, practice from your NCERT textbook's own diagram.

The concluding chart of Task 1's steps: after the population table has been read, cleaned and reduced to the three age categories, one plotting call turns it into a grouped bar chart with three bars per category — TotalPop, MalePop and FemalePop, keyed by the legend. Two reading skills are exercised here. The '1e8' offset printed above the y-axis means every value is in hundreds of millions, matplotlib's scientific notation for large numbers. And the bar pattern carries the findings: the 10-14 group is the lar …