Skip to content
Exercises · Q16

Q.Visit data.gov.in , search for the following in “catalogs” option of the website:
• Final population Totals, India and states
• State Wise literacy rate
Download them and create a CSV file containing population data and literacy rate of the respective state. Also add a column Region to the CSV file that should contain the values East, West, North and South. Plot a scatter plot for each region where X axis should be population and Y axis should be Literacy rate. Change the marker to a diamond and size as the square root of the literacy rate.
Group the data on the column region and display a bar chart depicting average literacy rate for each region.

Punjab PsebTextbookSubjective· 5mImportance★★★★★
84% · 37/44 Questions
🔒 Locked · start free trial →

You're viewing a preview — the full solution, concept, methods & PYQ mapping are locked.

Start your 14-day free trial to unlock the full solution →

Download population and literacy data from data.gov.in, create a CSV with a Region column, plot scatter plots (population vs. literacy) for each region with diamond markers sized by √literacy, then plot a bar chart of average literacy by region.


This is a data visualization task that combines data acquisition, preprocessing, and plotting using Pandas and Matplotlib. The question asks you to work with real government data, structure it, and produce two types of charts: scatter plots (one per region) and a grouped bar chart.

Why this approach?

Data preparation is the foundation. You need to merge population and literacy data by state, add a Region column manually (since geographic classification isn't in the raw data), and save it as CSV. Scatter plots reveal correlation between two continuous variables (population and literacy); grouping by region and using different markers/sizes adds a third dimension (literacy magnitude). Bar charts summarize categorical data — here, average literacy per region — making regional comparisons immediate.

The marker size proportional to literacy\sqrt{\text{literacy}} ensures visual differentiation without overwhelming the plot (raw literacy values would produce negligible size variation).


Step 1: Data Acquisition and CSV Creation

Since the question directs you to data.gov.in, you would download the two datasets. For this solution, I'll assume you've created a CSV file state_data.csv with the following structure (a representative sample):

State,Population,Literacy_Rate,Region
Kerala,33406061,93.91,South
Bihar,104099452,61.80,East
Maharashtra,112374333,82.34,West
Uttar Pradesh,199812341,67.68,North
Tamil Nadu,72147030,80.09,South
West Bengal,91276115,76.26,East
Rajasthan,68548437,66.11,North
Karnataka,61095297,75.36,South
Gujarat,60439692,78.03,West
Madhya Pradesh,72626809,69.32,North
Odisha,41974218,72.87,East
Punjab,27743338,75.84,North
Assam,31205576,72.19,East
Haryana,25351462,75.55,North
Goa,1458545,88.70,West
Note

The Region column is added manually based on geographic knowledge: Kerala, Tamil Nadu, Karnataka, Andhra Pradesh → South; Maharashtra, Gujarat, Goa → West; Bihar, West Bengal, Odisha, Assam → East; Uttar Pradesh, Rajasthan, Madhya Pradesh, Punjab, Haryana, Himachal Pradesh → North.


Step 2: Scatter Plot for Each Region

We'll read the CSV, group by Region, and plot a scatter for each group with diamond markers sized by Literacy Rate\sqrt{\text{Literacy Rate}}.

import pandas as pd
import matplotlib.pyplot as plt
import numpy as np

# Read the CSV
df = pd.read_csv('state_data.csv')

# Get unique regions
regions = df['Region'].unique()

# Create a figure with subplots (2x2 grid for 4 regions)
fig, axes = plt.subplots(2, 2, figsize=(14, 10))
axes = axes.flatten()

for i, region in enumerate(regions):
    # Filter data for the region
    region_data = df[df['Region'] == region]
    
    # Calculate marker sizes: square root of literacy rate, scaled for visibility
    sizes = np.sqrt(region_data['Literacy_Rate']) * 20  # Scale factor for visibility
    
    # Plot scatter
    axes[i].scatter(region_data['Population'], 
                   region_data['Literacy_Rate'],
                   marker='D',  # Diamond marker
                   s=sizes,
                   alpha=0.6,
                   c=f'C{i}')  # Different color per region
    
    axes[i].set_xlabel('Population', fontsize=10)
    axes[i].set_ylabel('Literacy Rate (%)', fontsize=10)
    axes[i].set_title(f'{region} Region', fontsize=12, fontweight='bold')
    axes[i].grid(True, alpha=0.3)
    
    # Add state labels
    for idx, row in region_data.iterrows():
        axes[i].annotate(row['State'], 
                        (row['Population'], row['Literacy_Rate']),
                        fontsize=7, alpha=0.7)

plt.tight_layout()
plt.savefig('scatter_by_region.png', dpi=150)
plt.show()

Key lines explained:

  • marker='D': Diamond-shaped markers as specified.
  • sizes = np.sqrt(region_data['Literacy_Rate']) * 20: Marker size is literacy\sqrt{\text{literacy}}, scaled by 20 for visibility (raw square roots would be too small).
  • axes[i].scatter(...): Each region gets its own subplot in a 2×2 grid.
  • annotate(...): Labels each point with the state name for clarity.

Expected output:

A 2×2 grid of scatter plots. Each subplot shows:

  • X-axis: Population (ranging from ~1.5M for Goa to ~200M for Uttar Pradesh)
  • Y-axis: Literacy Rate (ranging from ~62% for Bihar to ~94% for Kerala)
  • Markers: Diamonds, larger for states with higher literacy (Kerala's marker is largest in the South plot)
  • Title: "East Region", "West Region", "North Region", "South Region"
  • States like Kerala (high literacy, moderate population) appear in the upper-middle; Uttar Pradesh (low literacy, huge population) appears far right and lower.
Watch out

If you don't scale the marker sizes (the * 20 factor), the diamonds will be nearly invisible. The square root of 70–90 is only 8–9 pixels, which is too small. Adjust the scaling factor based on your plot size.


Step 3: Bar Chart of Average Literacy by Region

Group the data by Region and compute the mean literacy rate, then plot a bar chart.

# Group by Region and calculate average literacy rate
avg_literacy = df.groupby('Region')['Literacy_Rate'].mean().sort_values(ascending=False)

# Plot bar chart
plt.figure(figsize=(10, 6))
bars = plt.bar(avg_literacy.index, avg_literacy.values, 
               color=['#2E86AB', '#A23B72', '#F18F01', '#C73E1D'],
               edgecolor='black', linewidth=1.2)

plt.xlabel('Region', fontsize=12, fontweight='bold')
plt.ylabel('Average Literacy Rate (%)', fontsize=12, fontweight='bold') …

Unlock everything free for 14 days

  • Full step-by-step solutions
  • Concept-first explanations
  • Methods, shortcuts & mistakes
  • PYQ mapping + timed mock tests

Full access for 14 days. No credit card required.