Q.Collect and store data related to the screen time of students in your class separately for boys and girls and present it using a boxplot.
You're viewing a preview — the full solution, concept, methods & PYQ mapping are locked.
Start your 14-day free trial to unlock the full solution →This solution demonstrates how to collect simulated screen time data for boys and girls using pandas and numpy, and then visualize its distribution using a boxplot with seaborn and matplotlib.
Data visualization is a powerful tool for understanding patterns, distributions, and relationships within data. When you have numerical data (like screen time) categorized by a qualitative factor (like gender), a boxplot is an excellent choice to compare their distributions visually.
A boxplot, also known as a box-and-whisker plot, effectively summarizes the five-number summary of a set of data:
- Minimum: The smallest value (excluding outliers).
- First Quartile (Q1): 25% of the data falls below this value.
- Median (Q2): The middle value, where 50% of the data falls below it.
- Third Quartile (Q3): 75% of the data falls below this value.
- Maximum: The largest value (excluding outliers).
It also clearly shows potential outliers, which are data points that fall significantly outside the typical range of the rest of the data. By plotting boxplots for boys and girls side-by-side, we can easily compare their central tendencies (medians), spread (interquartile range), and identify any differences in their screen time habits.
Here's how you can collect, store, and visualize this data using Python:
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
# --- 1. Data Collection and Storage ---
# Since the question doesn't provide specific data, we will simulate it.
# In a real scenario, you would collect this data through surveys, apps, etc.
# Generate random screen time data for boys and girls
# Let's assume screen time is in hours per day.
np.random.seed(42) # For reproducibility
# Simulate screen time for boys (e.g., slightly higher mean and spread)
boys_screen_time = np.random.normal(loc=4.5, scale=1.5, size=50)
# Ensure screen time is non-negative
boys_screen_time[boys_screen_time < 0] = 0
# Simulate screen time for girls (e.g., slightly lower mean and spread)
girls_screen_time = np.random.normal(loc=3.8, scale=1.2, size=50)
girls_screen_time[girls_screen_time < 0] = 0
# Create a DataFrame to store the data
# A pandas DataFrame is ideal for tabular data, allowing easy manipulation and plotting.
data = {
'Gender': ['Boy'] * len(boys_screen_time) + ['Girl'] * len(girls_screen_time),
'ScreenTime_Hours': np.concatenate((boys_screen_time, girls_screen_time))
}
df = pd.DataFrame(data)
print("Sample of the collected data:")
print(df.head())
print("\nData Info:")
df.info()
# --- 2. Data Visualization using Boxplot ---
# Set the style for the plots for better aesthetics
sns.set_style("whitegrid")
# Create the boxplot
plt.figure(figsize=(8, 6)) # Set the figure size for better readability
# Use seaborn.boxplot to create the plot
# 'x' specifies the categorical variable (Gender) on the x-axis.
# 'y' specifies the numerical variable (ScreenTime_Hours) on the y-axis.
# 'data' specifies the DataFrame to use.
sns.boxplot(x='Gender', y='ScreenTime_Hours', data=df, palette='viridis')
# Add title and labels for clarity
plt.title('Distribution of Daily Screen Time for Boys and Girls', fontsize=16)
plt.xlabel('Gender', fontsize=12)
plt.ylabel('Screen Time (Hours/Day)', fontsize=12)
# Show the plot
plt.show()
Explanation of Key Lines:
-
import pandas as pd,import numpy as np,import matplotlib.pyplot as plt,import seaborn as sns: These lines import the necessary libraries.pandasis used for creating and managing the DataFrame, which is a tabular data structure.numpyis used here to generate random numerical data for simulation.matplotlib.pyplotis the foundational plotting library in Python, used for displaying the plot.seabornis a high-level data visualization library built onmatplotlib, providing a simpler interface for creating attractive statistical graphics like boxplots.
-
boys_screen_time = np.random.normal(loc=4.5, scale=1.5, size=50): This line usesnumpyto generate 50 random numbers following a normal distribution.loc(mean) is set to 4.5 hours, representing the average screen time for boys.scale(standard deviation) is set to 1.5 hours, indicating the spread of screen times.sizeis the number of data points to generate.- A similar line generates data for girls with slightly different parameters to show a potential difference.
-
data = {'Gender': ['Boy'] * len(boys_screen_time) + ['Girl'] * len(girls_screen_time), 'ScreenTime_Hours': np.concatenate((boys_screen_time, girls_screen_time))}: This dictionary prepares the data for the DataFrame.- It creates a
Genderlist by repeating 'Boy' and 'Girl' labels to match the number of screen time entries. np.concatenatecombines the screen time arrays for boys and girls into a single array.
- It creates a
-
df = pd.DataFrame(data): This creates apandasDataFrame nameddffrom thedatadictionary. Each key in the dictionary becomes a column in the DataFrame. -
sns.set_style("whitegrid"): This sets the visual style of the plots to a clean white background with a grid, which is often preferred for statistical plots. -
plt.figure(figsize=(8, 6)): This creates a new figure for the plot and sets its size to 8 inches wide and 6 inches tall, ensuring the plot is well-proportioned and readable. -
sns.boxplot(x='Gender', y='ScreenTime_Hours', data=df, palette='viridis'): This is the core line for creating the boxplot.x='Gender': Tellsseabornto place the 'Gender' column on the x-axis, creating separate boxes for each gender. …
Unlock everything free for 14 days
- Full step-by-step solutions
- Concept-first explanations
- Methods, shortcuts & mistakes
- PYQ mapping + timed mock tests
Full access for 14 days. No credit card required.