Q.A school having 500 students wants to identify beneficiaries of the merit-cum means scholarship, achieving more than 75% for two consecutive years and having family income less than 5 lakh per annum. Briefly describe data processing steps to be taken by the to beneficial prepare the list of school.
This is a classic data processing pipeline: collect the raw facts (marks of two years + family income), organise them into one table, clean them, then filter using three conditions joined by AND — which is really a set intersection of three groups of students. The rows that survive all three conditions are the beneficiaries.
The data processing steps
-
Collection — for each of the 500 students, gather: roll number, name, percentage in Year 1, percentage in Year 2, and declared annual family income (with income certificate as proof).
-
Storage / organisation — put all of it in a single structured table (a spreadsheet, a database table, or a pandas DataFrame), one row per student.
-
Cleaning — remove or correct rows with missing marks, invalid percentages (e.g., > 100) or missing income proof, so the selection is fair.
-
Processing (filtering) — apply the three eligibility conditions together:
- percentage in Year 1 > 75, and
- percentage in Year 2 > 75 (that is what "two consecutive years" means), and
- family income < ₹5,00,000 per annum.
In set terms, the beneficiaries are the intersection: (students >75% in Y1) ∩ (students >75% in Y2) ∩ (students with income < 5 lakh).
-
Output — generate the sorted beneficiary list and publish it.
The same pipeline as runnable code
import pandas as pd
df = pd.DataFrame({
'Roll': [1, 2, 3, 4, 5],
'Name': ['Asha', 'Bilal', 'Chetan', 'Diya', 'Esha'],
'Perc_Y1': [82, 76, 91, 68, 79],
'Perc_Y2': [85, 74, 88, 80, 77],
'Income_L': [4.2, 3.5, 6.0, 2.8, 4.9] # family income in lakhs
})
beneficiaries = df[(df.Perc_Y1 > 75) & (df.Perc_Y2 > 75) & (df.Income_L < 5)]
print(beneficiaries[['Roll', 'Name']])
Expected output:
| Roll | Name | |
|---|---|---|
| 0 | 1 | Asha |
| 4 | 5 | Esha |
Why the key line works: each comparison like df.Perc_Y1 > 75 produces a True/False column; the & operator keeps only rows where all three are True — exactly the intersection the scholarship rule demands. (Bilal fails Year-2, Chetan's income is 6 lakh, Diya fails Year-1 — so only Asha and Esha qualify.)
Steps: (1) collect each student's two-year percentages and family income, (2) organise them in one table, (3) clean invalid/missing entries, (4) filter with the combined condition percentage > 75 in both years AND income < ₹5 lakh — the intersection of the three qualifying sets, (5) output the beneficiary list.
Unlock everything free for 14 days
- Full step-by-step solutions
- Concept-first explanations
- Methods, shortcuts & mistakes
- PYQ mapping + timed mock tests
Full access for 14 days. No credit card required.