Skip to content
Exercises · Q3

Q.A bank 'xyz' wants to know about its popularity among the residents of a city 'ABC' on the basis of number of bank accounts each family has and the average monthly account balance of each person. Briefly describe the steps to be taken for collecting data and what results can be checked through processing of the collected data.

Haryana BsehTextbookSubjective· 3mImportance★★★★★est
10% · 10/97 Questions
✓ Free question

The bank must first define its target population (families in city ABC), then collect data through a survey or census on the number of accounts per family and average monthly balance per person, and finally process this data using statistical measures (mean, median, mode, range, standard deviation) and visualizations (histograms, bar charts) to understand account distribution and balance patterns.

This is a plain theory question — it asks you to describe the steps for data collection and the results that can be obtained from processing that data. There is no code, no MCQ, no case study. The answer should be structured as clear explanatory paragraphs.

The Idea: Data Collection and Processing as a Two-Phase Process

Data collection and data processing are two sides of the same coin. Collection is about gathering raw facts; processing is about extracting meaning from them. For the bank 'xyz', the goal is to understand how many accounts families hold and what their average balances are — this tells the bank about its market penetration and the financial health of its customers.

Phase 1: Data Collection Steps

Step 1: Define the target population.

The bank must clearly specify that the target is all families residing in city 'ABC'. A "family" needs an operational definition — for example, a group of people living together and sharing a common kitchen. This avoids ambiguity.

Step 2: Choose the data collection method.

Two main options exist:

  • Census: Collect data from every family in the city. This gives perfect accuracy but is expensive and time-consuming.
  • Sample survey: Select a representative subset of families (using random sampling, stratified sampling by locality, etc.). This is more practical for a large city.

Step 3: Design the data collection instrument.

A questionnaire or interview schedule should capture:

  • Family identifier (to avoid duplication)
  • Number of bank accounts held by the family (across all banks, or specifically at 'xyz'? The question says "number of bank accounts each family has" — likely all accounts, but the bank may want to distinguish its own accounts)
  • For each person in the family, the average monthly account balance (over, say, the last 6 months)

Step 4: Pilot test the instrument.

Test the questionnaire on a small group to catch confusing questions, missing options, or data entry errors.

Step 5: Train the data collectors.

Ensure field staff understand how to ask questions neutrally and record responses accurately.

Step 6: Execute the data collection.

Deploy the survey, monitor progress, and handle non-response (families that refuse or are unreachable).

Step 7: Validate and clean the data.

Check for inconsistencies — for example, a family reporting 10 accounts but zero balance, or a person's balance being negative. Remove or correct such records.

Phase 2: Processing and Results

Once the raw data is collected, the bank can process it to answer specific questions. Here are the key results that can be checked:

Result 1: Distribution of number of accounts per family.

  • Calculate the mean (average accounts per family), median (the middle value), and mode (most common number of accounts).
  • A histogram showing the frequency of families with 1, 2, 3, … accounts reveals whether most families have a single account or multiple accounts.
  • If the mean is close to 1, the bank has low penetration; if it's higher, families are using multiple banks.

Result 2: Distribution of average monthly balances per person.

  • Compute the mean balance, median balance, and standard deviation (spread).
  • A histogram of balances (grouped into ranges like ₹0–₹1000, ₹1000–₹5000, etc.) shows the financial profile of residents.
  • The bank can identify what proportion of people have very low balances (potential churn risk) versus high balances (potential for premium services).

Result 3: Relationship between number of accounts and average balance.

  • A scatter plot with "number of accounts per family" on the x-axis and "average balance per person" on the y-axis can reveal patterns.
  • For example, families with more accounts might have higher average balances (wealthier families diversify), or the opposite (people open multiple accounts to keep small balances in each).

Result 4: Comparison across localities or demographics.

  • If the data includes locality or age group, the bank can compute averages per zone. This helps decide where to open new branches or target marketing.

Result 5: Identification of outliers.

  • Families with an unusually high number of accounts (say, 10+) or extremely high balances might be businesses or high-net-worth individuals — the bank can follow up for relationship management.
Watch out

A common mistake is to confuse data collection with data processing. Collection is about gathering raw facts (e.g., "Family X has 3 accounts, person Y has ₹12,000 average balance"). Processing is about summarizing and analyzing those facts (e.g., "The average number of accounts per family is 2.1"). The question asks for both steps — do not skip either phase.

Tip

For a real-world bank, the most actionable result is often the correlation between account count and balance. If families with more accounts also have higher balances, the bank might launch a "consolidate your accounts" campaign to attract those balances. If the correlation is negative, the bank needs to improve its service to retain customers who spread their money across banks.

✓Final answer

The steps for data collection are: define the target population (families in city ABC), choose a census or sample survey method, design a questionnaire capturing number of accounts per family and average monthly balance per person, pilot test, train collectors, execute the survey, and validate/clean the data. The results from processing include: distribution of accounts per family (mean, median, mode, histogram), distribution of balances per person (mean, median, standard deviation, histogram), relationship between account count and balance (scatter plot), locality-wise comparisons, and identification of outliers.

Unlock everything free for 14 days

  • Full step-by-step solutions
  • Concept-first explanations
  • Methods, shortcuts & mistakes
  • PYQ mapping + timed mock tests

Full access for 14 days. No credit card required.