Skip to content

Computer Science · Ch 7 — Understanding Data

Data Collection

7.2

Data Collection

Data collection is the first step in any data processing task. Before you can analyse or store data, you must first gather it. This does not always mean creating new data from scratch — it can also mean identifying data that already exists and bringing it into a usable form.

The textbook gives a clear example: a grocery store's sales data. Three different situations show what "data collection" can actually look like in practice.

First, the shopkeeper may have handwritten sales records in a diary or register. In this case, data collection means converting that physical record into a digital format — for instance, typing it into a spreadsheet. The data exists, but it is not yet machine-readable.

Second, the data may already be in a digital format, such as a CSV (comma separated values) file. Here, no conversion is needed; the data is ready for processing. You simply need to access the file.

Third, the shopkeeper may have never recorded any data at all, but now wants a software system to maintain sales and accounts. In this scenario, data collection involves designing a system — perhaps using Python — that can store and retrieve data from a CSV file or a database management system like MySQL. The data is generated and collected as the software is used.

Tip

Think and Reflect: When we click a photograph using our digital or mobile camera, does it have some metadata associated with it?

Data is continuously being generated from many sources. Every interaction with a digital medium — a click, a like, a purchase — adds to the growing volume of data. Hospitals collect patient data to improve their services. Shopping malls collect data on what items people buy. When analysed, this data can reveal useful patterns. For example, if analysis shows that bedsheets and groceries are frequently bought together, the shop owner might place bedsheets near the grocery section to increase sales. …