Informatics Practices · Ch 2 — Emerging Trends
Big Data
Big Data
Technology has made its way into almost every sphere of our lives, and one consequence is that data is now being produced at a colossal rate. The core idea of this section is simple: when data sets grow so enormous in volume and so complex in nature that traditional data processing tools can no longer process and analyse them, we call such data Big Data. Big data is not simply "a lot of data" — it is data of a scale and a character that breaks the tools we used before.
Where the flood comes from
Three facts set the scene:
- There are today over a billion Internet users in the world.
- A majority of the world's web traffic comes from smartphones — the device most of us carry everywhere is also the device generating most of the traffic.
- At the current pace, around 2.5 quintillion bytes of data are created each day (a quintillion is a one followed by eighteen zeros — a number far beyond everyday intuition). Figure 2.8 illustrates the sources of this flood, showing how much data familiar platforms generate in a single minute.
And the pace is not steady — it keeps increasing with the continuous evolution of the Internet of Things (IoT), as more and more connected devices come online and generate data of their own.
Why traditional tools fail
Two properties together defeat conventional data processing tools:
- The data is voluminous. The sheer quantity outstrips what traditional data processing systems were designed to handle.
- The data is unstructured. Much of it does not sit in neat rows and columns the way a conventional database expects. Consider what we actually produce online: social media posts, instant messages and chats, the photographs we share through various sites, our tweets, blog articles, news items, opinion polls and the comments under them, audio and video chats. None of this has a fixed, predictable structure, so a tool built for orderly tables cannot digest it.
Either property alone is a nuisance; together they define a genuinely new kind of data problem.
Big data means big challenges
Big data does not describe only the data itself — it also names the whole family of difficulties that come with working at this scale. The challenges involved include:
- integration — bringing data together from many different sources into one usable whole
- storage — keeping such enormous quantities of data
- analysis — extracting meaning from it
- searching and querying — finding the right piece of information inside it
- processing and transfer — computing over the data and moving it from place to place
- visualisation — presenting it in a form people can actually grasp …
The figure is a circular infographic designed to look like a stopwatch — a deliberate visual choice, because everything it reports happens within a single sweep of sixty seconds. At the centre sits a circle carrying the key sentence: "In 60 Seconds these many data are generated". Arranged around this centre, like the markings of a watch face, are radial segments, each dedicated to one online platform. Every segment pairs the platform's logo with an approximate figure for what that platform generates in just one minute.
Reading around the dial, the segments report: over 104,300 Skype calls; 1.8 million snaps created on Snapchat; 204 million emails sent; 2.4 million search queries on Google; 46,200 posts uploaded on Instagram; 2.78 million video views on YouTube; 5 lakh page views on Wikipedia; 70,017 hours watched on Netflix; 342,000 apps downloaded from the App Store and Google Play; more than 120 new LinkedIn accounts; 15,000 GIFs sent via Messenger; 547,200 new tweets on Twitter; 44.4 million messages sent on WhatsApp; and 293,000 status updates on Facebook. As the caption itself says, all of these numbers are approximate. …