Skip to content

Informatics Practices · Ch 2 — Emerging Trends

Characteristics of Big Data

2.3.1

Characteristics of Big Data

What makes a data set "big data" rather than merely a large database? Big data is distinguished from traditional data by five characteristics, shown in Figure 2.9 and easily remembered as the five V's: Volume, Velocity, Variety, Veracity and Value. A data set exhibiting these properties demands the special tools and methods of big data; one that lacks them can be handled by conventional means.

Volume

The most prominent characteristic of big data is its enormous size. The working test is practical: if a particular data set is of such large size that it is difficult to process with traditional DBMS tools, it can be termed big data. Volume is the first thing the very name "big data" points to.

Velocity

Velocity is the rate at which the data under consideration is being generated and stored. Big data is not only large — it grows at an exponentially higher rate than traditional data sets. A shop's sales register gains a few entries a day; the streams of posts, messages and shared media that make up big data pour in continuously, and ever faster.

Variety

Variety asserts that the data set contains varied kinds of data: structured, semi-structured and unstructured. Text, images, videos and web pages are all examples. Traditional systems expect data in one predictable format; big data arrives in many formats at once, and a large share of it has no fixed structure at all.

Veracity

Veracity refers to the trustworthiness of the data. Big data can sometimes be inconsistent, biased or noisy; there can be abnormalities in the data, or issues with the methods used to collect it. This matters because processing incorrect data gives wrong results and can mislead the interpretations drawn from it — an analysis is only as reliable as the data feeding it.

Value

Big data is not just a big pile of data; it can also hold hidden patterns and useful knowledge of high business value. But extracting that knowledge is not free — processing big data requires an investment of resources. The sensible discipline is therefore to make a preliminary enquiry first: assess whether the data set genuinely has potential for value discovery before committing resources, or else the whole effort could be in vain.

The five V's at a glance

CharacteristicWhat it means for big data
VolumeEnormous size — difficult to process with traditional DBMS tools
VelocityGenerated and stored at an exponentially higher rate than traditional data
VarietyA mix of structured, semi-structured and unstructured data (text, images, videos, web pages)
VeracityTrustworthiness — the data may be inconsistent, biased, noisy or abnormal
Figure 2.9Characteristics of big data

The figure is a cluster diagram built around one central circle labelled "BIG DATA". Five smaller circles press directly against this centre, carrying the labels Volume, Velocity, Variety, Veracity and Value. There are no arrows and no ordering among them — each circle touches the hub equally — so the layout itself makes the point: these are five simultaneous, defining characteristics that a data set exhibits together, not stages of a process. Collectively, they are what separate big data from traditional data.

Volume, the most prominent characteristic, is about sheer size: when a data set becomes so large that traditional DBMS tools find it difficult to process, it can be called big data. Velocity is the rate at which that data is being generated and stored — big data grows at an exponentially higher rate than traditional data sets ever did.

Variety records that the data is mixed in form — structured, semi-structured and unstructured — covering text, images, videos, web pages and similar kinds. Veracity concerns trustworthiness: big data can be inconsistent, biased or noisy, can contain abnormalities, or can suffer from problems in the way it was collected, and processing such faulty data yields wrong results or misleading interpretations. …