Compression: The Art of Squeezing Information
Imagine you have a huge suitcase to pack for a trip. You could just throw everything in loose — clothes, shoes, books — and it would barely close. Or you could fold the clothes neatly, roll the socks into the toes of the shoes, and stack the books flat. Suddenly, the same suitcase holds everything with room to spare. You haven't changed what you're carrying; you've just arranged it more efficiently.
That is compression.
In the world of data, compression is the process of representing the same information using fewer bits (the 0s and 1s a computer uses). The goal is to reduce the size of a file — a text, an image, a video, a song — without losing the essential meaning.
The Core Intuition: Redundancy
Why can we compress anything at all? Because most data contains redundancy — patterns, repetitions, or predictable structures that don't carry new information.
Consider this sentence: AAAAABBBBBCCCCC. It's 15 characters long. But you could describe it much more compactly: "5 A's, then 5 B's, then 5 C's." That description is a compressed version. The original had a lot of repetition (redundancy), and we exploited it.
Compression works by finding and removing redundancy. If a file had zero redundancy — every bit was completely unpredictable — it would be impossible to compress. That's why compressed files (like .zip) can't be compressed again to any meaningful degree.
Two Flavors of Compression
There are two main families, and they differ in one crucial way: do you get back exactly what you started with?
1. Lossless Compression
You get back the exact original data. Every single bit is restored. This is mandatory for text, code, spreadsheets, and medical images — where a single wrong bit could be catastrophic.
How it works: It finds and exploits statistical patterns. The most famous example is Run-Length Encoding (RLE). Instead of storing AAAAABBBBBCCCCC, you store 5A5B5C. Another is Huffman coding, where common symbols (like the letter 'e' in English) get short codes, and rare symbols (like 'z') get longer codes. The result is a smaller file on average.
Lossless compression ratio = compressed sizeoriginal size
A ratio of 2 means the compressed file is half the size.
2. Lossy Compression
You do not get back the exact original. Some information is thrown away permanently. In exchange, you get much smaller file sizes. This is used for photos (JPEG), music (MP3), and video (MP4) — where the human eye or ear won't notice the missing details.
How it works: It exploits the limitations of human perception. For an image, tiny color variations that your eye can't see are discarded. For audio, frequencies outside your hearing range are removed. The result is a file that looks or sounds nearly identical to the original but is drastically smaller.
Lossy compression is irreversible. Once you save a JPEG at low quality, you can never recover the original pixels. Always keep a lossless master copy (like a RAW photo or a WAV audio file) if you might need the full quality later.
The Precise Statement …