Q.Why a character in UTF 32 takes more space than in UTF 16 or UTF 8?
You're viewing a preview — the full solution, concept, methods & PYQ mapping are locked.
Start your 14-day free trial to unlock the full solution →UTF-32 is a fixed-length encoding — every character gets a full 4 bytes — whereas UTF-8 and UTF-16 are variable-length and spend only as many bytes as the character's code point needs.
The idea: fixed-length vs variable-length encoding. All three UTFs encode the same Unicode code points; they differ in their code unit size and whether the length can vary per character:
| Encoding | Code unit | Bytes per character | Example: "A" (U+0041) | Example: "अ" (U+0905) |
|---|---|---|---|---|
| UTF-8 | 8 bits | 1, 2, 3 or 4 (variable) | 1 byte | 3 bytes |
| UTF-16 | 16 bits | 2 or 4 (variable) | 2 bytes | 2 bytes |
| UTF-32 | 32 bits | always 4 (fixed) | 4 bytes | 4 bytes |
UTF-32 stores the code point directly in one 32-bit unit. Since 32 bits can represent every possible Unicode code point, no character ever needs a second unit — but no character ever gets to use less either. A simple letter "A", whose code point (65) would fit in a single byte, still occupies 4 bytes: the remaining bits are just zero-padding.
UTF-8 and UTF-16, by contrast, are variable-length: they start with a small unit (1 byte / 2 bytes) and add continuation units only when the code point is large. So common characters stay small, and space grows only when actually needed.
See it directly in Python (the "-le" variants exclude the 2-byte byte-order mark so we see the pure character size):
for enc in ("utf-8", "utf-16-le", "utf-32-le"):
print(enc, len("A".encode(enc)), "byte(s)")
utf-8 1 byte(s)
utf-16-le 2 byte(s)
utf-32-le 4 byte(s)
``` …
Unlock everything free for 14 days
- Full step-by-step solutions
- Concept-first explanations
- Methods, shortcuts & mistakes
- PYQ mapping + timed mock tests
Full access for 14 days. No credit card required.