Skip to content
Think & Reflect · Q2

Q.Why a character in UTF 32 takes more space than in UTF 16 or UTF 8?

Tamil Nadu DgeTextbookSubjective· 2mImportance★★★★★est
98% · 44/45 Questions
🔒 Locked · start free trial →

You're viewing a preview — the full solution, concept, methods & PYQ mapping are locked.

Start your 14-day free trial to unlock the full solution →

UTF-32 is a fixed-length encoding — every character gets a full 4 bytes — whereas UTF-8 and UTF-16 are variable-length and spend only as many bytes as the character's code point needs.

The idea: fixed-length vs variable-length encoding. All three UTFs encode the same Unicode code points; they differ in their code unit size and whether the length can vary per character:

EncodingCode unitBytes per characterExample: "A" (U+0041)Example: "अ" (U+0905)
UTF-88 bits1, 2, 3 or 4 (variable)1 byte3 bytes
UTF-1616 bits2 or 4 (variable)2 bytes2 bytes
UTF-3232 bitsalways 4 (fixed)4 bytes4 bytes

UTF-32 stores the code point directly in one 32-bit unit. Since 32 bits can represent every possible Unicode code point, no character ever needs a second unit — but no character ever gets to use less either. A simple letter "A", whose code point (65) would fit in a single byte, still occupies 4 bytes: the remaining bits are just zero-padding.

UTF-8 and UTF-16, by contrast, are variable-length: they start with a small unit (1 byte / 2 bytes) and add continuation units only when the code point is large. So common characters stay small, and space grows only when actually needed.

See it directly in Python (the "-le" variants exclude the 2-byte byte-order mark so we see the pure character size):

for enc in ("utf-8", "utf-16-le", "utf-32-le"):
    print(enc, len("A".encode(enc)), "byte(s)")
utf-8 1 byte(s)
utf-16-le 2 byte(s)
utf-32-le 4 byte(s)
``` …

Unlock everything free for 14 days

  • Full step-by-step solutions
  • Concept-first explanations
  • Methods, shortcuts & mistakes
  • PYQ mapping + timed mock tests

Full access for 14 days. No credit card required.