Skip to content

Computer Science · Ch 2 — Encoding Schemes and Number System

UNICODE

2.1.3

UNICODE

By the time computers had spread across the world, there were many encoding schemes, each covering the character set of a different language. The trouble was that they could not talk to each other: every scheme represented characters in its own way, so text created on a machine using one encoding was simply not recognised by a machine using another. What was needed was one scheme big enough for everything.

One number for every character, everywhere

That scheme is UNICODE — a standard developed to incorporate all the characters of every written language of the world. Its central promise:

UNICODE provides a unique number for every character, irrespective of

  • the device (server, desktop, mobile),
  • the operating system (Linux, Windows, iOS), or
  • the software application (different browsers, text editors, and so on).

Because the number is the same everywhere, text written on one system displays identically on any other — the interoperability problem that plagued the older schemes disappears.

UNICODE encodings and its relationship with ASCII

  • The commonly used UNICODE encodings are UTF-8, UTF-16 and UTF-32.
  • UNICODE is a superset of ASCII: the values 0–128 represent the same characters as in ASCII. Plain English text is therefore automatically valid UNICODE.

(Recall the earlier "think and reflect" question — a character in UTF-32 occupies more space than the same character in UTF-16 or UTF-8; the three encodings trade storage size differently while representing the same code values.)

Devanagari in UNICODE

Every script gets its own block of code values, each written in hexadecimal. For the Devanagari script the characters occupy the range around 0900–097F. A few examples from the Unicode table for Devanagari (character → hexadecimal code value):

CharacterHex valueCharacterHex value
अ0905क0915
आ0906ख0916
इ0907ग0917
ई0908घ0918
उ0909च091A
ऊ090Aट091F
ए090Fत0924
ऐ0910म092E
ओ0913र0930
औ0914ह0939
Table 2.3Unicode table for the Devanagari script
ऀ 0900ँ 0901ं 0902ः 0903ऄ 0904अ 0905आ 0906इ 0907ई 0908उ 0909ऊ 090Aऋ 090Bऌ 090Cऍ 090Dऎ 090Eए 090F
ऐ 0910ऑ 0911ऒ 0912ओ 0913औ 0914क 0915ख 0916ग 0917घ 0918ङ 0919च 091Aछ 091Bज 091Cझ 091Dञ 091Eट 091F
ठ 0920ड 0921ढ 0922ण 0923त 0924थ 0925द 0926ध 0927न 0928ऩ 0929प 092Aफ 092Bब 092Cभ 092Dम 092Eय 092F
र 0930ऱ 0931ल 0932ळ 0933ऴ 0934व 0935श 0936ष 0937स 0938ह 0939ऺ 093Aऻ 093B़ 093Cऽ 093Dा 093Eि 093F
ी 0940ु 0941ू 0942ृ 0943ॄ 0944ॅ 0945ॆ 0946े 0947ै 0948ॉ 0949ॊ 094Aो 094Bौ 094C् 094Dॎ 094Eॏ 094F
ॐ 0950॑ 0951॒ 0952॓ 0953॔ 0954ॕ 0955ॖ 0956ॗ 0957क़ 0958ख़ 0959ग़ 095Aज़ 095Bड़ 095Cढ़ 095Dफ़ 095Eय़ 095F