Character Revelation — The Intuition
Imagine you are solving a puzzle. You have a box, and inside that box is a number. You cannot see the number, but you can ask yes-or-no questions about it. Each question gives you one piece of information — one bit — about what is inside.
Now suppose the box contains a single letter from the English alphabet. There are 26 possibilities. If you ask "Is it A?" and the answer is no, you have eliminated only one option. That is a very inefficient question — it gives you very little information. A better question might be "Is it in the first half of the alphabet?" That splits the 26 possibilities into two roughly equal groups of 13. The answer cuts down the possibilities by half. That is a much more informative question.
The idea of character revelation is about how much information you gain when you learn the exact identity of something — when the box is opened and you see the letter itself. It is the ultimate answer to all your questions.
The Precise Statement
In information theory, character revelation is the amount of information gained when an unknown outcome is fully revealed — that is, when you learn exactly which event from a set of possibilities actually occurred.
If you have a random variable X that can take values x1,x2,…,xn with probabilities p1,p2,…,pn, then the information gained when you learn that X=xi is:
I(xi)=log2(pi1)=−log2(pi)
This quantity is measured in bits.
The logarithm base 2 is used because we measure information in binary digits (bits). If you used natural log, the unit would be nats; if base 10, hartleys.
Why This Makes Sense
Think back to the alphabet example. If the letter is equally likely to be any of the 26, then pi=261 for each letter. The information gained when you learn the letter is:
I=−log2(261)=log2(26)≈4.7 bits
That means learning the exact letter gives you about 4.7 bits of information. Why 4.7? Because you could have asked 4–5 well-chosen yes-or-no questions to identify it — and 4.7 is the average number of such questions needed if you ask optimally.
Now consider a case where one outcome is almost certain. Suppose a biased coin lands heads with probability 0.99 and tails with probability 0.01. If you see heads, the information gained is:
I(heads)=−log2(0.99)≈0.014 bits
That is very little — you already expected heads, so learning it happened is hardly surprising. But if you see tails:
I(tails)=−log2(0.01)≈6.64 bits
That is a lot of information — a rare event carries much more surprise.
A common mistake is to think that character revelation is the same as entropy. Entropy is the average character revelation over all possible outcomes — it is the expected value of I(xi). Character revelation is the information from a specific outcome, not the average.
The Core Idea in One Sentence
Character revelation measures how surprising an outcome is: the less likely the outcome, the more information you gain when you learn it happened.