Prerequisite
An understanding of the Expected Value of discrete random variables. If you know that expected value is simply a probability-weighted average of all possible outcomes, you are ready to derive entropy from scratch.
1. Certainty and Zero Surprise
Imagine a coin-tossing scenario with the following outcomes and corresponding probabilities:
| Outcome | Probability |
|---|---|
| Heads (H) | 1.0 |
| Tails (T) | 0.0 |
These values indicate that the coin always shows up Heads (H).
If we know in advance that the outcome will always be Heads, we experience zero surprise when we see the actual outcome. It is always H. There is no uncertainty and no information gained when the coin lands.
Test this baseline in the simulation below. Notice what happens when you flip a coin with p = 1.0:
Experienced Surprise: 0.000 nats
2. Generalizing Surprise
More generally, say p is the probability of outcome H.
If we use X to denote a random variable which records the outcome of a coin toss, then X takes values in {H, T}. Then:
Pr(X = H) = p and Pr(X = T) = 1 − p
| X | Pr(X) |
|---|---|
| H | p |
| T | 1 − p |
How do we now generalize the "surprise"?
First, the surprise is now potentially non-zero, as the outcome is not pre-determined.
There could be any number of ways to quantify surprise, but we intuit some properties it must exhibit:
1. When an outcome is unlikely: the surprise upon its occurring should be
high.
2. When an outcome is quite likely: the surprise must be low.
3. In the extreme case where p = 1.0: the outcome H is certain, and the
surprise associated with it must be zero.
To satisfy these properties, we use ln(1/p) to quantify the surprise associated with an outcome of probability p.
The curve below shows how surprise varies with probability:
If coin lands Heads.
If coin lands Tails.
3. Surprise of Repeated Tosses
Given this formulation, over the course of many coin tosses, we experience:
• A surprise S(H) = ln(1/p) whenever the coin shows up Heads.
• A surprise S(T) = ln(1/(1−p)) whenever it shows up Tails.
| X | Pr(X) | S(X) [nats] |
|---|---|---|
| H | p | ln(1/p) |
| T | 1 − p | ln(1/(1−p)) |
Over repeated trials, what is the average surprise we experience?
4. Expected Surprise
To calculate the average surprise, we use the expected value of a discrete random variable:
E[Y] = ∑ y · Pr(Y = y)
Applying this to the surprise variable S(X):
This expected surprise is what is defined as entropy.
5. Why Entropy is a Measure of Chaos
Consider how this formula behaves across different values of p:
• When p = 1.0 (Guaranteed Heads): Entropy = −[ 1 · ln(1) + 0 ] = 0 nats.
• When p = 0.0 (Guaranteed Tails): Entropy = −[ 0 + 1 · ln(1) ] = 0 nats.
• When p = 0.5 (Fair Coin): Both outcomes are equally unpredictable. The formula
gives:
The table below shows how entropy varies with probability:
| Probability (p) | Entropy (nats) | System State |
|---|---|---|
| 0.0 | 0.000 | Certain Tails (Zero Chaos) |
| 0.1 | 0.325 | Biased (Low Chaos) |
| 0.5 | 0.693 | Fair Coin (Maximum Chaos) |
| 0.9 | 0.325 | Biased (Low Chaos) |
| 1.0 | 0.000 | Certain Heads (Zero Chaos) |
The simulation below plots this relationship across the full range of probabilities:
Maximum Chaos (Peak Uncertainty)
Entropy is a measure of the inherent chaos of the system, rather than an observer's knowledge. When a system is predictable (p = 1 or p = 0), there is no uncertainty and entropy is zero. When outcomes are evenly balanced (p = 0.5), uncertainty is at its maximum.
6. Generalizing Beyond the Coin
This principle extends directly to any discrete random variable with n possible outcomes:
To connect this to everyday intuition, consider searching for an object in a house—such as a book.
Suppose the book can be in one of four locations:
1. Bookshelf • 2. Under Sofa • 3. Kitchen • 4. Bathroom
Consider three different scenarios:
• Scenario A (Chaotic Environment): If the book is equally likely to be in any of the four spots (p = 0.25 each), the environment is unpredictable. You have no indication where to begin searching. Because the probability is uniformly distributed across all 4 locations, entropy is maximized:
• Scenario B (Orderly Environment): If the book is on the bookshelf 94% of the time (p = 0.94), with only a 2% chance of being in any other room, the environment is orderly. You can reliably look on the shelf first. Entropy drops to approximately 0.309 nats.
• Scenario C (Certain Location): If the book is always under the sofa with certainty (p = 1.0), the arrangement might be unconventional, but mathematically it contains no uncertainty. Since you know where it is in advance, surprise is zero, and entropy is strictly 0 nats.
In the simulation below, adjust the probabilities across the four locations to observe how the distribution affects entropy and inspect the exact terms contributing to the sum:
Sum of state contributions.
Uniform limit.
Summary
Reviewing the progression:
1. Certain outcomes deliver zero surprise. If an outcome is known in advance, observing it
provides no new information.
2. Rarity dictates surprise. An outcome with probability p carries surprise S(p) = ln(1/p) nats.
3. Entropy is expected surprise. Across repeated trials, the average surprise experienced is
−∑ p · ln(p) nats.
4. Entropy measures the inherent chaos of the system. Because the outcome of a
non-deterministic process cannot be determined in advance, each event carries some surprise.
This randomness is an intrinsic property of the system itself and does not depend on the observer. When outcomes are equally likely, entropy reaches its maximum value. When the system is biased or deterministic, entropy decreases toward zero.
In real-world data science and machine learning applications, we often do not know the true probability distribution of a system in advance and must approximate it using models. Measuring how far an estimated distribution diverges from the true distribution naturally leads to the concept of Cross Entropy.