Series 02 · Information & Entropy

An Intuitive Guide To Entropy

From coin flips to lost books: how quantifying surprise turns the mystery of entropy into an intuitive measure of uncertainty.

Prerequisite

An understanding of the Expected Value of discrete random variables. If you know that expected value is simply a probability-weighted average of all possible outcomes, you are ready to derive entropy from scratch.

1. Certainty and Zero Surprise

Imagine a coin-tossing scenario with the following outcomes and corresponding probabilities:

Outcome Probability
Heads (H) 1.0
Tails (T) 0.0

These values indicate that the coin always shows up Heads (H).

If we know in advance that the outcome will always be Heads, we experience zero surprise when we see the actual outcome. It is always H. There is no uncertainty and no information gained when the coin lands.

Test this baseline in the simulation below. Notice what happens when you flip a coin with p = 1.0:

SIMULATION 01 · CERTAINTY BASELINE
Coin Toss and Surprise
Adjust the probability of Heads from 1.0 to 0.0 to observe how surprise changes as certainty varies.
Last Flip Result 0 tosses
Heads (H)

Experienced Surprise: 0.000 nats

Probability of Heads: Pr(X = H) 100% (1.00)

2. Generalizing Surprise

More generally, say p is the probability of outcome H.

If we use X to denote a random variable which records the outcome of a coin toss, then X takes values in {H, T}. Then:

Pr(X = H) = p      and      Pr(X = T) = 1 − p

X Pr(X)
H p
T 1 − p

How do we now generalize the "surprise"?

First, the surprise is now potentially non-zero, as the outcome is not pre-determined.

There could be any number of ways to quantify surprise, but we intuit some properties it must exhibit:

1. When an outcome is unlikely: the surprise upon its occurring should be high.
2. When an outcome is quite likely: the surprise must be low.
3. In the extreme case where p = 1.0: the outcome H is certain, and the surprise associated with it must be zero.

To satisfy these properties, we use ln(1/p) to quantify the surprise associated with an outcome of probability p.

Quantifying Surprise
S(p) = ln( 1 / p ) = −ln( p )  [nats]
For guaranteed outcomes with p = 1.0, ln(1/1) = 0. Outcomes with small values of p result in a large surprise. Using the natural logarithm measures surprise in nats.

The curve below shows how surprise varies with probability:

SIMULATION 02 · QUANTIFYING SURPRISE
The Surprise Curve: S(p) = ln(1/p)
Move the probability slider to observe how surprise scales inversely with probability.
Surprise S(H) p
0.223 nats

If coin lands Heads.

Surprise S(T) 1 − p
1.609 nats

If coin lands Tails.

Selected Probability (p) 0.80
S(p) = −ln(p)  •  S(1.0) = 0.000 nats

3. Surprise of Repeated Tosses

Given this formulation, over the course of many coin tosses, we experience:

• A surprise S(H) = ln(1/p) whenever the coin shows up Heads.
• A surprise S(T) = ln(1/(1−p)) whenever it shows up Tails.

X Pr(X) S(X) [nats]
H p ln(1/p)
T 1 − p ln(1/(1−p))

Over repeated trials, what is the average surprise we experience?

4. Expected Surprise

To calculate the average surprise, we use the expected value of a discrete random variable:

E[Y] = ∑ y · Pr(Y = y)

Applying this to the surprise variable S(X):

Expected Surprise & Shannon Entropy
E[ S(X) ] = p · ln( 1 / p ) + (1 − p) · ln( 1 / (1 − p) )
⟹ H(X) = − [ p · ln(p) + (1 − p) · ln(1 − p) ]  [nats]
Using ln(1/a) = −ln(a), the expected surprise defines the binary entropy function.

This expected surprise is what is defined as entropy.

5. Why Entropy is a Measure of Chaos

Consider how this formula behaves across different values of p:

• When p = 1.0 (Guaranteed Heads): Entropy = −[ 1 · ln(1) + 0 ] = 0 nats.
• When p = 0.0 (Guaranteed Tails): Entropy = −[ 0 + 1 · ln(1) ] = 0 nats.
• When p = 0.5 (Fair Coin): Both outcomes are equally unpredictable. The formula gives:

Fair Coin Maximum Entropy
H = − [ 0.5 · ln(0.5) + 0.5 · ln(0.5) ] = ln(2) ≈ 0.693 nats
At p = 0.5, entropy reaches its maximum value of ln(2) ≈ 0.693 nats (1 bit).

The table below shows how entropy varies with probability:

Probability (p) Entropy (nats) System State
0.0 0.000 Certain Tails (Zero Chaos)
0.1 0.325 Biased (Low Chaos)
0.5 0.693 Fair Coin (Maximum Chaos)
0.9 0.325 Biased (Low Chaos)
1.0 0.000 Certain Heads (Zero Chaos)

The simulation below plots this relationship across the full range of probabilities:

SIMULATION 03 · EXPECTED SURPRISE
The Entropy Curve: H(p)
Adjust the probability of Heads to observe where uncertainty is maximized.
Total System Entropy H(X) Expected Surprise
0.693 nats

Maximum Chaos (Peak Uncertainty)

Probability of Heads (p) 0.50
H = −∑ p · ln(p)

Entropy is a measure of the inherent chaos of the system, rather than an observer's knowledge. When a system is predictable (p = 1 or p = 0), there is no uncertainty and entropy is zero. When outcomes are evenly balanced (p = 0.5), uncertainty is at its maximum.

6. Generalizing Beyond the Coin

This principle extends directly to any discrete random variable with n possible outcomes:

Generalized Shannon Entropy
H(X) = − ∑ p_i · ln(p_i)  [nats]
Where p_i is the probability of each outcome, and the sum of all probabilities equals 1.

To connect this to everyday intuition, consider searching for an object in a house—such as a book.

Suppose the book can be in one of four locations:

1. Bookshelf  •  2. Under Sofa  •  3. Kitchen  •  4. Bathroom

Consider three different scenarios:

• Scenario A (Chaotic Environment): If the book is equally likely to be in any of the four spots (p = 0.25 each), the environment is unpredictable. You have no indication where to begin searching. Because the probability is uniformly distributed across all 4 locations, entropy is maximized:

Uniform 4-State Distribution
H = − 4 · [ 0.25 · ln(0.25) ] = ln(4) ≈ 1.386 nats
With 4 equally likely locations, each state contributes 0.25 × 1.386 = 0.347 nats, reaching the maximum possible entropy of ln(4) ≈ 1.386 nats.

• Scenario B (Orderly Environment): If the book is on the bookshelf 94% of the time (p = 0.94), with only a 2% chance of being in any other room, the environment is orderly. You can reliably look on the shelf first. Entropy drops to approximately 0.309 nats.

• Scenario C (Certain Location): If the book is always under the sofa with certainty (p = 1.0), the arrangement might be unconventional, but mathematically it contains no uncertainty. Since you know where it is in advance, surprise is zero, and entropy is strictly 0 nats.

In the simulation below, adjust the probabilities across the four locations to observe how the distribution affects entropy and inspect the exact terms contributing to the sum:

SIMULATION 04 · MULTI-STATE SYSTEM
Household Locations and Entropy
Adjust probabilities across locations to observe how uniformity maximizes entropy.
Current Entropy H
1.386 nats

Sum of state contributions.

Max Possible ln(4)
1.386 nats

Uniform limit.

Tracked Item
Distribution Presets
Live Component Summation: H(X) = ∑ p · S(p)
Uniform distribution maximizes entropy across the house (ln 4 ≈ 1.386 nats)

Summary

Reviewing the progression:

1. Certain outcomes deliver zero surprise. If an outcome is known in advance, observing it provides no new information.
2. Rarity dictates surprise. An outcome with probability p carries surprise S(p) = ln(1/p) nats.
3. Entropy is expected surprise. Across repeated trials, the average surprise experienced is −∑ p · ln(p) nats.
4. Entropy measures the inherent chaos of the system. Because the outcome of a non-deterministic process cannot be determined in advance, each event carries some surprise.

This randomness is an intrinsic property of the system itself and does not depend on the observer. When outcomes are equally likely, entropy reaches its maximum value. When the system is biased or deterministic, entropy decreases toward zero.

In real-world data science and machine learning applications, we often do not know the true probability distribution of a system in advance and must approximate it using models. Measuring how far an estimated distribution diverges from the true distribution naturally leads to the concept of Cross Entropy.