Mathematical Foundations

entropy

/ EN-truh-pee /

Entropy measures how uncertain or surprising a situation is, on average. A coin you already know will land heads carries zero entropy — no suspense, nothing to learn. A fair coin carries more, because either outcome is a genuine surprise. A fair die more still. In short, entropy is the average amount of surprise you should expect before an outcome is revealed, and the more evenly the possibilities are spread, the higher it climbs.

Information theory makes this precise and measures it in bits. One bit is exactly the uncertainty of a single fair coin flip — the answer to one well-chosen yes/no question. A choice among four equally likely things carries two bits; among eight, three. Rare, surprising outcomes carry more information than common, expected ones — which is why telling someone 'the sun rose today' conveys almost nothing, while 'it snowed in the desert' conveys a great deal.

Claude Shannon introduced entropy in 1948, founding the entire field of information theory, and it quietly underpins all of digital life. Entropy sets the hard floor on how much you can compress a file without losing anything — you cannot squeeze out uncertainty that genuinely exists. In machine learning it is the parent concept behind cross-entropy and KL divergence, the loss functions that train most classifiers. A frequent confusion worth clearing up: this information-theory entropy is a cousin, not a twin, of the entropy in physics; they share deep mathematical roots but answer different questions.

Imagine guessing tomorrow's weather. In a place where it is sunny 99 days out of 100, the forecast holds almost no surprise — low entropy; you barely need to ask. In a place split evenly among sun, rain, cloud, and snow, every day is a genuine four-way toss-up — high entropy, two full bits of uncertainty to resolve.

Near-certain weather carries almost no entropy; a true four-way coin flip carries the most — uncertainty spread evenly is uncertainty maximized.

Entropy peaks when outcomes are all equally likely and drops to zero when one outcome is certain — it scores the uncertainty of a situation, not how 'disordered' something looks. Shannon's 1948 paper named this and launched the digital age.

Also called
信息熵資訊熵Shannon entropy香农熵