Branching Processes & Coalescents

Kingman's coalescent

/ KING-man /

Branching processes run a population forward in time. But geneticists ask the reverse question: given the population alive today, what does its genealogy look like going backward — when did present-day individuals share common ancestors? Kingman's coalescent is the canonical answer: a Markov process, run backward in time, that starts from n separate lineages and progressively merges them in pairs until all coalesce into a single most recent common ancestor. It is the universal genealogy of a large, neutrally-evolving population, and it revolutionized population genetics by putting the data (a present-day sample) at the centre of the model.

The n-coalescent is a continuous-time Markov chain on the set partitions of {1, ..., n}, started from the partition into singletons. Each pair of distinct blocks (lineages) merges at rate 1, independently; so when there are k blocks, the total merger rate is C(k, 2) = k(k-1)/2 and a uniformly-chosen pair coalesces. The waiting time T_k while there are k lineages is therefore Exponential(C(k,2)) with mean 2/(k(k-1)). Two clean consequences: the total time back to the most recent common ancestor is T_MRCA = sum_(k=2)^n T_k, with E[T_MRCA] = 2(1 - 1/n), which is bounded by 2 even for huge samples — the deepest split happens fast in genealogical time, but most of the waiting is spent in the last few merges (E[T_2] = 1 alone, the time while just two lineages remain). The total branch length L_n = sum_(k=2)^n k T_k has E[L_n] = 2 sum_(j=1)^(n-1) 1/j about 2 log n, which (multiplied by the mutation rate) predicts the number of segregating sites in a sample. Crucially, only PAIRS merge at a time: simultaneous multiple mergers have probability zero, the defining feature of the Kingman (as opposed to Lambda or Xi) coalescent.

Why it matters: Kingman's coalescent is the backbone of modern statistical population genetics — it is the model under which one interprets DNA sequence variation, estimates effective population sizes, dates the most recent common ancestor (including 'mitochondrial Eve'), and detects selection and population growth as departures from its predictions. It arises as the universal genealogical limit (after a suitable time rescaling by the population size) of a vast class of forward models — Wright-Fisher, Moran, and many Cannings exchangeable models — provided offspring variance is finite and no single individual has a non-negligible fraction of the offspring. That last proviso is the honest hypothesis: when reproduction is highly skewed (a few individuals can have enormously many offspring, as in some marine species or strongly-selected sweeps), the binary-merger assumption fails and the correct limit is a Lambda- or Xi-coalescent with multiple and simultaneous mergers.

Sample n = 100 individuals. The expected time back to their single common ancestor is E[T_MRCA] = 2(1 - 1/100) = 1.98 coalescent-time units, barely more than for a sample of 10 (1.8). But more than half of that time, on average, is the final stretch when only two lineages remain (E[T_2] = 1) — most of the genealogy's depth lies in its last merge.

Lineages merge in pairs at rate C(k,2); the time to the common ancestor is bounded by 2 even for huge samples.

Kingman's coalescent allows only pairwise mergers (simultaneous multiple mergers have probability zero), and arises only when offspring variance is finite and no individual dominates reproduction. Highly skewed reproduction breaks it and requires Lambda/Xi coalescents.

Also called
the coalescentn-coalescentgenealogy backward in time金曼聚合過程聚合過程