the space of probability measures
To say 'mu_n converges to mu' you must regard the measures themselves as points in some space, and to do convergence and compactness arguments you want that space to be a decent topological, even metric, space. So we promote the collection of all probability measures on a metric space S to a space in its own right, written P(S), and study its geometry. This shift of viewpoint, from individual measures to a space OF measures, is what lets the powerful machinery of topology and functional analysis act on limit theorems.
P(S) is the set of all Borel probability measures on S. We topologize it with the topology of weak convergence: a sequence (or net) mu_n -> mu in P(S) iff integral f dmu_n -> integral f dmu for all f in C_b(S). The fundamental structural fact is that this topology inherits the good properties of S. If S is separable and metrizable, then P(S) is separable and metrizable (for instance by the Levy-Prokhorov metric or the bounded-Lipschitz metric). If S is a Polish space (separable, completely metrizable), then P(S) is itself Polish. So the class of nice spaces is closed under the operation of forming the space of measures, which is exactly what you need to iterate constructions (measures on measures, as in de Finetti theory or empirical-measure limits).
P(S) is a convex set, mixtures t*mu + (1-t)*nu are again probability measures, and its extreme points are the Dirac masses delta_x, with x running over S. In fact the map x |-> delta_x embeds S homeomorphically into P(S), so S sits inside P(S) as the extreme boundary. Compactness in P(S) is governed by tightness via Prokhorov's theorem, and this is the practical payoff: questions about convergence of laws become questions about (relative) compactness of subsets of one fixed space P(S).
On S = R, the empirical measure of n iid samples, L_n = (1/n) sum_{i=1}^n delta_{X_i}, is a random point in P(R). The Glivenko-Cantelli theorem says L_n converges (almost surely, in P(R)) to the true law, and Sanov's theorem describes the large deviations of L_n away from it, both statements that only make sense once P(R) is itself a space.
The empirical measure as a random element of P(S): treating measures as points unlocks LLN and large-deviation statements about laws.
Metrizability of P(S) needs S separable; for non-separable S the weak topology can fail to be metrizable, and one must work with nets or restrict to separable supports. In probability we almost always assume Polish S precisely to avoid this.