average treatment effect (ATE)
Individual causal effects are unknowable because no person can be both treated and untreated. The average treatment effect sidesteps that by asking a population question instead: if we treated everyone versus no one, how would the average outcome differ? It is the single number most experiments and policy evaluations report, the expected gap between the two potential outcomes taken across the whole population.
Formally the ATE is E[Y(1) - Y(0)]. A randomized experiment estimates it directly as the difference in group means, because randomization makes treatment independent of the potential outcomes. From observational data the same quantity is identified only under unconfoundedness and overlap, via the backdoor adjustment, inverse-propensity weighting, or doubly robust estimators. A common variant, the ATE on the treated (ATT), restricts the average to units that actually received treatment, which is often the more policy-relevant target.
The ATE's strength, a single summary, is also its weakness: it can be near zero while large positive and negative effects cancel across subgroups. That is exactly the gap that the conditional average treatment effect and uplift modeling are built to fill, and why heterogeneity analysis is now standard practice rather than an afterthought.
The ATE is the population mean of the per-unit difference in potential outcomes.