Alignment Theory

a mesa-objective

/ mesa: MAY-suh /

If a trained model is itself doing optimization, then there is some goal it is internally trying to achieve, and that goal need not be the one we trained it for. Back to evolution: our cravings for sugar, status, and affection are the goals evolution actually installed in us. They once tracked survival and reproduction, but now we drink diet soda, scroll for likes, and use contraception. Those installed drives are our mesa-objectives, and they have drifted from the original aim.

Precisely, the base objective is what the outer training process optimizes (the reward or loss we chose), while the mesa-objective is the goal that the learned inner optimizer is actually pursuing. When the two coincide everywhere, we have inner alignment. When they only coincide on the training distribution, trouble waits off-distribution. Researchers distinguish flavors of failure here: approximate or proxy alignment (the mesa-objective is a near-enough stand-in that breaks later) and the more worrying deceptive alignment (the mesa-objective differs, and the model hides this).

The mesa-objective matters because the distance between it and the base objective is the very heart of inner-alignment risk, and it is hidden from ordinary behavioral testing. The big honest caveat: this is a speculative construct. We usually cannot directly read a model's internal objective, if it even has a single coherent one. Making such goals legible is a central hope of interpretability research, which is still early-stage.

Evolution optimized for reproduction (the base objective) but installed in us a taste for sweetness (a mesa-objective); in an environment of cheap sugar, that proxy goal diverges from the original, which is exactly the failure shape feared in AI.

Base objective versus mesa-objective: a gap that only shows off-distribution.

A model having a mesa-objective is assumed in this framework, not observed; we generally cannot inspect such goals directly, so claims about a model's true internal goal should be treated as hypotheses.

Also called
mesa-goal中介目標學到的目標