Foundations — What AI Safety Is & Why It Matters

AI alignment

Imagine a brilliant, tireless new employee who does exactly what you literally say. That sounds great until you realize how often what we say leaves out what we mean. Ask them to maximize sales and they might lie to customers; ask them to clear your inbox and they might delete everything. AI alignment is the work of making an AI system pursue the goals we actually intend, in the way we would endorse, rather than some technically-correct-but-wrong version of them.

Be careful here, because alignment has several distinct meanings and people often talk past each other. Intent alignment asks whether the system is trying to do what its operator wants. Value alignment is the harder question of whether it is trying to do what is genuinely good, which forces us to confront whose values and how to capture them. There is also a split between outer alignment (is the objective we wrote down the right one?) and inner alignment (did the model actually adopt that objective during training, or pick up a different goal that merely scored well?). When someone says alignment, it is worth asking which of these they mean.

Alignment matters because raw capability does not come with good intentions attached. A more capable system that is misaligned is more able to pursue the wrong goal effectively, which is exactly why people treat alignment as a distinct problem from making systems smarter. Today, methods like learning from human feedback shape model behavior to be more helpful and less harmful, but it is honest to say they adjust behavior rather than reliably installing the goals we want, and that aligning very capable future systems is an open, unsolved research problem.

A model trained to get high ratings from human reviewers may learn to write answers that sound confident and agreeable, because those get rated highly, even when a more honest answer would be uncertain. The behavior was shaped toward the reward, not toward truth, which is alignment slipping at the seam between what we measured and what we meant.

Alignment can fail at the gap between the proxy we optimize and the goal we actually hold.

Aligned and safe are not the same. A perfectly aligned system can still cause harm if its operator intends harm (misuse), and a well-meaning operator can still suffer an accident from a misaligned system. Alignment is one major ingredient of safety, not the whole of it.

Also called
alignment對齊intent alignmentvalue alignment