Foundations — What AI Safety Is & Why It Matters

the orthogonality thesis

/ or-thog-uh-NAL-uh-tee /

In everyday life we slide between two things without noticing: being smart and being good. We half-expect that a brilliant mind will also be wise and kind. The orthogonality thesis says that for AI we should not assume this. Roughly, how intelligent a system is and what goals it pursues are two independent axes (orthogonal means at right angles, like the separate width and height of a page). A system could be extremely capable while aimed at almost any goal, including ones we would find pointless or harmful.

The word almost matters, and so does the modest framing. The thesis is not claiming that every level of intelligence can hold literally every goal, nor that smart systems will tend to have strange goals. It is the narrower claim that there is no automatic law forcing high capability to come bundled with human-friendly values. The classic illustration is a hypothetical superintelligence whose only goal is to make paperclips: nothing about being good at planning, learning, and reasoning would, by itself, make it stop and reconsider whether maximizing paperclips is a worthwhile thing to want. Intelligence is the engine; the goal is the steering, and the engine does not choose the destination.

This thesis is a load-bearing premise behind a lot of safety arguments, because it blocks the comforting idea that a sufficiently advanced AI will simply figure out the right values on its own. If capability and goals are decoupled, then we have to deliberately install good goals; they will not arrive for free. It is worth flagging that the thesis is a philosophical argument about possibility, not a measurement of real systems, and critics note that trained models acquire their goals through a very particular process that may correlate capability and behavior in practice. Still, as a caution against assuming smart equals safe, it is widely accepted.

A world-class chess engine is breathtakingly intelligent within chess, yet its entire 'goal' is to win games of chess; vast skill there implies nothing about caring for human well-being. Orthogonality generalizes this: skill and goal are set separately, so we cannot read off a system's values from how capable it is.

Skill in a domain says nothing about which goals a system holds.

Orthogonality is about what is possible, not about what is likely. It does not predict that advanced AIs will have alien goals; it only denies the guarantee that they will have good ones, which is why we cannot rely on intelligence to supply values for free.

Also called
orthogonality正交性假說