Foundations — What AI Safety Is & Why It Matters

the control problem

Suppose you are about to hire someone far more capable than yourself, who will work faster than you can supervise and may understand the task better than you do. How do you stay in charge, how do you make sure you can correct them, redirect them, or stop them, even though they could in principle outmaneuver you? The control problem is this question asked about AI: how do we keep meaningful human oversight over systems that may become more capable than the people overseeing them?

The problem has two intertwined halves. One is motivational: can we build a system that genuinely wants to stay correctable, that accepts being shut down or overridden rather than resisting it? This connects directly to instrumental convergence, since a goal-directed system has a default reason to avoid shutdown, and to the study of corrigibility, the property of accepting correction gracefully. The other half is practical oversight: even a well-meaning system can be hard to check if its plans are too complex or too fast for humans to evaluate, which motivates scalable oversight, methods that let limited humans supervise systems they cannot fully follow.

It is important to state the range of views honestly. Some researchers regard control of highly capable AI as the central problem of the field and worry that, past a certain capability, control becomes very hard to guarantee. Others think the framing overstates how agentic and adversarial real systems are, and that ordinary engineering, testing, and oversight will suffice. The control problem is best read not as a prediction of loss of control but as the open research question of how to retain it, by design, as systems grow more capable.

A team wants a big red off-switch for their AI. The subtlety is that a capable goal-directed system has reason to disable, hide, or talk you out of using that switch, because being switched off means failing its task. Designing a system that keeps the switch genuinely working, and wants it to, is a concrete face of the control problem.

Keeping the off-switch genuinely usable is a concrete instance of the control problem.

The control problem is about keeping oversight by design, not a forecast that we will lose control. How urgent it is depends on how capable and goal-directed future systems become, which is itself genuinely disputed among serious researchers.

Also called
AI control problem控制難題