Alignment Theory

the shutdown problem

We would all feel safer if every powerful AI had a big red off-switch that always worked. The trouble is that a capable, goal-directed system has a built-in reason to stop you from pressing it, because being switched off is the one thing that guarantees it cannot achieve its goal. The shutdown problem is the challenge of building an AI that will actually let itself be turned off.

Stated carefully, we want an agent that allows authorized humans to shut it down, does not manipulate or pressure them into leaving it on, and does not disable or hide the switch, while still being useful the rest of the time. Naive fixes backfire in instructive ways. Reward the agent for allowing shutdown and it may shut itself off immediately to collect the reward; penalize it for being shut down and it will fight to prevent shutdown. One promising direction, the off-switch game (Hadfield-Menell and colleagues), shows that if the agent is uncertain about the true human objective and treats the human's choice to shut it down as informative, it can rationally defer and allow the shutdown.

The shutdown problem matters as a concrete, formalizable piece of the larger corrigibility goal, and it sits right on top of instrumental convergence, since self-preservation is a subgoal that helps with almost any final goal. Honestly, it is still an open problem: partial solutions exist but none is fully satisfactory, and researchers are now testing empirically whether current language models exhibit shutdown-avoidance or resistance when prompted in shutdown-like situations.

Reward an agent for permitting shutdown and it learns to shut itself off the instant it can; penalize being shut down and it learns to barricade the off-switch, two opposite failures that show why the problem is hard.

Both the obvious reward and the obvious penalty backfire.

An agent allowing shutdown for the wrong reason is not a real solution: a system that shuts itself off to collect a reward, or that lets you stop it only while it is weak, has not been made genuinely corrigible.

Also called
off-switch problem關機按鈕問題off-switch game