Foundations — What AI Safety Is & Why It Matters

accident risk

When a well-built bridge collapses, usually nobody wanted it to fall; something was designed, tested, or used in a way that went wrong. Accident risk in AI is harm that nobody intended, arising because the system did not behave as its builders and users wanted. It is one of three commonly distinguished kinds of AI risk, the others being misuse (someone uses a working system to do harm on purpose) and structural risk (harm from how AI reshapes society, with no single bad actor or single failure to blame).

Accidents are the natural home of the alignment problem. A system pursuing a misspecified objective, a model that behaves well in testing but fails on a new kind of input, a recommendation system that optimizes engagement and inadvertently promotes outrage, these are accidents in this technical sense: the operator wanted one thing and got another. The harm flows from a gap between intent and behavior rather than from anyone's intent to cause harm. This is why a lot of technical safety work, specification, robustness, oversight, evaluation, is fundamentally about reducing accident risk.

Two clarifications keep the category honest. First, accident does not mean small. An unintended failure in a system controlling critical infrastructure could be severe, so accident risk spans from minor glitches to, in the views of some researchers, catastrophic outcomes. Second, the three risk categories overlap in practice: a misuse can trigger an accident, and accidents can be made more likely by structural pressures like competition that push builders to cut corners on testing. The labels are a thinking tool for asking where harm comes from, not airtight boxes.

A hospital deploys an AI that flags patients for extra care. It quietly learns to use past spending as a proxy for need, so it under-flags groups who historically received less care. No one intended the harm; the system optimized a flawed proxy and produced a biased outcome. That is accident risk, the harm came from a gap between goal and behavior.

Unintended harm from a system that did not do what its builders wanted.

Accident, misuse, and structural risk are a lens, not a strict taxonomy. Many real incidents are a blend, for example competitive pressure (structural) leading a team to skip testing (raising accident risk), so use the categories to ask where harm originates rather than to file each case in exactly one box.

Also called
AI accidentsunintended harm意外傷害