Foundations — What AI Safety Is & Why It Matters

safety-washing

You may have heard of greenwashing, where a company markets itself as environmentally friendly while doing little to back it up. Safety-washing is the same move in AI: presenting an organization or a product as more safe, careful, or responsible than it actually is, using the language and imagery of safety for reassurance or reputation rather than for real risk reduction. The label is borrowed from the gap between what is said and what is done.

It shows up in recognizable ways. A lab might publish a glossy responsibility framework with no binding commitments, point to a safety team that has little authority over launch decisions, or tout a benchmark score on a test that mostly measures general capability while branding it a safety result. A subtle and well-documented trap is exactly this last one: if a so-called safety metric is strongly correlated with raw capability, then simply building a more capable model raises the number, and the apparent safety gain is an artifact of capability, not evidence that the system is meaningfully safer.

Naming safety-washing matters because safety claims are hard for outsiders to verify, which creates room for theater. The honest difficulty is that the same activities can be genuine or performative, publishing a framework, running evaluations, having a safety team, so the accusation can be unfair as well as warranted. The practical antidotes are the things that are hard to fake: binding commitments such as responsible scaling policies, independent third-party evaluation, and transparency that lets others check the claims, rather than taking reassuring language at face value.

A company headlines that its new model 'scores 30 percent higher on safety.' On inspection, the benchmark mostly rewards general fluency and reasoning, which the bigger model improved anyway. The safety number rose as a side effect of capability, not because the model is harder to misuse, a textbook case of a metric that invites safety-washing.

When a 'safety' score really tracks capability, the apparent gain is an illusion.

Calling something safety-washing can be a fair critique or an unfair smear, since genuine and performative safety work can look alike from outside. The reliable distinguisher is verifiability, binding commitments and independent checks, not the warmth of the language used.

Also called
safetywashingethics-washing安全漂白安全作秀