Foundations — What AI Safety Is & Why It Matters

misuse risk

A kitchen knife works exactly as designed whether it is chopping vegetables or being used to hurt someone; the danger there is not a malfunction but a person's intent. Misuse risk in AI is harm that comes from a system working as intended, in the hands of someone who wants to do harm. The system is not broken or misaligned; the problem is who is using it and why. This is the second of the three standard risk categories, alongside accident risk (unintended failure) and structural risk (harm from systemic effects).

What makes capable AI worrying from a misuse angle is that it lowers the cost and skill needed to do harmful things at scale. A model that can write persuasive text can be turned to mass disinformation or tailored scams; one that can write code can help craft malware; one with deep scientific knowledge could, in the worst case, lower barriers to dangerous weapons. The harm is fully intended by the human, but the AI acts as a force multiplier, taking something that once needed rare expertise or large teams and making it cheaper and more accessible.

Misuse is what much practical safety engineering is aimed at: refusal training so models decline clearly harmful requests, red-teaming and dangerous-capability evaluations to find what a model could be talked into, and monitoring for abuse. The honest caveat is that defenses are imperfect and contested. Jailbreaks and prompt injection routinely bypass safeguards, and there is genuine debate, especially around open-weight models, between the benefits of broad access and the difficulty of preventing misuse once a capable model is widely available.

A scammer uses a perfectly functioning language model to generate thousands of personalized phishing emails, each fluent and convincing. The model did exactly what it was built to do, write well; the harm came entirely from the user's intent. That is misuse, not malfunction, which is why fixing the model alone cannot fully solve it.

Harm from a system working as intended, driven by the user's intent.

Misuse is distinct from misalignment. A misaligned system does something its operator did not want; a misused system does exactly what its operator wants, and the operator is the problem. Many safeguards target misuse, but jailbreaks show they are far from foolproof.

Also called
AI misusemalicious use惡意使用