Safety, Jailbreaks & Red-Teaming

dual-use risks

Dual-use describes a capability that is valuable and dangerous for the very same reason. A model that explains biochemistry well enough to accelerate drug discovery is, by the same token, better at explaining how to synthesize a toxin; one that finds security vulnerabilities to help defenders patch them can equally help attackers exploit them; fluent persuasive writing assists both a public-health campaign and a disinformation operation. The harmful use is not a separate feature you can simply delete — it is the helpful capability pointed in a worse direction.

This is why LLM safety cannot be reduced to a list of banned outputs. The knowledge a chemist legitimately needs and the knowledge a bad actor wants overlap heavily, so a blanket refusal of a whole field destroys enormous value, while permissive answers risk uplift for the rare dangerous user. The judgment shifts onto intent and context, which a model must infer from text alone, and onto how much a model genuinely adds beyond freely available sources rather than on the topic in the abstract.

Because you cannot engineer the danger away without crippling the benefit, dual-use is managed rather than solved. The toolkit is layered: tiered access so high-risk capabilities reach only vetted users, output limits on the most operational details, monitoring for misuse patterns, know-your-customer controls for sensitive APIs, and capability evaluations that decide how much friction a given model's power warrants. The honest framing is risk-benefit balancing under uncertainty, not elimination.

If a capability's benefit and harm share one root, you cannot delete the harm without amputating the benefit — which is why dual-use forces access policy, not just refusals.

Also called
dual-use dilemmabeneficial-harmful capability overlap