Governance & Catastrophic Risk

an AI safety institute

When new medicines or aircraft appear, governments do not just trust the manufacturer's word that they are safe; they stand up public bodies with the expertise to test and check. An AI safety institute is the equivalent for advanced AI: a government-backed organization whose job is to understand, measure, and help manage the risks of the most capable systems.

Concretely, these institutes typically build technical teams that evaluate frontier models for dangerous capabilities (for example, whether a model could meaningfully help someone build a weapon or run a cyberattack), develop shared testing standards and methods, run red-teaming exercises, and advise their governments. Several were created starting in 2023, notably in the United Kingdom and the United States, and a number of countries have since formed similar bodies and begun coordinating across borders. Crucially, most are evaluators and standard-setters rather than regulators: they often test and advise but do not, by themselves, license or ban a model.

AI safety institutes matter because they give the public side of governance real technical muscle, instead of leaving evaluation entirely to the companies that build the models. But their power and independence vary a lot. Some have privileged early access to models; others rely on voluntary cooperation. They face hard problems: hiring scarce talent away from industry, getting deep enough access to test properly, avoiding capture by the firms they assess, and the fact that passing their tests shows safety on what was tested, not a guarantee of safety in the wild.

Before a lab releases its newest model, a national AI safety institute is given early access, runs structured tests for biology and cyber misuse, finds one worrying capability, and reports it so the lab can add safeguards before launch.

An institute gives governments hands-on technical capacity, so safety claims can be checked, not just believed.

Most AI safety institutes evaluate and advise; they are usually not the body that licenses or bans a model. Their independence and access vary widely, and a clean bill of health covers the tested distribution, not every real-world use.

Also called
AISIAI safety institute