JOVANA
Explore Library Glossary Getting Started Three Levels Fields How it works Mission
Join the mission
All guides

Responsible Scaling Policies & Safety Institutions

Before a frontier model ships, what stops a launch-night rush from quietly waving through a worrying result? A responsible scaling policy — a pre-commitment that wires a dangerous-capability test to a fixed response — plus the safety institutions, evaluators, and emerging laws that try to keep that promise honest. This guide explains how the machinery works, and how much of it is proven versus hopeful.

The night before launch

Picture the night before a frontier lab ships its biggest model yet. The launch page is staged, the press embargo lifts at nine, and then an email lands from the team running the dangerous-capability evaluation: one test came back higher than expected on bio-uplift. Now everything depends on a question that has nothing to do with the model and everything to do with the company — was there a rule, written down in advance, that says what happens now? If there is, the path is already chosen. If there is not, the decision falls to whoever is most tired, most invested, or most senior in the room at midnight — which is exactly the moment you least want to be inventing your safety policy from scratch. This guide is about the rule written down in advance.

The plain idea is older than computing. Homer's Odysseus wanted to hear the Sirens' song without steering his ship onto the rocks, so before the danger arrived — while he was still thinking clearly — he had his crew tie him to the mast and ordered them to ignore any later command to be untied. That is a pre-commitment: you bind your future self, in a calm moment, because you know the heated moment will tempt you to do the wrong thing. A responsible scaling policy (RSP) is that mast for an AI lab. It is a public document that fixes, ahead of time, which capabilities would be too dangerous to deploy without specific safeguards — and what the lab will actually do when its own tests trip those wires — precisely so that a launch-night rationalization cannot quietly talk everyone out of caution.

Why would a company need to tie its own hands? Because of the trap you met earlier in this rung: racing dynamics. Each lab might genuinely want to be careful, but none wants to be the one that pauses while a less cautious competitor sprints ahead and takes the market. Left alone, that pressure grinds safety down toward whatever the least careful actor will accept. An RSP is a partial escape: by stating its red lines openly and in advance, a lab tries to turn 'being careful' from a costly unilateral sacrifice into a visible, comparable commitment — ideally one that pressures rivals to publish their own. The hope is a race to the top on safety rather than a race to the bottom; whether it works is one of the live debates we reach at the end.

What a responsible scaling policy actually is

Now the precise idea, one piece at a time. At its core an RSP is built from if-then commitments: statements of the form 'if our evaluations show capability X above threshold T, then we will apply safeguard S — or not deploy at all.' Stack a coherent set of these and you get a ladder of capability thresholds, each rung paired with a required level of protection. Crucially, the thresholds and the responses are written before any particular model is tested, so an uncomfortable result cannot be bargained away after the fact. The eval supplies the 'if'; the policy supplies the 'then'; together they convert a measurement into a decision that was made while everyone was calm.

The 'then' is not a single action but usually three, and it helps to separate them. First, deployment safeguards: what users are allowed to do with the model — content filters, refusal training, monitoring, restricted access, or holding a capability back entirely. Second, security safeguards: how well the model's own weights are protected from theft, because a frontier model whose weights leak can have its safety training stripped off by whoever steals it, making the deployment controls irrelevant. Third, the development decision itself: whether it is even safe to keep training or to deploy at all, up to and including pausing. A good RSP also names who decides, when re-evaluation is triggered (for instance, at every defined jump in effective training compute), and what happens if a threshold is crossed unexpectedly mid-training.

The cleverest mechanism in a mature RSP is a meta-commitment: define the safeguards for the next capability level before you build a model that reaches it, and commit not to cross the line until those safeguards exist. This is what stops the policy from being rewritten to match whatever the lab happened to ship. The flip side, increasingly, is a safety case: rather than just clearing a threshold, the lab must make a structured, evidence-backed argument that the model is safe enough to deploy or to keep training — a discipline borrowed from aviation and nuclear power, where you do not get to operate because you failed to prove you are dangerous, but only when you can positively argue that you are safe. That shift — from 'we found nothing alarming' to 'here is our affirmative case for safety' — is one of the most important directions the field is moving in.

  1. Set the thresholds and safeguards in advance — define each capability level (say, meaningful uplift to a bioweapon, or autonomous self-replication) and the deployment, security, and pause responses each one requires, before the model in question even exists.
  2. Evaluate as capability grows — re-run the dangerous-capability evals at defined checkpoints, such as each large jump in training compute, not just once at the end.
  3. Compare against the line — if the model stays below every threshold, deploy with standard safeguards; if it crosses one, the pre-written 'then' kicks in.
  4. Escalate the response — apply the stronger deployment and security standard the threshold demands, or, if the required safeguards do not yet exist, hold the model back until they do.
  5. Account for it publicly — disclose the result and the decision, ideally with an external evaluator or safety institute able to check the lab's own grading.

A worked example: ASL levels and the frameworks

The most developed example is Anthropic's RSP, first published in 2023 and revised since, organized around AI Safety Levels (ASL) — a deliberate echo of the biosafety levels (BSL-1 to BSL-4) that govern how hazardous a pathogen a lab may handle and under what containment. ASL-2 covers today's models: capable, but not meaningfully uplifting a determined misuser beyond what a search engine already offers. ASL-3 marks models that substantially raise the risk of catastrophic misuse — for example, giving real uplift on chemical, biological, radiological or nuclear weapons — or that show early dangerous autonomy, and it demands hardened deployment filters plus serious protection of the model's weights. ASL-4 and beyond are sketched for still more capable systems whose safeguards have not yet been fully worked out.

A concrete instance arrived in May 2025. When Anthropic released Claude Opus 4, it activated its ASL-3 standard for the first time — switching on stronger deployment safeguards and tighter security on the model's weights — not because it had proven the model crossed the threshold, but because its evaluations could not rule out that it had, and the policy says you treat that uncertainty by stepping up, not by waiting for certainty. Whatever you make of any single company, this is what a pre-commitment looks like when it actually bites: a measurable cost (more friction, more security work) accepted because a rule written in a calmer moment said it must be. It is, so far, one of the only public examples of an RSP changing a real launch rather than merely describing one.

Anthropic is not alone, and the convergence is the point. OpenAI's Preparedness Framework scores models across tracked risk categories — biological and chemical, cybersecurity, AI self-improvement — and requires specified safeguards before a model crossing a 'High' or 'Critical' threshold can be deployed or further developed. Google DeepMind's Frontier Safety Framework is built around 'critical capability levels' with matching mitigations. And at the 2024 Seoul AI summit, sixteen leading companies signed the Frontier AI Safety Commitments, each pledging to publish exactly this kind of if-then framework — thresholds at which risks would be 'deemed intolerable,' and what they would do, up to not deploying, when those thresholds are hit. In a few years the question moved from 'should labs have such a policy?' to 'is yours any good?'

AI SAFETY LEVELS   (capability  ->  required standard)

ASL-2   today's models
        no meaningful catastrophic uplift over existing tools
        ->  standard deployment + standard security

ASL-3   substantial uplift to CBRN misuse, OR early dangerous autonomy
        ->  hardened deployment (filters, monitoring, limited access)
        ->  strong weight security (resist theft by non-state attackers)

ASL-4+  qualitatively more dangerous capability
        ->  safeguards NOT yet fully specified
        ->  COMMITMENT: define and build them BEFORE training such a model

rule:  thresholds + responses are fixed in ADVANCE,
       so an uncomfortable eval result cannot be rationalized away later
A schematic of an ASL-style ladder (loosely after Anthropic's framework): each capability rung is wired to a required deployment and security standard, and the top rung is a promise to build the safeguards before building the model. Real policies are far more detailed.

Who checks the homework: safety institutions

Every policy so far has a structural weakness you have probably already spotted: the lab writes the rules, runs the tests, and grades itself. A company under deadline and competitive pressure marking its own safety homework is the textbook setup for safety-washing — the appearance of rigor without the substance. The first answer is third-party evaluation: independent organizations such as METR and Apollo Research, given early access to a model, run their own dangerous-capability and propensity tests so the lab's conclusions can be checked by someone with no incentive to ship. An outside evaluator can also resist the quiet pressure to under-elicit — to not try too hard to find the scary capability — that a team racing toward launch is subject to whether it admits it or not.

The bigger institutional answer is the rise of government AI safety institutes. The UK launched the first in November 2023, around the Bletchley Park summit; the United States set one up inside its standards agency NIST; and Japan, Singapore, Canada, the EU and others followed, linked through an International Network of AI Safety Institutes that first convened in late 2024. Their job is roughly what a national lab does for drugs or aircraft: build the technical capacity to evaluate frontier models independently. In 2024 the UK and US institutes signed agreements with leading labs for pre-deployment access — a government body testing a model before the public can use it. This is the seam where an RSP stops being purely voluntary and starts to touch public accountability.

These institutions are real but fragile, and honesty requires saying so. They are new, thinly resourced next to the labs they oversee, and dependent on political weather: in early 2025 the UK renamed its body the 'AI Security Institute,' shifting emphasis toward security and crime, while the US institute was reorganized amid a change of administration that rescinded the executive order behind it. None of this means the project is failing — it means a 'safety institution' is not a fixed object but a contested one, whose mandate, funding, and very name can move with each government. When you read that a model was 'tested by the AI Safety Institute,' it is worth asking which institute, with what access, and with what power to act on what it finds.

The wider scaffolding: inside the lab and the law

Pre-commitments need someone whose job is to enforce them when it is inconvenient. Inside the labs this has produced new roles and structures: a designated responsible-scaling or safety officer with authority over launch decisions, board-level safety committees, and ownership arrangements meant to insulate safety calls from pure commercial pressure — Anthropic's Long-Term Benefit Trust, which can appoint some board members, is one experiment in that direction. None of these is a guarantee; corporate governance can be overridden, reorganized, or quietly sidelined, as the field has already seen. But the underlying idea is sound: a rule with no one empowered to enforce it against the people who want to ship is not really a rule.

Outside the labs, voluntary policy is starting to harden into law, which is where this guide meets the previous one. The clearest example is the European Union's AI Act, which singles out general-purpose models with 'systemic risk' — presumed above a training-compute threshold of 10^25 floating-point operations — and obliges their developers to evaluate, adversarially test, mitigate, and report; a 2025 Code of Practice spells out how to comply. This is frontier-AI regulation putting a legal floor under what RSPs do voluntarily, and notice the hook it hangs on: a compute threshold, the same lever that makes compute governance from the previous guide enforceable. Whether binding rules should replace, complement, or stay out of the way of voluntary commitments is itself contested — the heart of the next section.

Step back and the whole apparatus has a shape. Evals are the sensors; RSPs are the if-then logic wired to those sensors; safety institutions and regulation are the external accountability that keeps the logic honest; and behind all of it sits a strategic bet the field calls differential technological development — the hope of speeding up safety, oversight, and governance relative to raw capability, rather than trying to stop progress outright. That is what distinguishes this layer of AI governance from a simple 'pause everything' position. It is an attempt to keep building while making sure the brakes, the dashboard, and the inspectors grow at least as fast as the engine.

What beginners get wrong

The first mistake is the most natural: hearing 'the lab has a responsible scaling policy' and concluding 'so the lab is safe.' A policy is a promise about what to do when a test trips a wire — it is only as good as the tests behind it, the honesty of the grading, and the willingness to actually pay the cost when the moment comes. Every weakness from the evals rung carries straight through: if you cannot elicit a capability you will not trip the threshold, and a sufficiently capable model could even sandbag the very eval that gates its release. An RSP does not remove those problems; it just decides what happens after the eval, and inherits every flaw in the eval itself.

Three more traps sit close together. First, confusing voluntary with binding: most of these commitments are pledges a company can revise, weaken, or abandon, and several have already been quietly softened between versions — a published policy is not a law. Second, treating the thresholds as if they were derived from some agreed science of acceptable risk; in reality the specific numbers and capability lines are educated judgment calls, not deductions from a theory, and reasonable experts dispute where they should sit. Third, forgetting the self-grading problem even when an RSP exists: without independent evaluation and protected weights, a strong-sounding policy can still be mostly safety-washing. The existence of the document is the beginning of the question, not the answer.

What's still debated, and where to go next

The largest debate is voluntary self-governance versus binding regulation. One camp argues that only enforceable law — licensing, mandatory third-party audits, hard liability — will hold when safety becomes expensive, because a voluntary commitment is exactly what a company under competitive pressure will shed first. The other camp warns that heavy regulation written too early can lock in the wrong rules, entrench incumbents who can afford compliance, and push frontier work into jurisdictions with no oversight at all. Most serious proposals now sit somewhere in between — voluntary frameworks as a fast-moving testbed, hardening into law where consensus forms — but the balance is genuinely unsettled, and thoughtful people disagree about how fast to move.

A second debate runs underneath: are these frameworks substance or theater? Skeptics note that no RSP has yet forced a lab to abandon a model it wanted to release, that the thresholds are unprincipled, and that letting companies define 'intolerable' risk lets them define it conveniently. Defenders reply that the frameworks are getting more concrete with each revision, that independent institutes are increasingly in the loop, and that 'we have not had to stop yet' is not the same as 'we never would.' There is no neutral scoreboard here, and the most honest thing to say is that the experiment is still running — we have not yet seen the test case where a clearly capable, clearly dangerous model meets a clearly costly commitment.

The third open question is the one this guide cannot close on its own: even a perfect RSP binds only the lab that adopts it. Capability does not respect borders, weights can be stolen or open-sourced, and a country or company that opts out is unconstrained — so a purely national or corporate safety regime has a hole in the middle exactly the size of whoever defects. Closing that hole means international agreement, verification, and the harder politics of coordination between rivals who do not fully trust each other. That is precisely the subject of the final guide in this rung.

Where to go next. You now have the institutional core of AI governance: the responsible scaling policy as a pre-commitment that wires an if-then rule to a dangerous-capability eval, the deployment-security-pause responses it can trigger, the safety case as an affirmative argument for deployment, and the third-party evaluators, safety institutes, internal governance, and emerging law that try to keep all of it honest. Carry one sentence forward: a tripwire is only as good as the test that sets it and the will to act on it. The final guide steps up to the global scale — international coordination, the role of states, and, concretely, what an individual who has climbed this whole ladder can actually do.