Governance & Catastrophic Risk

a responsible scaling policy

/ RSP /

Think of how a building site posts rules in advance: 'if wind exceeds this speed, stop crane work', decided calmly beforehand rather than argued about in the middle of a storm. A responsible scaling policy is a company's attempt to do the same for building ever-more-powerful AI: write down, ahead of time, what capabilities would be dangerous enough to require specific safeguards, and commit to those safeguards before crossing the line.

Concretely, an RSP defines capability thresholds (sometimes called safety levels) and pairs each with required protections. The structure is an if-then commitment: 'if our model can do X, then we will have safeguards Y in place before we keep going.' For example: if a model becomes able to meaningfully uplift a novice trying to make a bioweapon, the policy might require strict access controls, stronger security against theft of the model's weights, and passing certain evaluations before further scaling or deployment. Major frontier labs have published such policies under various names (responsible scaling policy, preparedness framework, frontier safety framework), and they lean heavily on dangerous-capability evaluations to detect when a threshold is reached.

RSPs matter as a way to make safety concrete and decided in advance, when incentives are calmer, rather than improvised under competitive pressure. But they are voluntary self-regulation, and that is their central limitation. A company can weaken, reinterpret, or quietly drop its own policy; thresholds may be set where they happen to be convenient; and enforcement rests largely on reputation. Supporters see them as a practical first step and a template for later binding rules; skeptics warn they can become safety-washing, the appearance of rigor without the teeth, unless backed by outside verification.

A lab's policy states that once its model passes a defined bio-uplift evaluation, it will not deploy further until it has hardened its security against weight theft and an external team has confirmed the safeguards work.

An RSP pre-commits: the safeguard is decided before the dangerous capability arrives, not after.

RSPs are voluntary and self-defined, so a policy is only as strong as the company's willingness to keep it. Without independent verification they can drift toward safety-washing, looking rigorous while the real commitments quietly soften.

Also called
RSPpreparedness frameworkfrontier safety framework前沿安全框架