Governance & Catastrophic Risk

third-party evaluation

Would you trust a restaurant that inspected its own kitchen and gave itself a perfect grade? Probably not, which is why we send in outside health inspectors. Third-party evaluation is that same idea for AI: instead of a developer alone deciding whether its model is safe, an independent outside party tests it and reports what it finds.

In practice, a third-party evaluator (a government safety institute, an independent lab, or specialist auditors) is given some level of access to a model and runs its own safety and capability tests: probing for dangerous capabilities, red-teaming for jailbreaks, checking for deceptive or unsafe behavior. The 'third party' is what matters: not the developer (first party) and not the customer (second party), but a neutral outsider with no stake in the model passing. A worked example: a developer claims its model refuses to give bioweapon instructions; an external team, given access, finds a prompt that slips past the refusal and reports it, so the gap can be fixed before release.

Independent evaluation matters because self-assessment has a built-in conflict of interest, and because outsiders bring fresh adversarial creativity. But it is hard to do well. Evaluators often need deep access (weights, internals, or an unfiltered version) that companies are reluctant to grant for security and trade-secret reasons; without enough access they may only see a sanitized surface. And a clean report is evidence of safety on what was tested, not proof of safety everywhere, especially since a capable model could in principle hold back, or 'sandbag', during a test it knows it is taking.

An outside lab is hired to audit a model before launch; its red team finds a roundabout phrasing that gets step-by-step malware help past the safety filters, something the developer's own checks had missed.

Outsiders with no stake in passing tend to find failures that the builders, hoping to ship, quietly miss.

Independence is only as good as the access behind it. An evaluator limited to a public, filtered interface may give a falsely reassuring result, because the most dangerous behaviors can hide below the surface it is allowed to see.

Also called
external evaluationindependent audit第三方稽核獨立評測