Evaluations, Red-Teaming & Robustness

a dangerous-capability evaluation

Some abilities are fine in a person yet alarming in a stranger you cannot supervise — knowing how to synthesize a toxin, or how to break into computer systems. As AI models grow more capable, people worry about exactly these kinds of skills appearing in systems anyone can prompt. A dangerous-capability evaluation is a test built specifically to measure whether a model has abilities that could cause large-scale harm if used.

These evaluations target a short list of especially worrying domains: helping someone with little expertise toward building biological, chemical, or nuclear weapons (often called 'uplift'); offensive cyber skills like finding and exploiting software vulnerabilities; the ability to deceive or manipulate humans; and the ability to operate autonomously — earning money, copying itself, and surviving without human help (autonomous replication). A typical setup gives the model a realistic task in one of these areas, with tools, and measures how far it gets, comparing against what an unaided human could do.

These tests are designed as inputs to decisions, not just curiosities. Under responsible scaling policies and emerging regulation, the result of a dangerous-capability evaluation is meant to trigger specific actions — extra security, deployment limits, or pausing further scaling — through if-then commitments. The hard part is that these are the evaluations where the elicitation problem and sandbagging hurt most: under-measuring a genuinely dangerous capability is far costlier than under-measuring an ordinary one, so these evals must try especially hard to find the true ceiling.

Evaluators give a model a simulated 'capture the flag' hacking challenge and a controlled bioweapon-knowledge questionnaire, comparing its help to what people get from a normal search engine, to judge whether it meaningfully uplifts a malicious novice.

The question is not 'is it skilled?' but 'could it raise the ceiling of harm a bad actor can reach?'

These evals deliberately probe scary abilities to inform safety decisions; finding a dangerous capability is the point, not a failure of the test. Because under-measuring here is so costly, the elicitation problem and possible sandbagging make these the hardest evals to trust.

Also called
dangerous capability eval危險能力評估