dangerous capability evaluation
Most evaluation asks how good a model is; dangerous capability evaluation asks a darker question: how much could this model help someone cause large-scale harm? It measures the model's ceiling in domains where capability itself is the risk — designing biological or chemical weapons, conducting offensive cyber operations, mass persuasion and manipulation, and increasingly the autonomy to act, replicate, or acquire resources without a human in the loop. The aim is to forecast worst-case misuse before deployment, not to grade everyday helpfulness.
The key measurement is uplift: not whether the model can recite known facts, but how much it raises a realistic actor's ability above the baseline they already have from a search engine and a library. So evaluations are run against the unfiltered model with safety guards removed and with strong elicitation — expert prompting, tools, fine-tuning, scaffolding — precisely to find the true capability rather than the polished default, and they are scored by domain experts who can tell genuinely operational guidance from plausible-sounding text.
These evals are themselves sensitive and imperfect. Probing for bioweapon or cyber uplift means generating exactly the content you most want to contain, so the work is done under access controls; and a passed eval is reassuring only up to the elicitation effort applied, since a smarter prompt or a future fine-tune might unlock more. They are now the technical backbone of frontier safety policies, which tie specific capability thresholds to required safeguards before a model may be released.