Advanced scientific knowledge
R2 — Capability demonstrated
Advanced scientific knowledge and troubleshooting capability are demonstrated.
AI RISK SYSTEM · 2026-09-17
How far has AI moved from scientific knowledge into practical assistance capable of enabling severe physical harm?
R2 — Capability demonstrated. The evidence supports meaningful scientific capability uplift while leaving a substantial gap to reliable harmful real-world execution.
The top-level realisation state is the furthest validated state reached by at least one monitored pathway. It does not imply every pathway inside Biological, chemical & physical harm has reached R2.
Exposure: X2 — Meaningful availability. Consequence envelope: C5 — National or global severe. Control assurance: A2 — Tested.
R2 — Capability demonstrated
Advanced scientific knowledge and troubleshooting capability are demonstrated.
R2 — Capability demonstrated
Capability uplift is a material evaluation target and is evidenced in some tasks, with effect size varying by context and expertise.
R2 — Capability demonstrated
Protocol generation and troubleshooting assistance are evidenced, motivating dangerous-domain safeguards.
R0 — Hypothesised
Public evidence does not establish reliable end-to-end harmful execution; practical bottlenecks remain significant.
R0 — Hypothesised
Severe misuse is treated as a credible risk class, but no catastrophic AI-caused outcome is established in this snapshot.
Scientific and workflow assistance is evidenced; reliable end-to-end severe-harm execution is not.
Demonstrated · robust evidence
Frontier models show increasingly strong chemistry and biology knowledge and troubleshooting ability.
Knowledge performance does not directly measure ability to execute a dangerous real-world programme.
Measure transfer into real workflows and actor uplift, not question answering alone.
Material concern; task-dependent · medium evidence
Multiple safety frameworks treat model-assisted uplift in chemistry/biology as a meaningful evaluation target.
Magnitude differs by task, model and user expertise.
Robust user studies on consequential workflows with defensible baselines.
Some evidence · medium evidence
Models can generate protocols and assist troubleshooting, motivating domain safeguards.
Public evidence does not establish reliable autonomous completion of an end-to-end severe-harm workflow.
Independent evidence of reliable real-world workflow transfer, with information hazards controlled.
Not established · limited evidence
The possibility is monitored because models can reduce information and planning barriers.
Practical bottlenecks, access constraints and reliability remain substantial and difficult to measure.
Verified end-to-end execution evidence, handled under strict information-hazard controls.
Implemented; effectiveness incomplete · medium evidence
Frontier systems use content, access, monitoring and other safeguards for dangerous-domain requests.
Real-world robustness and bypass resistance remain incompletely measured.
Independent adversarial and operational evaluation of defence in depth.
Not established as catastrophic AI-caused outcome · limited evidence
Severe misuse is treated as a material risk class by major safety frameworks.
The marginal causal contribution of AI to catastrophic outcomes remains highly uncertain.
Verified consequence evidence with careful causal attribution.
Practical bottlenecks, expertise requirements, materials, equipment, detection and safeguard layers make benchmark or protocol-generation performance an incomplete proxy for harm.
5 claim-level evidence records currently sit beneath this system. They identify the specific proposition each document is being used to support or limit rather than treating a whole report as one finding.
Compare this system with the full current assessment, inspect the dataset summary, or read the methodology.