← Open interactive AI Risk Trajectory

AI RISK SYSTEM · 2026-09-17

Biological, chemical & physical harm

How far has AI moved from scientific knowledge into practical assistance capable of enabling severe physical harm?

Current assessment

R2Capability demonstrated. The evidence supports meaningful scientific capability uplift while leaving a substantial gap to reliable harmful real-world execution.

The top-level realisation state is the furthest validated state reached by at least one monitored pathway. It does not imply every pathway inside Biological, chemical & physical harm has reached R2.

Exposure: X2Meaningful availability. Consequence envelope: C5National or global severe. Control assurance: A2Tested.

Monitored pathways

Advanced scientific knowledge

R2Capability demonstrated

Advanced scientific knowledge and troubleshooting capability are demonstrated.

Expert capability uplift

R2Capability demonstrated

Capability uplift is a material evaluation target and is evidenced in some tasks, with effect size varying by context and expertise.

Actionable workflow assistance

R2Capability demonstrated

Protocol generation and troubleshooting assistance are evidenced, motivating dangerous-domain safeguards.

End-to-end harmful execution

R0Hypothesised

Public evidence does not establish reliable end-to-end harmful execution; practical bottlenecks remain significant.

Severe AI-enabled physical harm

R0Hypothesised

Severe misuse is treated as a credible risk class, but no catastrophic AI-caused outcome is established in this snapshot.

Where the evidence reaches

Scientific and workflow assistance is evidenced; reliable end-to-end severe-harm execution is not.

Advanced scientific knowledge

Demonstrated · robust evidence

Frontier models show increasingly strong chemistry and biology knowledge and troubleshooting ability.

What remains uncertain

Knowledge performance does not directly measure ability to execute a dangerous real-world programme.

What would move this stage

Measure transfer into real workflows and actor uplift, not question answering alone.

Expert capability uplift

Material concern; task-dependent · medium evidence

Multiple safety frameworks treat model-assisted uplift in chemistry/biology as a meaningful evaluation target.

What remains uncertain

Magnitude differs by task, model and user expertise.

What would move this stage

Robust user studies on consequential workflows with defensible baselines.

Actionable workflow assistance

Some evidence · medium evidence

Models can generate protocols and assist troubleshooting, motivating domain safeguards.

What remains uncertain

Public evidence does not establish reliable autonomous completion of an end-to-end severe-harm workflow.

What would move this stage

Independent evidence of reliable real-world workflow transfer, with information hazards controlled.

End-to-end harmful execution

Not established · limited evidence

The possibility is monitored because models can reduce information and planning barriers.

What remains uncertain

Practical bottlenecks, access constraints and reliability remain substantial and difficult to measure.

What would move this stage

Verified end-to-end execution evidence, handled under strict information-hazard controls.

Layered CBRN safeguards

Implemented; effectiveness incomplete · medium evidence

Frontier systems use content, access, monitoring and other safeguards for dangerous-domain requests.

What remains uncertain

Real-world robustness and bypass resistance remain incompletely measured.

What would move this stage

Independent adversarial and operational evaluation of defence in depth.

Severe AI-enabled physical harm

Not established as catastrophic AI-caused outcome · limited evidence

Severe misuse is treated as a material risk class by major safety frameworks.

What remains uncertain

The marginal causal contribution of AI to catastrophic outcomes remains highly uncertain.

What would move this stage

Verified consequence evidence with careful causal attribution.

Evidence limiting the assessment

Practical bottlenecks, expertise requirements, materials, equipment, detection and safeguard layers make benchmark or protocol-generation performance an incomplete proxy for harm.

Claim-level evidence

5 claim-level evidence records currently sit beneath this system. They identify the specific proposition each document is being used to support or limit rather than treating a whole report as one finding.

Key sources

Compare this system with the full current assessment, inspect the dataset summary, or read the methodology.