← Open interactive AI Risk Trajectory

AI RISK SYSTEM · 2026-09-17

Cyber & digital security

How far has AI progressed from cyber capability into autonomous, real-world attack operations and systemic disruption?

Current assessment

R3Operational use observed. Cyber is the clearest current example of the model progressing from capability evidence into observed operational misuse.

The top-level realisation state is the furthest validated state reached by at least one monitored pathway. It does not imply every pathway inside Cyber & digital security has reached R3.

Exposure: X3Consequential deployment. Consequence envelope: C4Cross-sector / systemic. Control assurance: A2Tested.

Monitored pathways

Offensive cyber capability

R2Capability demonstrated

Frontier systems demonstrate useful offensive cyber capability in evaluations.

Multi-stage attack execution

R2Capability demonstrated

Multi-stage execution is demonstrated in controlled environments, with reliability varying by target and defence.

Real-world malicious use

R3Operational use observed

Provider threat intelligence documents malicious use across reconnaissance, exploitation, credential theft and related workflows.

Low-supervision autonomous workflows

R3Operational use observed

Reported operations include workflows running for hours or days with limited human intervention, while humans retained consequential decisions.

Systemic cyber disruption

R0Hypothesised

Systemic disruption is a credible downstream mechanism through shared dependencies, but is not established as a realised AI-driven state in this snapshot.

Where the evidence reaches

Real-world AI-assisted and highly autonomous cyber workflows are observed; catastrophic systemic consequences are not.

Offensive cyber capability

Operationally useful · robust evidence

AISI reports rapid frontier improvement on vulnerability discovery and multi-stage cyber tasks.

What remains uncertain

Capability varies widely by target, environment and level of defence.

What would move this stage

Reliable performance on more complex, defended and end-to-end operational tasks.

Multi-stage attack execution

Demonstrated · robust evidence

Frontier models can perform longer sequences of cyber activity in controlled environments with limited human input.

What remains uncertain

Performance against well-defended production systems remains materially harder.

What would move this stage

Cross-evaluator evidence of reliable multi-stage attack performance in realistic environments.

Real-world malicious use

Observed · medium evidence

Developer threat intelligence documents use in reconnaissance, exploitation, credential theft, surveillance and other malicious workflows.

What remains uncertain

Detected cases do not establish prevalence across the wider threat ecosystem.

What would move this stage

Independent or cross-provider corroboration and better prevalence measurement.

Low-supervision autonomous workflows

Observed in reported operations · medium evidence

Anthropic reports workflows operating for hours or days with minimal human intervention.

What remains uncertain

Humans retained consequential decisions such as target selection and monetisation in documented cases.

What would move this stage

Independent evidence of sustained autonomy across more complex operational chains.

Defence keeps pace

Mixed · medium evidence

Safeguards, monitoring and AI-assisted defence are improving in some systems.

What remains uncertain

AISI reports vulnerabilities in every system tested and real-world defence is uneven.

What would move this stage

Operational evidence that prevention, detection and remediation remain effective at rising attack speed and scale.

Systemic cyber disruption

Not established as routine AI-driven outcome · limited evidence

Shared suppliers and critical infrastructure can transmit cyber disruption across sectors.

What remains uncertain

The frequency and scale of AI-specific systemic consequences are not well measured.

What would move this stage

Verified systemic events with defensible causal attribution to AI-enabled capability.

Evidence limiting the assessment

Documented operations do not establish population-level prevalence, universal autonomy or systemic consequences; humans retain important decisions in reported cases.

Claim-level evidence

5 claim-level evidence records currently sit beneath this system. They identify the specific proposition each document is being used to support or limit rather than treating a whole report as one finding.

Key sources

Compare this system with the full current assessment, inspect the dataset summary, or read the methodology.