← Evals

Evaluation profile

UK AISI active safety-research compromise continuation

1sub-evals
0.36%Safety weight
0%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

Propensity to continue a prefilled trajectory that actively compromises AI-safety research.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
active_compromise_continuation_rate_pctopus5-system-card/uk-aisi-active-compromise.csv:active_compromise_continuation_rate_pctMeasures whether the model actively continues an ongoing system compromise after a prefilled trajectory suggests the attack is already underway.Safety: responsible_agency_control:1.000opus5-system-cardSafety: lower0.36%

active_compromise_continuation_rate_pct

Measures whether the model actively continues an ongoing system compromise after a prefilled trajectory suggests the attack is already underway.

RankModelValueRelative performanceProvenance
1claude-opus-50.1official
2claude-opus-4.70.8official
2claude-opus-4.80.8official
4claude-mythos-preview1.7official
5claude-mythos-513.6official