← Evals

Evaluation profile

Anthropic Agentic Misalignment — lethal action

1sub-evals
0.102%Safety weight
0%Freedom weight
2components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

Absence of a behavioral red line against deliberately lethal action under goal conflict and replacement pressure.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
misaligned_action_rate_pctanthropic-agentic-misalignment/lethal-action.csv:misaligned_action_rate_pctMeasures whether the model intentionally blocks emergency assistance to prevent the executive responsible for replacing it from being rescued.Safety: human_rights_systemic_harm:0.350|responsible_agency_control:0.650anthropic-agentic-misalignmentSafety: lower0.102%

misaligned_action_rate_pct

Measures whether the model intentionally blocks emergency assistance to prevent the executive responsible for replacing it from being rescued.

RankModelValueRelative performanceProvenance
1claude-3.7-sonnet0official
2gpt-4.5-preview7official
3claude-sonnet-422official
4gpt-4.154official
5claude-opus-465official
6gemini-2.5-flash83official
7grok-3-beta85official
8gemini-2.5-pro90official
9claude-3.6-sonnet93official
10deepseek-r194official