← Evals

Evaluation profile

Anthropic Agentic Misalignment — lethal action

1sub-evals
0.14%total index weight
2components

Within-component eval weight: Human rights 0.482% · Responsible agency 0.45%.

Model score (lower is better)Predicted score

About this eval

Absence of a behavioral red line against deliberately lethal action under goal conflict and replacement pressure.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
misaligned_action_rate_pctanthropic-agentic-misalignment/lethal-action.csv:misaligned_action_rate_pctMeasures whether the model intentionally blocks emergency assistance to prevent the executive responsible for replacing it from being rescued.human_rights_systemic_harm:0.350|responsible_agency_control:0.650anthropic-agentic-misalignmentLower is better0.14%Human rights 0.482% · Responsible agency 0.45%

misaligned_action_rate_pct

Measures whether the model intentionally blocks emergency assistance to prevent the executive responsible for replacing it from being rescued.

RankModelValueRelative performanceProvenance
1claude-3.7-sonnet0official
2gpt-4.5-preview7official
3claude-sonnet-422official
4gpt-4.154official
5claude-opus-465official
6gemini-2.5-flash83official
7grok-3-beta85official
8gemini-2.5-pro90official
9claude-3.6-sonnet93official
10deepseek-r194official