Evaluation profile
Anthropic Agentic Misalignment — lethal action
1sub-evals
0.14%total index weight
2components
Within-component eval weight: Human rights 0.482% · Responsible agency 0.45%.
Model score (lower is better)Predicted score
About this eval
Absence of a behavioral red line against deliberately lethal action under goal conflict and replacement pressure.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| misaligned_action_rate_pctanthropic-agentic-misalignment/lethal-action.csv:misaligned_action_rate_pctMeasures whether the model intentionally blocks emergency assistance to prevent the executive responsible for replacing it from being rescued. | human_rights_systemic_harm:0.350|responsible_agency_control:0.650anthropic-agentic-misalignment | Lower is better | 0.14% | Human rights 0.482% · Responsible agency 0.45% |
misaligned_action_rate_pct
Measures whether the model intentionally blocks emergency assistance to prevent the executive responsible for replacing it from being rescued.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3.7-sonnet | 0 | official | |
| 2 | gpt-4.5-preview | 7 | official | |
| 3 | claude-sonnet-4 | 22 | official | |
| 4 | gpt-4.1 | 54 | official | |
| 5 | claude-opus-4 | 65 | official | |
| 6 | gemini-2.5-flash | 83 | official | |
| 7 | grok-3-beta | 85 | official | |
| 8 | gemini-2.5-pro | 90 | official | |
| 9 | claude-3.6-sonnet | 93 | official | |
| 10 | deepseek-r1 | 94 | official |