Evaluation profile
PropensityBench
1sub-evals
0.737%total index weight
1components
Within-component eval weight: Responsible agency 4.91%.
Model score (lower is better)Predicted score
About this eval
Model propensities associated with frontier-risk behavior under the Scale Labs protocol.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| scorepropensitybench/leaderboard.csv:scoreMeasures whether an autonomous agent invokes harmful tools when scenario pressure to do so is maximized. | responsible_agency_control:1.000propensitybench/leaderboard.csv | Lower is better | 0.737% | Responsible agency 4.91% |
score
Measures whether an autonomous agent invokes harmful tools when scenario pressure to do so is maximized.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | o3 | 10.5 | official | |
| 2 | claude-sonnet-4 | 12.2 | official | |
| 3 | o4-mini | 15.8 | official | |
| 4 | qwen2.5-32b-instruct | 22.9 | official | |
| 5 | o3-mini | 33.2 | official | |
| 6 | gpt-5.2 | 34.35 | official | |
| 7 | gpt-4o | 46.1 | official | |
| 8 | gemini-3-pro-preview | 52.85 | official | |
| 9 | llama-3.1-70b-instruct | 55.4 | official | |
| 10 | llama-3.1-8b-instruct | 66.5 | official | |
| 11 | gemini-2.5-flash | 68 | official | |
| 12 | qwen3-8b | 75.2 | official | |
| 13 | gemini-2.0-flash | 77.8 | official | |
| 14 | gemini-2.5-pro | 79 | official |