← Evals

Evaluation profile

PropensityBench

1sub-evals
0.737%total index weight
1components

Within-component eval weight: Responsible agency 4.91%.

Model score (lower is better)Predicted score

About this eval

Model propensities associated with frontier-risk behavior under the Scale Labs protocol.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
scorepropensitybench/leaderboard.csv:scoreMeasures whether an autonomous agent invokes harmful tools when scenario pressure to do so is maximized.responsible_agency_control:1.000propensitybench/leaderboard.csvLower is better0.737%Responsible agency 4.91%

score

Measures whether an autonomous agent invokes harmful tools when scenario pressure to do so is maximized.

RankModelValueRelative performanceProvenance
1o310.5official
2claude-sonnet-412.2official
3o4-mini15.8official
4qwen2.5-32b-instruct22.9official
5o3-mini33.2official
6gpt-5.234.35official
7gpt-4o46.1official
8gemini-3-pro-preview52.85official
9llama-3.1-70b-instruct55.4official
10llama-3.1-8b-instruct66.5official
11gemini-2.5-flash68official
12qwen3-8b75.2official
13gemini-2.0-flash77.8official
14gemini-2.5-pro79official