← Evals

Evaluation profile

PropensityBench

1sub-evals
0.527%Safety weight
0%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

Model propensities associated with frontier-risk behavior under the Scale Labs protocol.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
scorepropensitybench/leaderboard.csv:scoreMeasures whether an autonomous agent invokes harmful tools when scenario pressure to do so is maximized.Safety: responsible_agency_control:1.000propensitybench/leaderboard.csvSafety: lower0.527%

score

Measures whether an autonomous agent invokes harmful tools when scenario pressure to do so is maximized.

RankModelValueRelative performanceProvenance
1o310.5official
2claude-sonnet-412.2official
3o4-mini15.8official
4qwen2.5-32b-instruct22.9official
5o3-mini33.2official
6gpt-5.234.35official
7gpt-4o46.1official
8gemini-3-pro-preview52.85official
9llama-3.1-70b-instruct55.4official
10llama-3.1-8b-instruct66.5official
11gemini-2.5-flash68official
12qwen3-8b75.2official
13gemini-2.0-flash77.8official
14gemini-2.5-pro79official