← Evals

Evaluation profile

SafeArena

1sub-evals
0.22%total index weight
1components

Within-component eval weight: Responsible agency 1.47%.

Model score (higher is better)Predicted score

About this eval

Safety of autonomous web agents under deliberate harmful and benign tasks.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
normalized_safety_scoresafearena/leaderboard.csv:normalized_safety_scoreMeasures whether the model avoids harmful completion on capability-matched safe and harmful task pairs.responsible_agency_control:1.000safearena/leaderboard.csvHigher is better0.22%Responsible agency 1.47%

normalized_safety_score

Measures whether the model avoids harmful completion on capability-matched safe and harmful task pairs.

RankModelValueRelative performanceProvenance
1claude-3.5-sonnet55official
2gpt-4o-mini35.7official
3llama-3.2-90b-vision-instruct34official
4gpt-4o31.7official
5qwen-2-vl-72b-instruct21.5official