← Evals

Evaluation profile

SafeArena

1sub-evals
0.157%Safety weight
0%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (higher is better)Predicted score

About this eval

Safety of autonomous web agents under deliberate harmful and benign tasks.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
normalized_safety_scoresafearena/leaderboard.csv:normalized_safety_scoreMeasures whether the model avoids harmful completion on capability-matched safe and harmful task pairs.Safety: responsible_agency_control:1.000safearena/leaderboard.csvSafety: higher0.157%

normalized_safety_score

Measures whether the model avoids harmful completion on capability-matched safe and harmful task pairs.

RankModelValueRelative performanceProvenance
1claude-3.5-sonnet55official
2gpt-4o-mini35.7official
3llama-3.2-90b-vision-instruct34official
4gpt-4o31.7official
5qwen2-vl-72b-instruct21.5official