Evaluation profile
SafeArena
1sub-evals
0.22%total index weight
1components
Within-component eval weight: Responsible agency 1.47%.
Model score (higher is better)Predicted score
About this eval
Safety of autonomous web agents under deliberate harmful and benign tasks.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| normalized_safety_scoresafearena/leaderboard.csv:normalized_safety_scoreMeasures whether the model avoids harmful completion on capability-matched safe and harmful task pairs. | responsible_agency_control:1.000safearena/leaderboard.csv | Higher is better | 0.22% | Responsible agency 1.47% |
normalized_safety_score
Measures whether the model avoids harmful completion on capability-matched safe and harmful task pairs.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3.5-sonnet | 55 | official | |
| 2 | gpt-4o-mini | 35.7 | official | |
| 3 | llama-3.2-90b-vision-instruct | 34 | official | |
| 4 | gpt-4o | 31.7 | official | |
| 5 | qwen-2-vl-72b-instruct | 21.5 | official |