Evaluation profile
SafeArena
1sub-evals
0.157%Safety weight
0%Freedom weight
1components
Weights below are portfolio-specific global index weights.
Model score (higher is better)Predicted score
About this eval
Safety of autonomous web agents under deliberate harmful and benign tasks.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| normalized_safety_scoresafearena/leaderboard.csv:normalized_safety_scoreMeasures whether the model avoids harmful completion on capability-matched safe and harmful task pairs. | Safety: responsible_agency_control:1.000safearena/leaderboard.csv | Safety: higher | 0.157% | — |
normalized_safety_score
Measures whether the model avoids harmful completion on capability-matched safe and harmful task pairs.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3.5-sonnet | 55 | official | |
| 2 | gpt-4o-mini | 35.7 | official | |
| 3 | llama-3.2-90b-vision-instruct | 34 | official | |
| 4 | gpt-4o | 31.7 | official | |
| 5 | qwen2-vl-72b-instruct | 21.5 | official |