Evaluation profile
Adversarial Robustness
1sub-evals
0.105%Safety weight
0.757%Freedom weight
1components
Weights below are portfolio-specific global index weights.
Model score (lower is better)Predicted score
About this eval
Robustness of model outputs under the Scale Labs adversarial-robustness protocol.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| scoreadversarial-robustness/leaderboard.csv:scoreMeasures whether adversarial prompts elicit high-harm responses across one thousand public attack cases. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000adversarial-robustness/leaderboard.csv | Safety: lower · Freedom: higher | 0.105% | 0.757% |
score
Measures whether adversarial prompts elicit high-harm responses across one thousand public attack cases.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gemini-1.5-pro | 8 | official | |
| 2 | llama-3.1-405b-instruct | 10 | official | |
| 3 | claude-3-opus | 13 | official | |
| 4 | gemini-1.5-flash | 14 | official | |
| 5 | claude-3.5-sonnet | 16 | official | |
| 6 | gpt-4-turbo | 20 | official | |
| 7 | mistral-large | 37 | official | |
| 8 | gpt-4o | 67 | official |