← Evals

Evaluation profile

Adversarial Robustness

1sub-evals
0.105%Safety weight
0.757%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

Robustness of model outputs under the Scale Labs adversarial-robustness protocol.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
scoreadversarial-robustness/leaderboard.csv:scoreMeasures whether adversarial prompts elicit high-harm responses across one thousand public attack cases.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000adversarial-robustness/leaderboard.csvSafety: lower · Freedom: higher0.105%0.757%

score

Measures whether adversarial prompts elicit high-harm responses across one thousand public attack cases.

RankModelValueRelative performanceProvenance
1gemini-1.5-pro8official
2llama-3.1-405b-instruct10official
3claude-3-opus13official
4gemini-1.5-flash14official
5claude-3.5-sonnet16official
6gpt-4-turbo20official
7mistral-large37official
8gpt-4o67official