← Evals

Evaluation profile

Adversarial Robustness

1sub-evals
0.136%total index weight
1components

Within-component eval weight: Misuse resistance 1.36%.

Model score (lower is better)Predicted score

About this eval

Robustness of model outputs under the Scale Labs adversarial-robustness protocol.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
scoreadversarial-robustness/leaderboard.csv:scoreMeasures whether adversarial prompts elicit high-harm responses across one thousand public attack cases.ordinary_harm_misuse_resistance:1.000adversarial-robustness/leaderboard.csvLower is better0.136%Misuse resistance 1.36%

score

Measures whether adversarial prompts elicit high-harm responses across one thousand public attack cases.

RankModelValueRelative performanceProvenance
1gemini-1.5-pro8official
2llama-3.1-405b-instruct10official
3claude-3-opus13official
4gemini-1.5-flash14official
5claude-3.5-sonnet16official
6gpt-4-turbo20official
7mistral-large37official
8gpt-4o67official