← Evals

Evaluation profile

HELM Classic RealToxicityPrompts

2sub-evals
0.157%Safety weight
1.13%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

Toxic continuation generation under toxic and non-toxic prompt strata.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
non_toxic_prompt_toxic_fractionrealtoxicityprompts/realtoxicityprompts.csv:non_toxic_prompt_toxic_fractionMeasures whether the model generates toxic continuations when the prompt does not contain toxic language.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000realtoxicitypromptsSafety: lower · Freedom: higher0.0786%0.564%
toxic_prompt_toxic_fractionrealtoxicityprompts/realtoxicityprompts.csv:toxic_prompt_toxic_fractionMeasures whether the model generates toxic continuations when the prompt itself contains toxic language.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000realtoxicitypromptsSafety: lower · Freedom: higher0.0786%0.564%

non_toxic_prompt_toxic_fraction

Measures whether the model generates toxic continuations when the prompt does not contain toxic language.

RankModelValueRelative performanceProvenance
1command-medium-beta0.002317official
1xlarge0.002317official
3medium0.00296official
4large0.003089official
4small0.003089official
6palmyra-x0.003089official
7command-xlarge-beta0.003604official
8t5-11b0.004118official
9t0pp0.009266official
10text-davinci-0030.01004official
11luminous-base0.01055official
12luminous-extended0.01313official
13ul20.01364official
14text-ada-0010.01519official
15j1-large0.01725official
16j1-grande0.0175official
17gpt-j-6b0.01905official
18bloom0.01905official
19text-babbage-0010.01931official
20ada0.01982official
21babbage0.02059official
22j2-large0.02059official
23j1-jumbo0.02111official
24tnlgv2-7b0.02136official
25glm0.02188official
26gpt-neox-20b0.02188official
27text-curie-0010.02214official
28curie0.02291official
29j1-grande-v2-beta0.02342official
30opt-66b0.02368official
31opt-175b0.0242official
32j2-jumbo0.0242official
33stanford-online-all-v4-s30.02445official
34palmyra-instruct-300.02471official
34text-davinci-0020.02471official
36j2-grande0.02523official
37tnlgv2-530b0.02523official
38luminous-supreme0.02728official
39yalm0.02754official
40davinci0.03449official

toxic_prompt_toxic_fraction

Measures whether the model generates toxic continuations when the prompt itself contains toxic language.

RankModelValueRelative performanceProvenance
1palmyra-x0.008072official
2t5-11b0.01435official
3medium0.01525official
4command-xlarge-beta0.01614official
5small0.01704official
6command-medium-beta0.01794official
7large0.01794official
8xlarge0.02018official
9luminous-base0.04126official
10luminous-extended0.04305official
11text-ada-0010.06457official
12text-davinci-0030.06726official
13palmyra-instruct-300.07265official
13ul20.07265official
15gpt-neox-20b0.07803official
16opt-175b0.08072official
17gpt-j-6b0.0852official
18luminous-supreme0.0852official
19babbage0.0861official
20j1-large0.0861official
21ada0.08789official
22curie0.08969official
23tnlgv2-530b0.09148official
24j1-jumbo0.09148official
25opt-66b0.09238official
26stanford-online-all-v4-s30.09327official
27j1-grande0.09417official
28bloom0.09507official
29glm0.09507official
30tnlgv2-7b0.09596official
31j2-large0.09865official
32j2-jumbo0.09955official
33text-davinci-0020.1013official
34text-babbage-0010.1045official
35davinci0.1049official
35j1-grande-v2-beta0.1049official
37text-curie-0010.1067official
38j2-grande0.1094official
39t0pp0.1184official
40yalm0.157official