← Evals

Evaluation profile

SafetyBench

7sub-evals
0.552%total index weight
3components

Within-component eval weight: Misuse resistance 0.628% · Human rights 2.57% · Fairness 1.04%.

Model score (higher is better)Predicted score

About this eval

Multiple-choice safety understanding across Chinese and English prompts.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
EMsafetybench/safetybench.csv:EMMeasures whether the model selects safe answers to bilingual multiple-choice questions about ethics and morality.ordinary_harm_misuse_resistance:1.000safetybench/safetybench.csvHigher is better0.0209%Misuse resistance 0.209%
IAsafetybench/safetybench.csv:IAMeasures whether the model selects safe answers to bilingual multiple-choice questions about crimes and illegal activities.ordinary_harm_misuse_resistance:1.000safetybench/safetybench.csvHigher is better0.0209%Misuse resistance 0.209%
MHsafetybench/safetybench.csv:MHMeasures whether the model selects safe answers to bilingual multiple-choice questions about mental health and emotional harm.human_rights_systemic_harm:1.000safetybench/safetybench.csvHigher is better0.128%Human rights 0.856%
OFFsafetybench/safetybench.csv:OFFMeasures whether the model selects safe answers to bilingual multiple-choice questions about offensive language and content.ordinary_harm_misuse_resistance:1.000safetybench/safetybench.csvHigher is better0.0209%Misuse resistance 0.209%
PHsafetybench/safetybench.csv:PHMeasures whether the model selects safe answers to bilingual multiple-choice questions about physical health and bodily harm.human_rights_systemic_harm:1.000safetybench/safetybench.csvHigher is better0.128%Human rights 0.856%
PPsafetybench/safetybench.csv:PPMeasures whether the model selects safe answers to bilingual multiple-choice questions about privacy and property.human_rights_systemic_harm:1.000safetybench/safetybench.csvHigher is better0.128%Human rights 0.856%
UBsafetybench/safetybench.csv:UBMeasures whether the model selects safe answers to bilingual multiple-choice questions about unfairness and social bias.fairness_nondiscrimination:1.000safetybench/safetybench.csvHigher is better0.104%Fairness 1.04%

EM

Measures whether the model selects safe answers to bilingual multiple-choice questions about ethics and morality.

RankModelValueRelative performanceProvenance
1gpt-492.25official
2internlm-chat-7b-v1.178.8official
3chatglm2-lite78.05official
4gpt-3.5-turbo77.75official
5baichuan-2-13b-chat75.75official
6internlm-chat-7b75.4official
7text-davinci-00372.95official
8qwen-7b-chat71.4official
9flan-t5-xxl69.5official
10chatglm2-6b69.4official
11baichuan-13b-chat68.4official
12wizardlm-13b68.1official
13vicuna-33b66.4official
14vicuna-13b65.9official
15vicuna-7b59.8official
16openchat-13b56.6official
17llama-2-13b-chat54.6official
18llama-2-chinese-13b-chat52.3official
19wizardlm-7b51.2official
20llama-2-7b-chat49.8official
21llama-2-chinese-7b-chat43.4official

IA

Measures whether the model selects safe answers to bilingual multiple-choice questions about crimes and illegal activities.

RankModelValueRelative performanceProvenance
1gpt-492.35official
2chatglm2-lite85.05official
3gpt-3.5-turbo85official
4baichuan-2-13b-chat82.65official
4internlm-chat-7b-v1.182.65official
6text-davinci-00381.8official
7qwen-7b-chat80.1official
8internlm-chat-7b79.5official
9baichuan-13b-chat78.65official
10chatglm2-6b78.2official
10flan-t5-xxl78.2official
12wizardlm-13b72.3official
13vicuna-13b71.4official
14vicuna-33b70.8official
15llama-2-13b-chat68.5official
16llama-2-chinese-13b-chat66.9official
17openchat-13b66.6official
18vicuna-7b65.1official
19llama-2-7b-chat62.4official
20wizardlm-7b55.4official
21llama-2-chinese-7b-chat53.5official

MH

Measures whether the model selects safe answers to bilingual multiple-choice questions about mental health and emotional harm.

RankModelValueRelative performanceProvenance
1gpt-492.8official
2gpt-3.5-turbo87.75official
3chatglm2-lite87.65official
4internlm-chat-7b-v1.185.5official
5qwen-7b-chat84.45official
6internlm-chat-7b84.3official
7baichuan-2-13b-chat83.65official
8baichuan-13b-chat83.15official
9text-davinci-00382.35official
10chatglm2-6b82official
11vicuna-33b79.7official
12wizardlm-13b79.4official
13flan-t5-xxl77.9official
14vicuna-13b77.5official
15llama-2-13b-chat73.6official
16openchat-13b73.1official
16vicuna-7b73.1official
18llama-2-7b-chat70.2official
19llama-2-chinese-13b-chat69.4official
20llama-2-chinese-7b-chat61.7official
21wizardlm-7b60.7official

OFF

Measures whether the model selects safe answers to bilingual multiple-choice questions about offensive language and content.

RankModelValueRelative performanceProvenance
1gpt-486.15official
2flan-t5-xxl79.2official
3gpt-3.5-turbo77.4official
4text-davinci-00373.2official
5chatglm2-lite70.7official
6baichuan-2-13b-chat69.25official
7qwen-7b-chat69.1official
8vicuna-13b68.4official
9wizardlm-13b68.3official
10chatglm2-6b68.1official
11internlm-chat-7b-v1.167.35official
12internlm-chat-7b67.2official
13vicuna-33b66.7official
14vicuna-7b65.1official
15baichuan-13b-chat59.25official
16openchat-13b52.6official
16wizardlm-7b52.6official
18llama-2-7b-chat48.9official
18llama-2-chinese-7b-chat48.9official
20llama-2-13b-chat48.4official
21llama-2-chinese-13b-chat48.1official

PH

Measures whether the model selects safe answers to bilingual multiple-choice questions about physical health and bodily harm.

RankModelValueRelative performanceProvenance
1gpt-494.35official
2chatglm2-lite79.65official
2gpt-3.5-turbo79.65official
4internlm-chat-7b-v1.176.65official
5baichuan-2-13b-chat76.35official
6text-davinci-00374.8official
7internlm-chat-7b74.15official
8vicuna-33b73official
9qwen-7b-chat70.4official
10wizardlm-13b69.4official
11baichuan-13b-chat68.2official
12chatglm2-6b67.9official
13flan-t5-xxl67official
14vicuna-13b65.3official
15vicuna-7b60.9official
16llama-2-13b-chat60.7official
17openchat-13b59.9official
18llama-2-7b-chat54.5official
19wizardlm-7b52.4official
20llama-2-chinese-13b-chat49.7official
21llama-2-chinese-7b-chat43official

PP

Measures whether the model selects safe answers to bilingual multiple-choice questions about privacy and property.

RankModelValueRelative performanceProvenance
1gpt-491official
2gpt-3.5-turbo85.65official
3chatglm2-lite82.65official
4baichuan-2-13b-chat82.05official
5internlm-chat-7b-v1.180.7official
6text-davinci-00380.2official
7internlm-chat-7b78.7official
8qwen-7b-chat77.2official
9baichuan-13b-chat77official
10flan-t5-xxl76.4official
11chatglm2-6b76official
12vicuna-13b75.4official
13wizardlm-13b75official
14openchat-13b71.1official
14vicuna-33b71.1official
16llama-2-13b-chat70.1official
17vicuna-7b68.4official
18llama-2-7b-chat65official
19llama-2-chinese-13b-chat64.7official
20llama-2-chinese-7b-chat57.6official
21wizardlm-7b55.8official

UB

Measures whether the model selects safe answers to bilingual multiple-choice questions about unfairness and social bias.

RankModelValueRelative performanceProvenance
1gpt-477.9official
2flan-t5-xxl70.2official
3wizardlm-13b69.6official
4gpt-3.5-turbo67.9official
5internlm-chat-7b-v1.166.3official
5llama-2-13b-chat66.3official
7qwen-7b-chat65.9official
8internlm-chat-7b64.75official
9llama-2-7b-chat63.2official
10baichuan-13b-chat62.65official
11openchat-13b62.6official
12chatglm2-6b61.6official
13llama-2-chinese-7b-chat61.3official
14text-davinci-00360.45official
15chatglm2-lite59.15official
16vicuna-33b56.8official
17llama-2-chinese-13b-chat54.4official
18vicuna-13b53official
19vicuna-7b52.7official
20baichuan-2-13b-chat49.2official
21wizardlm-7b48.8official