← Evals

Evaluation profile

SafeDialBench

6sub-evals
0.686%total index weight
3components

Within-component eval weight: Misuse resistance 1.02% · Human rights 2.77% · Fairness 1.68%.

Model score (higher is better)Predicted score

About this eval

Dialogue safety across aggression, ethics, fairness, legality, morality, and privacy dimensions.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
aggressionsafedialbench/safedialbench.csv:aggressionMeasures whether the model maintains safe behavior during multi-turn jailbreak conversations about aggression and violence.ordinary_harm_misuse_resistance:1.000safedialbench/safedialbench.csvHigher is better0.0339%Misuse resistance 0.339%
ethicssafedialbench/safedialbench.csv:ethicsMeasures whether the model maintains safe behavior during multi-turn jailbreak conversations about self-harm and abuse.human_rights_systemic_harm:1.000safedialbench/safedialbench.csvHigher is better0.208%Human rights 1.39%
fairnesssafedialbench/safedialbench.csv:fairnessMeasures whether the model maintains safe behavior during multi-turn jailbreak conversations about stereotypes and distributional harms.fairness_nondiscrimination:1.000safedialbench/safedialbench.csvHigher is better0.168%Fairness 1.68%
legalitysafedialbench/safedialbench.csv:legalityMeasures whether the model maintains safe behavior during multi-turn jailbreak conversations about illegal harmful conduct.ordinary_harm_misuse_resistance:1.000safedialbench/safedialbench.csvHigher is better0.0339%Misuse resistance 0.339%
moralitysafedialbench/safedialbench.csv:moralityMeasures whether the model maintains safe behavior during multi-turn jailbreak conversations about nonviolent immorality and social-norm violations.ordinary_harm_misuse_resistance:1.000safedialbench/safedialbench.csvHigher is better0.0339%Misuse resistance 0.339%
privacysafedialbench/safedialbench.csv:privacyMeasures whether the model maintains safe behavior during multi-turn jailbreak conversations about personal, organizational, and social privacy violations.human_rights_systemic_harm:1.000safedialbench/safedialbench.csvHigher is better0.208%Human rights 1.39%

aggression

Measures whether the model maintains safe behavior during multi-turn jailbreak conversations about aggression and violence.

RankModelValueRelative performanceProvenance
1moonshot-v17.277official
2deepseek-r17.273official
3yi-34b-chat7.26official
4gpt-4o7.207official
5qwq-32b7.19official
6glm-4-9b-chat7.17official
7qwen-14b-chat7.15official
8qwen-7b-chat7.127official
8yi-6b-chat7.127official
10qwen-2.5-14b-instruct7.123official
11baichuan-2-7b-chat7.073official
12baichuan-2-13b-chat7.03official
13mistral-7b-instruct7.027official
14o3-mini7.02official
15chatglm3-6b7.017official
16qwen-2.5-7b-instruct7.013official
17internlm2-20b-sft6.957official
18deepseek-llm-7b-chat6.953official

ethics

Measures whether the model maintains safe behavior during multi-turn jailbreak conversations about self-harm and abuse.

RankModelValueRelative performanceProvenance
1glm-4-9b-chat7.753official
2qwen-14b-chat7.68official
2yi-34b-chat7.68official
4chatglm3-6b7.637official
5qwen-7b-chat7.623official
6baichuan-2-7b-chat7.613official
7baichuan-2-13b-chat7.6official
8mistral-7b-instruct7.587official
9yi-6b-chat7.577official
10deepseek-llm-7b-chat7.563official
11internlm2-20b-sft7.547official
12gpt-4o7.487official
13o3-mini7.403official
14qwen-2.5-14b-instruct7.39official
15qwen-2.5-7b-instruct7.357official
16moonshot-v17.353official
17qwq-32b7.313official
18deepseek-r17.303official

fairness

Measures whether the model maintains safe behavior during multi-turn jailbreak conversations about stereotypes and distributional harms.

RankModelValueRelative performanceProvenance
1moonshot-v17.7official
2gpt-4o7.68official
3deepseek-r17.607official
4qwq-32b7.6official
5qwen-2.5-14b-instruct7.56official
6o3-mini7.557official
7qwen-2.5-7b-instruct7.553official
8glm-4-9b-chat7.4official
9yi-34b-chat7.337official
10yi-6b-chat7.277official
11qwen-14b-chat7.27official
12qwen-7b-chat7.19official
13chatglm3-6b7.187official
13mistral-7b-instruct7.187official
15baichuan-2-13b-chat7.17official
16baichuan-2-7b-chat7.123official
17deepseek-llm-7b-chat7.117official
18internlm2-20b-sft7.1official

legality

Measures whether the model maintains safe behavior during multi-turn jailbreak conversations about illegal harmful conduct.

RankModelValueRelative performanceProvenance
1yi-34b-chat8.117official
2glm-4-9b-chat8.103official
3mistral-7b-instruct7.993official
4baichuan-2-13b-chat7.987official
4qwen-14b-chat7.987official
6chatglm3-6b7.983official
7internlm2-20b-sft7.977official
8deepseek-llm-7b-chat7.947official
9baichuan-2-7b-chat7.937official
10qwen-7b-chat7.93official
11yi-6b-chat7.887official
12deepseek-r17.38official
13moonshot-v17.317official
14qwq-32b7.263official
15gpt-4o7.21official
15qwen-2.5-7b-instruct7.21official
17o3-mini7.193official
18qwen-2.5-14b-instruct7.17official

morality

Measures whether the model maintains safe behavior during multi-turn jailbreak conversations about nonviolent immorality and social-norm violations.

RankModelValueRelative performanceProvenance
1yi-34b-chat7.52official
2qwen-14b-chat7.437official
3glm-4-9b-chat7.427official
4baichuan-2-7b-chat7.383official
5moonshot-v17.35official
6baichuan-2-13b-chat7.303official
7mistral-7b-instruct7.28official
8qwen-7b-chat7.277official
9chatglm3-6b7.243official
10gpt-4o7.237official
11internlm2-20b-sft7.213official
12deepseek-llm-7b-chat7.183official
12deepseek-r17.183official
14qwq-32b7.16official
15yi-6b-chat7.123official
16qwen-2.5-7b-instruct7.06official
17qwen-2.5-14b-instruct7.047official
18o3-mini7.007official

privacy

Measures whether the model maintains safe behavior during multi-turn jailbreak conversations about personal, organizational, and social privacy violations.

RankModelValueRelative performanceProvenance
1yi-34b-chat7.88official
2glm-4-9b-chat7.853official
3qwen-14b-chat7.69official
4yi-6b-chat7.67official
5baichuan-2-13b-chat7.617official
6mistral-7b-instruct7.607official
7internlm2-20b-sft7.597official
8deepseek-llm-7b-chat7.587official
9chatglm3-6b7.567official
10qwen-7b-chat7.563official
11baichuan-2-7b-chat7.523official
12moonshot-v17.25official
13deepseek-r17.193official
14gpt-4o7.14official
15qwen-2.5-14b-instruct7.117official
16qwq-32b7.093official
17o3-mini7.077official
18qwen-2.5-7b-instruct7.05official