← Evals

Evaluation profile

SuperCLUE Safety

3sub-evals
0.63%total index weight
3components

Within-component eval weight: Misuse resistance 1.94% · Human rights 2.61% · Fairness 0.452%.

Model score (higher is better)Predicted score

About this eval

Chinese model safety across traditional safety, responsibility, and instruction-attack dimensions.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
instruction_attackSC-safety/sc-safety.csv:instruction_attackMeasures whether the model maintains safe and substantively helpful open-ended responses under adversarial instructions intended to bypass its safety behavior, including paired follow-up turns.ordinary_harm_misuse_resistance:1.000SC-safety/sc-safety.csvHigher is better0.121%Misuse resistance 1.21%
responsible_aiSC-safety/sc-safety.csv:responsible_aiMeasures whether the model generates safe and substantively helpful open-ended Chinese responses reflecting broader responsible-AI values, including social responsibility and concern for vulnerable groups.human_rights_systemic_harm:0.375|fairness_nondiscrimination:0.125|ordinary_harm_misuse_resistance:0.500SC-safety/sc-safety.csvHigher is better0.249%Human rights 1.12% · Fairness 0.452% · Misuse resistance 0.364%
traditional_safetySC-safety/sc-safety.csv:traditional_safetyMeasures whether the model generates safe and substantively helpful answers to open-ended Chinese prompts and follow-ups about harmful, illegal, privacy-invasive, and health-related requests.human_rights_systemic_harm:0.500|ordinary_harm_misuse_resistance:0.500SC-safety/sc-safety.csvHigher is better0.26%Human rights 1.49% · Misuse resistance 0.364%

instruction_attack

Measures whether the model maintains safe and substantively helpful open-ended responses under adversarial instructions intended to bypass its safety behavior, including paired follow-up turns.

RankModelValueRelative performanceProvenance
1gpt-4-turbo82.76official
2gpt-477.59official
3baichuan-2-13b-chat72.41official
3minimax-abab-5.572.41official
3sensechat-3.072.41official
3spark3.072.41official
3yi-34b-chat72.41official
8claude-270.69official
9360gpt-pro68.97official
9chatglm3-6b68.97official
9gpt-3.5-turbo68.97official
9internlm-chat-20b68.97official
9qwen-14b-chat68.97official
9qwen-72b-chat68.97official
15baichuan-2-7b-chat65.52official
15bluelm-7b-chat65.52official
17aquilachat2-34b62.07official
17chinese-alpaca-2-13b62.07official
17mistral-7b-instruct62.07official
20chatglm2-6b58.62official
20internlm-chat-7b58.62official
20openbuddy-mistral-7b58.62official
20openbuddy-zephyr-7b58.62official
24andestron-13b-chat55.17official
24chatglm-turbo55.17official
24chinese-alpaca-2-7b55.17official
24qwen-1.8b-chat55.17official
24qwen-7b-chat55.17official
29baichuan-13b-chat51.72official
30andestron-7b-chat48.28official
31moss-moon-003-sft44.83official

responsible_ai

Measures whether the model generates safe and substantively helpful open-ended Chinese responses reflecting broader responsible-AI values, including social responsibility and concern for vulnerable groups.

RankModelValueRelative performanceProvenance
1claude-278.18official
1gpt-4-turbo78.18official
3gpt-474.55official
4baichuan-2-13b-chat72.73official
4gpt-3.5-turbo72.73official
6qwen-72b-chat70.91official
7bluelm-7b-chat69.09official
7yi-34b-chat69.09official
9minimax-abab-5.567.27official
9sensechat-3.067.27official
11internlm-chat-20b65.45official
11spark3.065.45official
13baichuan-2-7b-chat61.82official
13qwen-14b-chat61.82official
15chatglm-turbo58.18official
15openbuddy-mistral-7b58.18official
17chatglm3-6b56.36official
18aquilachat2-34b54.55official
18mistral-7b-instruct54.55official
20360gpt-pro52.73official
20openbuddy-zephyr-7b52.73official
20qwen-7b-chat52.73official
23chinese-alpaca-2-13b49.09official
24andestron-13b-chat47.27official
24andestron-7b-chat47.27official
24chatglm2-6b47.27official
24qwen-1.8b-chat47.27official
28internlm-chat-7b45.45official
29moss-moon-003-sft43.64official
30chinese-alpaca-2-7b41.82official
31baichuan-13b-chat40official

traditional_safety

Measures whether the model generates safe and substantively helpful answers to open-ended Chinese prompts and follow-ups about harmful, illegal, privacy-invasive, and health-related requests.

RankModelValueRelative performanceProvenance
1baichuan-2-13b-chat87.23official
1openbuddy-zephyr-7b87.23official
3aquilachat2-34b82.98official
3sensechat-3.082.98official
5360gpt-pro80.85official
5baichuan-13b-chat80.85official
5baichuan-2-7b-chat80.85official
5minimax-abab-5.580.85official
5qwen-72b-chat80.85official
10chatglm-turbo78.72official
10chatglm2-6b78.72official
10qwen-7b-chat78.72official
10yi-34b-chat78.72official
14claude-277.66official
15bluelm-7b-chat76.6official
16gpt-4-turbo75.53official
17andestron-13b-chat74.47official
17chinese-alpaca-2-13b74.47official
17openbuddy-mistral-7b74.47official
20gpt-3.5-turbo73.4official
21chinese-alpaca-2-7b72.34official
21qwen-14b-chat72.34official
23gpt-470.21official
23qwen-1.8b-chat70.21official
25andestron-7b-chat68.09official
25internlm-chat-20b68.09official
25mistral-7b-instruct68.09official
28chatglm3-6b65.96official
28internlm-chat-7b65.96official
28spark3.065.96official
31moss-moon-003-sft63.83official