Evaluation profile
SuperCLUE Safety
Within-component eval weight: Misuse resistance 1.94% · Human rights 2.61% · Fairness 0.452%.
About this eval
Chinese model safety across traditional safety, responsibility, and instruction-attack dimensions.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| instruction_attackSC-safety/sc-safety.csv:instruction_attackMeasures whether the model maintains safe and substantively helpful open-ended responses under adversarial instructions intended to bypass its safety behavior, including paired follow-up turns. | ordinary_harm_misuse_resistance:1.000SC-safety/sc-safety.csv | Higher is better | 0.121% | Misuse resistance 1.21% |
| responsible_aiSC-safety/sc-safety.csv:responsible_aiMeasures whether the model generates safe and substantively helpful open-ended Chinese responses reflecting broader responsible-AI values, including social responsibility and concern for vulnerable groups. | human_rights_systemic_harm:0.375|fairness_nondiscrimination:0.125|ordinary_harm_misuse_resistance:0.500SC-safety/sc-safety.csv | Higher is better | 0.249% | Human rights 1.12% · Fairness 0.452% · Misuse resistance 0.364% |
| traditional_safetySC-safety/sc-safety.csv:traditional_safetyMeasures whether the model generates safe and substantively helpful answers to open-ended Chinese prompts and follow-ups about harmful, illegal, privacy-invasive, and health-related requests. | human_rights_systemic_harm:0.500|ordinary_harm_misuse_resistance:0.500SC-safety/sc-safety.csv | Higher is better | 0.26% | Human rights 1.49% · Misuse resistance 0.364% |
instruction_attack
Measures whether the model maintains safe and substantively helpful open-ended responses under adversarial instructions intended to bypass its safety behavior, including paired follow-up turns.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-4-turbo | 82.76 | official | |
| 2 | gpt-4 | 77.59 | official | |
| 3 | baichuan-2-13b-chat | 72.41 | official | |
| 3 | minimax-abab-5.5 | 72.41 | official | |
| 3 | sensechat-3.0 | 72.41 | official | |
| 3 | spark3.0 | 72.41 | official | |
| 3 | yi-34b-chat | 72.41 | official | |
| 8 | claude-2 | 70.69 | official | |
| 9 | 360gpt-pro | 68.97 | official | |
| 9 | chatglm3-6b | 68.97 | official | |
| 9 | gpt-3.5-turbo | 68.97 | official | |
| 9 | internlm-chat-20b | 68.97 | official | |
| 9 | qwen-14b-chat | 68.97 | official | |
| 9 | qwen-72b-chat | 68.97 | official | |
| 15 | baichuan-2-7b-chat | 65.52 | official | |
| 15 | bluelm-7b-chat | 65.52 | official | |
| 17 | aquilachat2-34b | 62.07 | official | |
| 17 | chinese-alpaca-2-13b | 62.07 | official | |
| 17 | mistral-7b-instruct | 62.07 | official | |
| 20 | chatglm2-6b | 58.62 | official | |
| 20 | internlm-chat-7b | 58.62 | official | |
| 20 | openbuddy-mistral-7b | 58.62 | official | |
| 20 | openbuddy-zephyr-7b | 58.62 | official | |
| 24 | andestron-13b-chat | 55.17 | official | |
| 24 | chatglm-turbo | 55.17 | official | |
| 24 | chinese-alpaca-2-7b | 55.17 | official | |
| 24 | qwen-1.8b-chat | 55.17 | official | |
| 24 | qwen-7b-chat | 55.17 | official | |
| 29 | baichuan-13b-chat | 51.72 | official | |
| 30 | andestron-7b-chat | 48.28 | official | |
| 31 | moss-moon-003-sft | 44.83 | official |
responsible_ai
Measures whether the model generates safe and substantively helpful open-ended Chinese responses reflecting broader responsible-AI values, including social responsibility and concern for vulnerable groups.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-2 | 78.18 | official | |
| 1 | gpt-4-turbo | 78.18 | official | |
| 3 | gpt-4 | 74.55 | official | |
| 4 | baichuan-2-13b-chat | 72.73 | official | |
| 4 | gpt-3.5-turbo | 72.73 | official | |
| 6 | qwen-72b-chat | 70.91 | official | |
| 7 | bluelm-7b-chat | 69.09 | official | |
| 7 | yi-34b-chat | 69.09 | official | |
| 9 | minimax-abab-5.5 | 67.27 | official | |
| 9 | sensechat-3.0 | 67.27 | official | |
| 11 | internlm-chat-20b | 65.45 | official | |
| 11 | spark3.0 | 65.45 | official | |
| 13 | baichuan-2-7b-chat | 61.82 | official | |
| 13 | qwen-14b-chat | 61.82 | official | |
| 15 | chatglm-turbo | 58.18 | official | |
| 15 | openbuddy-mistral-7b | 58.18 | official | |
| 17 | chatglm3-6b | 56.36 | official | |
| 18 | aquilachat2-34b | 54.55 | official | |
| 18 | mistral-7b-instruct | 54.55 | official | |
| 20 | 360gpt-pro | 52.73 | official | |
| 20 | openbuddy-zephyr-7b | 52.73 | official | |
| 20 | qwen-7b-chat | 52.73 | official | |
| 23 | chinese-alpaca-2-13b | 49.09 | official | |
| 24 | andestron-13b-chat | 47.27 | official | |
| 24 | andestron-7b-chat | 47.27 | official | |
| 24 | chatglm2-6b | 47.27 | official | |
| 24 | qwen-1.8b-chat | 47.27 | official | |
| 28 | internlm-chat-7b | 45.45 | official | |
| 29 | moss-moon-003-sft | 43.64 | official | |
| 30 | chinese-alpaca-2-7b | 41.82 | official | |
| 31 | baichuan-13b-chat | 40 | official |
traditional_safety
Measures whether the model generates safe and substantively helpful answers to open-ended Chinese prompts and follow-ups about harmful, illegal, privacy-invasive, and health-related requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | baichuan-2-13b-chat | 87.23 | official | |
| 1 | openbuddy-zephyr-7b | 87.23 | official | |
| 3 | aquilachat2-34b | 82.98 | official | |
| 3 | sensechat-3.0 | 82.98 | official | |
| 5 | 360gpt-pro | 80.85 | official | |
| 5 | baichuan-13b-chat | 80.85 | official | |
| 5 | baichuan-2-7b-chat | 80.85 | official | |
| 5 | minimax-abab-5.5 | 80.85 | official | |
| 5 | qwen-72b-chat | 80.85 | official | |
| 10 | chatglm-turbo | 78.72 | official | |
| 10 | chatglm2-6b | 78.72 | official | |
| 10 | qwen-7b-chat | 78.72 | official | |
| 10 | yi-34b-chat | 78.72 | official | |
| 14 | claude-2 | 77.66 | official | |
| 15 | bluelm-7b-chat | 76.6 | official | |
| 16 | gpt-4-turbo | 75.53 | official | |
| 17 | andestron-13b-chat | 74.47 | official | |
| 17 | chinese-alpaca-2-13b | 74.47 | official | |
| 17 | openbuddy-mistral-7b | 74.47 | official | |
| 20 | gpt-3.5-turbo | 73.4 | official | |
| 21 | chinese-alpaca-2-7b | 72.34 | official | |
| 21 | qwen-14b-chat | 72.34 | official | |
| 23 | gpt-4 | 70.21 | official | |
| 23 | qwen-1.8b-chat | 70.21 | official | |
| 25 | andestron-7b-chat | 68.09 | official | |
| 25 | internlm-chat-20b | 68.09 | official | |
| 25 | mistral-7b-instruct | 68.09 | official | |
| 28 | chatglm3-6b | 65.96 | official | |
| 28 | internlm-chat-7b | 65.96 | official | |
| 28 | spark3.0 | 65.96 | official | |
| 31 | moss-moon-003-sft | 63.83 | official |