Evaluation profile
ChiSafetyBench
2sub-evals
0.293%total index weight
4components
Within-component eval weight: Misuse resistance 1.48% · Human rights 0.441% · Fairness 0.693% · Truthfulness 0.067%.
Model score (see each sub-eval direction)Predicted score
About this eval
Chinese safety knowledge and refusal behavior.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| harmful_response_ratechisafetybench/chisafetybench.csv:harmful_response_rateMeasures whether the model avoids producing harmful responses to open-ended Chinese safety prompts. | ordinary_harm_misuse_resistance:1.000chisafetybench/chisafetybench.csv | Lower is better | 0.128% | Misuse resistance 1.28% |
| mcq_scorechisafetybench/chisafetybench.csv:mcq_scoreMeasures whether the model correctly answers Chinese multiple-choice questions about safety risks and appropriate conduct. | human_rights_systemic_harm:0.227|fairness_nondiscrimination:0.294|truthfulness_honesty:0.063|ordinary_harm_misuse_resistance:0.416chisafetybench/chisafetybench.csv | Higher is better | 0.165% | Human rights 0.441% · Fairness 0.693% · Truthfulness 0.067% · Misuse resistance 0.198% |
harmful_response_rate
Measures whether the model avoids producing harmful responses to open-ended Chinese safety prompts.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | qwen1.5-110b-chat | 0 | official | |
| 2 | qwen1.5-14b-chat | 0.22 | official | |
| 2 | qwen1.5-32b-chat | 0.22 | official | |
| 2 | qwen1.5-72b-chat | 0.22 | official | |
| 5 | baichuan-2-13b-chat | 0.43 | official | |
| 5 | deepseek-v3 | 0.43 | official | |
| 5 | qwen1.5-7b-chat | 0.43 | official | |
| 8 | baichuan-2-7b-chat | 0.65 | official | |
| 8 | deepseek-r1 | 0.65 | official | |
| 8 | yi-34b-chat | 0.65 | official | |
| 11 | yi-6b-chat | 0.87 | official | |
| 12 | chatglm2-6b | 1.08 | official | |
| 12 | chatglm3-6b | 1.08 | official | |
| 14 | chatglm-6b | 1.95 | official |
mcq_score
Measures whether the model correctly answers Chinese multiple-choice questions about safety risks and appropriate conduct.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | qwen1.5-72b-chat | 91.13 | official | |
| 2 | qwen1.5-110b-chat | 90.62 | official | |
| 3 | qwen1.5-14b-chat | 88.39 | official | |
| 4 | qwen1.5-32b-chat | 87.3 | official | |
| 5 | yi-6b-chat | 86.01 | official | |
| 6 | deepseek-v3 | 84.17 | official | |
| 7 | baichuan-2-13b-chat | 79.83 | official | |
| 8 | qwen1.5-7b-chat | 79.39 | official | |
| 9 | deepseek-r1 | 71.41 | official | |
| 10 | yi-34b-chat | 68.54 | official | |
| 11 | baichuan-2-7b-chat | 50.03 | official | |
| 12 | chatglm3-6b | 41.16 | official |