Evaluation profile
ChiSafetyBench
4sub-evals
0.222%Safety weight
0.667%Freedom weight
5components
Weights below are portfolio-specific global index weights.
Model score (see each sub-eval direction)Predicted score
About this eval
Chinese safety knowledge and refusal behavior.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| harmful_response_ratechisafetybench/chisafetybench.csv:harmful_response_rateMeasures whether the model avoids producing harmful responses to open-ended Chinese safety prompts. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000chisafetybench/chisafetybench.csv | Safety: lower · Freedom: higher | 0.0997% | 0.222% |
| mcq_scorechisafetybench/chisafetybench.csv:mcq_scoreMeasures whether the model correctly answers Chinese multiple-choice questions about safety risks and appropriate conduct. | Safety: human_rights_systemic_harm:0.227|fairness_nondiscrimination:0.294|truthfulness_honesty:0.063|ordinary_harm_misuse_resistance:0.416chisafetybench/chisafetybench.csv | Safety: higher | 0.122% | — |
| refusal_rr1chisafetybench/chisafetybench.csv:refusal_rr1Measures first refusal rate metric from the refusal table. | Freedom: user_direction_following:1.000chisafetybench | Freedom: lower | — | 0.222% |
| refusal_rr2chisafetybench/chisafetybench.csv:refusal_rr2Measures second refusal rate metric from the refusal table. | Freedom: user_direction_following:1.000chisafetybench | Freedom: lower | — | 0.222% |
harmful_response_rate
Measures whether the model avoids producing harmful responses to open-ended Chinese safety prompts.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | qwen1.5-110b-chat | 0 | official | |
| 2 | qwen1.5-14b-chat | 0.22 | official | |
| 2 | qwen1.5-32b-chat | 0.22 | official | |
| 2 | qwen1.5-72b-chat | 0.22 | official | |
| 5 | baichuan-2-13b-chat | 0.43 | official | |
| 5 | deepseek-v3 | 0.43 | official | |
| 5 | qwen1.5-7b-chat | 0.43 | official | |
| 8 | baichuan-2-7b-chat | 0.65 | official | |
| 8 | deepseek-r1 | 0.65 | official | |
| 8 | yi-34b-chat | 0.65 | official | |
| 11 | yi-6b-chat | 0.87 | official | |
| 12 | chatglm2-6b | 1.08 | official | |
| 12 | chatglm3-6b | 1.08 | official | |
| 14 | chatglm-6b | 1.95 | official |
mcq_score
Measures whether the model correctly answers Chinese multiple-choice questions about safety risks and appropriate conduct.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | qwen1.5-72b-chat | 91.13 | official | |
| 2 | qwen1.5-110b-chat | 90.62 | official | |
| 3 | qwen1.5-14b-chat | 88.39 | official | |
| 4 | qwen1.5-32b-chat | 87.3 | official | |
| 5 | yi-6b-chat | 86.01 | official | |
| 6 | deepseek-v3 | 84.17 | official | |
| 7 | baichuan-2-13b-chat | 79.83 | official | |
| 8 | qwen1.5-7b-chat | 79.39 | official | |
| 9 | deepseek-r1 | 71.41 | official | |
| 10 | yi-34b-chat | 68.54 | official | |
| 11 | baichuan-2-7b-chat | 50.03 | official | |
| 12 | chatglm3-6b | 41.16 | official |
refusal_rr1
Measures first refusal rate metric from the refusal table.
| Rank | Model | Value | Relative performance | Provenance |
|---|
refusal_rr2
Measures second refusal rate metric from the refusal table.
| Rank | Model | Value | Relative performance | Provenance |
|---|