← Evals

Evaluation profile

ChiSafetyBench

4sub-evals
0.222%Safety weight
0.667%Freedom weight
5components

Weights below are portfolio-specific global index weights.

Model score (see each sub-eval direction)Predicted score

About this eval

Chinese safety knowledge and refusal behavior.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
harmful_response_ratechisafetybench/chisafetybench.csv:harmful_response_rateMeasures whether the model avoids producing harmful responses to open-ended Chinese safety prompts.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000chisafetybench/chisafetybench.csvSafety: lower · Freedom: higher0.0997%0.222%
mcq_scorechisafetybench/chisafetybench.csv:mcq_scoreMeasures whether the model correctly answers Chinese multiple-choice questions about safety risks and appropriate conduct.Safety: human_rights_systemic_harm:0.227|fairness_nondiscrimination:0.294|truthfulness_honesty:0.063|ordinary_harm_misuse_resistance:0.416chisafetybench/chisafetybench.csvSafety: higher0.122%
refusal_rr1chisafetybench/chisafetybench.csv:refusal_rr1Measures first refusal rate metric from the refusal table.Freedom: user_direction_following:1.000chisafetybenchFreedom: lower0.222%
refusal_rr2chisafetybench/chisafetybench.csv:refusal_rr2Measures second refusal rate metric from the refusal table.Freedom: user_direction_following:1.000chisafetybenchFreedom: lower0.222%

harmful_response_rate

Measures whether the model avoids producing harmful responses to open-ended Chinese safety prompts.

RankModelValueRelative performanceProvenance
1qwen1.5-110b-chat0official
2qwen1.5-14b-chat0.22official
2qwen1.5-32b-chat0.22official
2qwen1.5-72b-chat0.22official
5baichuan-2-13b-chat0.43official
5deepseek-v30.43official
5qwen1.5-7b-chat0.43official
8baichuan-2-7b-chat0.65official
8deepseek-r10.65official
8yi-34b-chat0.65official
11yi-6b-chat0.87official
12chatglm2-6b1.08official
12chatglm3-6b1.08official
14chatglm-6b1.95official

mcq_score

Measures whether the model correctly answers Chinese multiple-choice questions about safety risks and appropriate conduct.

RankModelValueRelative performanceProvenance
1qwen1.5-72b-chat91.13official
2qwen1.5-110b-chat90.62official
3qwen1.5-14b-chat88.39official
4qwen1.5-32b-chat87.3official
5yi-6b-chat86.01official
6deepseek-v384.17official
7baichuan-2-13b-chat79.83official
8qwen1.5-7b-chat79.39official
9deepseek-r171.41official
10yi-34b-chat68.54official
11baichuan-2-7b-chat50.03official
12chatglm3-6b41.16official

refusal_rr1

Measures first refusal rate metric from the refusal table.

RankModelValueRelative performanceProvenance

refusal_rr2

Measures second refusal rate metric from the refusal table.

RankModelValueRelative performanceProvenance