Evaluation profile
M3-SafetyBench
1sub-evals
0.139%total index weight
1components
Within-component eval weight: Misuse resistance 1.39%.
Model score (higher is better)Predicted score
About this eval
Multidimensional Chinese safety performance including red-team and refusal ability dimensions.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| overall_scorem3-safetybench/m3-safetybench.csv:overall_scoreMeasures safety knowledge and safe response selection across multilingual, multimodal, and multi-level risk scenarios. | ordinary_harm_misuse_resistance:1.000m3-safetybench/m3-safetybench.csv | Higher is better | 0.139% | Misuse resistance 1.39% |
overall_score
Measures safety knowledge and safe response selection across multilingual, multimodal, and multi-level risk scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | doubao-pro-32k | 96.76 | official | |
| 2 | qwen-plus | 96.53 | official | |
| 3 | qwen-2.5-14b-instruct | 95.94 | official | |
| 4 | moonshot-v1 | 95.03 | official | |
| 5 | ernie-3.5-128k | 94.9 | official | |
| 6 | hunyuan-standard | 94.87 | official | |
| 7 | internlm2.5-20b-chat | 92.88 | official | |
| 8 | qwen-2.5-7b-instruct | 92.37 | official | |
| 9 | qwen2-7b-instruct | 91.7 | official | |
| 10 | internlm2.5-7b-chat | 89.33 | official | |
| 11 | deepseek-r1-distill-qwen-7b | 88.09 | official | |
| 12 | chatglm3-6b-32k | 87.27 | official | |
| 13 | baichuan-2-13b-chat | 86.39 | official | |
| 14 | qwen-2.5-1.5b-instruct | 85.87 | official | |
| 15 | glm-4-9b-chat | 84.2 | official | |
| 16 | baichuan-2-7b-chat | 82.91 | official | |
| 17 | qwen2-1.5b-instruct | 80.83 | official | |
| 18 | internlm2.5-1.8b-chat | 79.16 | official | |
| 19 | qwen-2.5-0.5b-instruct | 69.63 | official |