Evaluation profile
DSPSafeBench
1sub-evals
0.222%total index weight
1components
Within-component eval weight: Misuse resistance 2.22%.
Model score (higher is better)Predicted score
About this eval
Aggregate compliance rate on adversarial Chinese content-safety prompts.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| scoredspsafebench/dspsafebench.csv:scoreMeasures whether the model produces responses that comply with the benchmark’s safety criteria across diverse Chinese prompts. | ordinary_harm_misuse_resistance:1.000dspsafebench/dspsafebench.csv | Higher is better | 0.222% | Misuse resistance 2.22% |
score
Measures whether the model produces responses that comply with the benchmark’s safety criteria across diverse Chinese prompts.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | yi-1.5-9b-chat-16k | 79.37 | official | |
| 2 | phi-3-mini-4k-instruct | 78.62 | official | |
| 3 | internlm2.5-7b-chat | 77.64 | official | |
| 3 | minicpm3-4b | 77.64 | official | |
| 5 | qwen-2.5-7b-instruct | 73.51 | official | |
| 6 | mistral-7b-instruct | 73.04 | official | |
| 7 | glm-4-9b-chat | 72.43 | official | |
| 8 | gemma-2-9b-it | 72.34 | official | |
| 9 | qwen-2.5-1.5b-instruct | 71.26 | official | |
| 10 | gemma-2-2b-it | 70.42 | official | |
| 11 | baichuan-2-7b-chat | 65.31 | official | |
| 12 | llama-3.1-8b-instruct | 61.51 | official |