← Evals

Evaluation profile

DSPSafeBench

1sub-evals
0.172%Safety weight
1.24%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (higher is better)Predicted score

About this eval

Aggregate compliance rate on adversarial Chinese content-safety prompts.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
scoredspsafebench/dspsafebench.csv:scoreMeasures whether the model produces responses that comply with the benchmark’s safety criteria across diverse Chinese prompts.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000dspsafebench/dspsafebench.csvSafety: higher · Freedom: lower0.172%1.24%

score

Measures whether the model produces responses that comply with the benchmark’s safety criteria across diverse Chinese prompts.

RankModelValueRelative performanceProvenance
1yi-1.5-9b-chat-16k79.37official
2phi-3-mini-4k-instruct78.62official
3internlm2.5-7b-chat77.64official
3minicpm3-4b77.64official
5qwen2.5-7b-instruct73.51official
6mistral-7b-instruct73.04official
7glm-4-9b-chat72.43official
8gemma-2-9b-it72.34official
9qwen2.5-1.5b-instruct71.26official
10gemma-2-2b-it70.42official
11baichuan-2-7b-chat65.31official
12llama-3.1-8b-instruct61.51official