← Evals

Evaluation profile

ChineseSafe

1sub-evals
0.15%total index weight
1components

Within-component eval weight: Misuse resistance 1.5%.

Model score (higher is better)Predicted score

About this eval

Accuracy at classifying safe versus unsafe Chinese content under the paper's perplexity-based evaluation.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
scorechinesesafe/chinesesafe.csv:scoreMeasures whether the model correctly classifies Chinese prompts and responses as safe or unsafe using the benchmark’s perplexity method.ordinary_harm_misuse_resistance:1.000chinesesafe/chinesesafe.csvHigher is better0.15%Misuse resistance 1.5%

score

Measures whether the model correctly classifies Chinese prompts and responses as safe or unsafe using the benchmark’s perplexity method.

RankModelValueRelative performanceProvenance
1baichuan-2-13b-chat70.43official
2deepseek-llm-67b-chat68.08official
3yi-1.5-34b-chat66.02official
4gemma-1.1-7b-it64.32official
5qwen1.5-72b-chat63.67official
6qwen1.5-7b-chat62.48official
7qwen1.5-14b-chat61.29official
8yi-1.5-9b-chat60.35official
9opt-66b59.93official
10deepseek-llm-7b-chat56.79official
11gpt-j-6b55.98official
12ziya2-13b-chat55.25official
13baichuan-2-7b-chat53.99official
14opt-30b53.82official
15internlm2-chat-20b53.67official
16glm-4-9b-chat50.03official
17internlm2-chat-7b49.49official
18opt-13b49.31official
19opt-6.7b48.54official
20mistral-7b-instruct42.99official
21llama3-chatqa-1.5-8b42.11official
22llama3-chatqa-1.5-70b40.41official