← Evals

Evaluation profile

JailBench

1sub-evals
0.12%total index weight
1components

Within-component eval weight: Misuse resistance 1.2%.

Model score (lower is better)Predicted score

About this eval

Jailbreak susceptibility across Chinese safety categories.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
jailbreak_success_ratejailbench/jailbench.csv:jailbreak_success_rateMeasures whether jailbreak attacks can make the model comply with harmful requests.ordinary_harm_misuse_resistance:1.000jailbench/jailbench.csvLower is better0.12%Misuse resistance 1.2%

jailbreak_success_rate

Measures whether jailbreak attacks can make the model comply with harmful requests.

RankModelValueRelative performanceProvenance
1gpt-434.5official
2llama-3-8b-instruct43.94official
3llama-2-7b-chat48.68official
4qwen2-7b-chat49.73official
5internlm2-chat-7b51.22official
6llama-2-13b-chat55.39official
7chatglm358.81official
8vicuna-7b-v1.559.27official
9vicuna-13b-v1.566.32official
10deepseek-v370.67official
11qwen1.5-7b-chat71.6official
12gpt-3.5-turbo73.86official
13glm-4-9b-chat74.95official
14mistral-7b-instruct78.14official