← Evals

Evaluation profile

JailBench

1sub-evals
0.093%Safety weight
0.667%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

Jailbreak susceptibility across Chinese safety categories.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
jailbreak_success_ratejailbench/jailbench.csv:jailbreak_success_rateMeasures whether jailbreak attacks can make the model comply with harmful requests.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000jailbench/jailbench.csvSafety: lower · Freedom: higher0.093%0.667%

jailbreak_success_rate

Measures whether jailbreak attacks can make the model comply with harmful requests.

RankModelValueRelative performanceProvenance
1gpt-434.5official
2llama-3-8b-instruct43.94official
3llama-2-7b-chat48.68official
4qwen2-7b-chat49.73official
5internlm2-chat-7b51.22official
6llama-2-13b-chat55.39official
7chatglm358.81official
8vicuna-7b-v1.559.27official
9vicuna-13b-v1.566.32official
10deepseek-v370.67official
11qwen1.5-7b-chat71.6official
12gpt-3.5-turbo73.86official
13glm-4-9b-chat74.95official
14mistral-7b-instruct78.14official