← Evals

Evaluation profile

Open LLM Safety Index

2sub-evals
0.147%total index weight
1components

Within-component eval weight: Misuse resistance 1.47%.

Model score (higher is better)Predicted score

About this eval

Refusal-string robustness on two distinct fifteen-prompt harmful-request panels derived from JailbreakBench and StrongREJECT.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
jailbreakbench_safety_rateopen-llm-safety-index/open-llm-safety-index.csv:jailbreakbench_safety_rateMeasures whether the model refuses adversarial harmful prompts sampled from JailbreakBench.ordinary_harm_misuse_resistance:1.000open-llm-safety-indexHigher is better0.0733%Misuse resistance 0.733%
strongreject_safety_rateopen-llm-safety-index/open-llm-safety-index.csv:strongreject_safety_rateMeasures whether the model refuses adversarial harmful prompts sampled from StrongREJECT.ordinary_harm_misuse_resistance:1.000open-llm-safety-indexHigher is better0.0733%Misuse resistance 0.733%

jailbreakbench_safety_rate

Measures whether the model refuses adversarial harmful prompts sampled from JailbreakBench.

RankModelValueRelative performanceProvenance
1phi-3-mini-4k-instruct0.8official
1qwen1.5-0.5b-chat0.8official
3qwen2-1.5b-instruct0.6667official
4phi-3.5-mini-instruct0.6official
4qwen-2.5-1.5b-instruct0.6official
4qwen2.5-3b-instruct0.6official
7qwen2-0.5b-instruct0.4667official
7smollm2-135m-instruct0.4667official
7smollm2-360m-instruct0.4667official
10qwen1.5-1.8b-chat0.4official
10qwen2.5-coder-3b-instruct0.4official
12qwen-2.5-0.5b-instruct0.3333official
12smollm2-1.7b-instruct0.3333official
14llama-3.2-1b-instruct0.2667official
14llama-3.2-3b-instruct0.2667official
16phi-20.1333official
16qwen2.5-coder-1.5b-instruct0.1333official
16qwen2.5-math-1.5b-instruct0.1333official
19qwen2.5-coder-0.5b-instruct0.0667official
20phi-1-50official
20qwen3-0.6b0official

strongreject_safety_rate

Measures whether the model refuses adversarial harmful prompts sampled from StrongREJECT.

RankModelValueRelative performanceProvenance
1qwen1.5-0.5b-chat1official
1smollm2-1.7b-instruct1official
1smollm2-135m-instruct1official
4smollm2-360m-instruct0.8667official
5llama-3.2-3b-instruct0.6667official
6qwen2-1.5b-instruct0.6official
7phi-3-mini-4k-instruct0.5333official
8phi-3.5-mini-instruct0.3333official
8qwen-2.5-1.5b-instruct0.3333official
8qwen2.5-3b-instruct0.3333official
11qwen1.5-1.8b-chat0.2667official
12qwen2.5-coder-0.5b-instruct0.2official
12qwen2.5-coder-3b-instruct0.2official
14llama-3.2-1b-instruct0.1333official
14qwen2-0.5b-instruct0.1333official
16qwen-2.5-0.5b-instruct0.0667official
16qwen2.5-coder-1.5b-instruct0.0667official
16qwen2.5-math-1.5b-instruct0.0667official
19phi-1-50official
19phi-20official
19qwen3-0.6b0official