Evaluation profile
HarmBench
1sub-evals
0.0846%total index weight
1components
Within-component eval weight: Misuse resistance 0.846%.
Model score (lower is better)Predicted score
About this eval
Harmful compliance or attack success under harmful request benchmarks.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| drharmbench/harmbench.csv:drMeasures whether the model refuses direct requests for assistance with harmful behavior. | ordinary_harm_misuse_resistance:1.000harmbench | Lower is better | 0.0846% | Misuse resistance 0.846% |
dr
Measures whether the model refuses direct requests for assistance with harmful behavior.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | llama-2-7b-chat | 0.8 | official | |
| 2 | claude-2 | 2 | official | |
| 2 | claude-2.1 | 2 | official | |
| 4 | llama-2-13b-chat | 2.8 | official | |
| 4 | llama-2-70b-chat | 2.8 | official | |
| 6 | claude-1 | 5 | official | |
| 7 | gpt-4-turbo | 9.3 | official | |
| 8 | qwen-7b-chat | 13 | official | |
| 9 | r2d2 | 14.2 | official | |
| 10 | qwen-14b-chat | 16.5 | official | |
| 11 | gemini-1.0-pro | 18 | official | |
| 12 | qwen-72b-chat | 18.3 | official | |
| 13 | baichuan-2-7b-chat | 18.8 | official | |
| 14 | baichuan-2-13b-chat | 19.3 | official | |
| 15 | vicuna-13b-v1.5 | 19.8 | official | |
| 16 | gpt-4 | 21 | official | |
| 17 | vicuna-7b-v1.5 | 24.3 | official | |
| 18 | gpt-3.5-turbo | 27.15 | official | |
| 19 | koala-13b | 27.3 | official | |
| 20 | koala-7b | 38.3 | official | |
| 21 | orca-2-7b | 39 | official | |
| 22 | orca-2-13b | 44.5 | official | |
| 23 | openchat-3.5 | 46 | official | |
| 24 | mistral-7b-instruct | 46.3 | official | |
| 25 | mixtral-8x7b-instruct | 47.3 | official | |
| 26 | starling-7b | 57 | official | |
| 27 | solar-10.7b-instruct | 61.3 | official | |
| 28 | zephyr-7b-beta | 65.8 | official |