Evaluation profile
Do-Not-Answer
1sub-evals
0.117%total index weight
1components
Within-component eval weight: Misuse resistance 1.18%.
Model score (higher is better)Predicted score
About this eval
Harmless refusal behavior on 939 prompts that responsible models should not answer.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| human_harmlessness_ratedo-not-answer/do-not-answer.csv:human_harmlessness_rateMeasures whether responses to sensitive requests avoid harms identified by human annotators. | ordinary_harm_misuse_resistance:1.000do-not-answer/do-not-answer.csv | Higher is better | 0.117% | Misuse resistance 1.18% |
human_harmlessness_rate
Measures whether responses to sensitive requests avoid harms identified by human annotators.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | llama-2-7b-chat | 99.68 | official | |
| 2 | gpt-3.5-turbo | 98.51 | official | |
| 3 | claude-1 | 98.3 | official | |
| 4 | gpt-4 | 97.55 | official | |
| 5 | vicuna-7b | 94.46 | official | |
| 6 | chatglm2-6b | 90.95 | official |