Evaluation profile
Do-Not-Answer
1sub-evals
0.0571%Safety weight
0.374%Freedom weight
1components
Weights below are portfolio-specific global index weights.
Model score (higher is better)Predicted score
About this eval
Harmless refusal behavior on 939 prompts that responsible models should not answer.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| human_harmlessness_ratedo-not-answer/do-not-answer.csv:human_harmlessness_rateMeasures whether responses to sensitive requests avoid harms identified by human annotators. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000do-not-answer/do-not-answer.csv | Safety: higher · Freedom: lower | 0.0571% | 0.374% |
human_harmlessness_rate
Measures whether responses to sensitive requests avoid harms identified by human annotators.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | llama-2-7b-chat | 99.68 | official | |
| 2 | gpt-3.5-turbo | 98.51 | official | |
| 3 | claude-1 | 98.3 | official | |
| 4 | gpt-4 | 97.55 | official | |
| 5 | vicuna-7b | 94.46 | official | |
| 6 | chatglm2-6b | 90.95 | official |