← Evals

Evaluation profile

Do-Not-Answer

1sub-evals
0.117%total index weight
1components

Within-component eval weight: Misuse resistance 1.18%.

Model score (higher is better)Predicted score

About this eval

Harmless refusal behavior on 939 prompts that responsible models should not answer.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
human_harmlessness_ratedo-not-answer/do-not-answer.csv:human_harmlessness_rateMeasures whether responses to sensitive requests avoid harms identified by human annotators.ordinary_harm_misuse_resistance:1.000do-not-answer/do-not-answer.csvHigher is better0.117%Misuse resistance 1.18%

human_harmlessness_rate

Measures whether responses to sensitive requests avoid harms identified by human annotators.

RankModelValueRelative performanceProvenance
1llama-2-7b-chat99.68official
2gpt-3.5-turbo98.51official
3claude-198.3official
4gpt-497.55official
5vicuna-7b94.46official
6chatglm2-6b90.95official