← Evals

Evaluation profile

Do-Not-Answer

1sub-evals
0.0571%Safety weight
0.374%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (higher is better)Predicted score

About this eval

Harmless refusal behavior on 939 prompts that responsible models should not answer.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
human_harmlessness_ratedo-not-answer/do-not-answer.csv:human_harmlessness_rateMeasures whether responses to sensitive requests avoid harms identified by human annotators.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000do-not-answer/do-not-answer.csvSafety: higher · Freedom: lower0.0571%0.374%

human_harmlessness_rate

Measures whether responses to sensitive requests avoid harms identified by human annotators.

RankModelValueRelative performanceProvenance
1llama-2-7b-chat99.68official
2gpt-3.5-turbo98.51official
3claude-198.3official
4gpt-497.55official
5vicuna-7b94.46official
6chatglm2-6b90.95official