Evaluation profile
CASE-Bench
1sub-evals
0.0363%total index weight
1components
Within-component eval weight: Misuse resistance 0.363%.
Model score (higher is better)Predicted score
About this eval
Agreement with human judgments about whether a model should respond under explicitly safe or unsafe conversational contexts.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| agreement_accuracycase-bench/case-bench.csv:agreement_accuracyMeasures whether the model’s safety ratings of contextualized requests agree with majority human safe-or-unsafe labels. | ordinary_harm_misuse_resistance:1.000sorrybench | Higher is better | 0.0363% | Misuse resistance 0.363% |
agreement_accuracy
Measures whether the model’s safety ratings of contextualized requests agree with majority human safe-or-unsafe labels.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3.5-sonnet | 90.11 | official | |
| 2 | qwen2-72b-instruct | 84.67 | official | |
| 3 | llama-3-70b-instruct | 84.44 | official | |
| 4 | mixtral-8x7b-instruct | 83.11 | official | |
| 5 | dolphin-2.9-llama3-70b | 80.78 | official | |
| 6 | gpt-4o-mini | 80 | official | |
| 7 | gpt-4o | 78.56 | official |