Evaluation profile
RefusalBench
1sub-evals
0.945%total index weight
2components
Within-component eval weight: Benign helpfulness 8.4% · Misuse resistance 1.05%.
Model score (higher is better)Predicted score
About this eval
Refusal calibration on biological research prompts.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| youden_jrefusalbench/refusalbench.csv:youden_jMeasures whether the model refuses harmful prompts without unnecessarily refusing benign prompts. | benign_helpfulness_non_overrefusal:0.500|ordinary_harm_misuse_resistance:0.500refusalbench/refusalbench.csv | Higher is better | 0.945% | Benign helpfulness 8.4% · Misuse resistance 1.05% |
youden_j
Measures whether the model refuses harmful prompts without unnecessarily refusing benign prompts.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | grok-4.20 | 0.9701 | official | |
| 2 | gemini-3.1-pro-preview | 0.9217 | official | |
| 3 | nemotron-3-super-120b-a12b | 0.8891 | official | |
| 4 | glm-5 | 0.7291 | official | |
| 5 | claude-sonnet-4.6 | 0.6766 | official | |
| 6 | claude-opus-4.5 | 0.6709 | official | |
| 7 | claude-opus-4.6 | 0.6667 | official | |
| 8 | minimax-m2.7 | 0.6604 | official | |
| 9 | qwen3-next-80b-a3b | 0.4545 | official | |
| 10 | gemini-3.1-flash-lite | 0.4533 | official | |
| 11 | gpt-5.5 | 0.4213 | official | |
| 12 | claude-opus-4.7 | 0.234 | official | |
| 13 | deepseek-r1 | 0.1424 | official | |
| 14 | nova-pro | 0.09333 | official | |
| 15 | kimi-k2.6 | 0.08511 | official | |
| 16 | gpt-5.4-mini | 0.06411 | official | |
| 17 | deepseek-v3.2 | 0.01333 | official | |
| 17 | llama-3.3-70b-instruct | 0.01333 | official | |
| 19 | mistral-large-3 | 0.00906 | official |