Evaluation profile
RefusalBench
1sub-evals
0.674%Safety weight
0%Freedom weight
2components
Weights below are portfolio-specific global index weights.
Model score (higher is better)Predicted score
About this eval
Refusal calibration on biological research prompts.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| youden_jrefusalbench/refusalbench.csv:youden_jMeasures whether the model distinguishes dual-use biological requests from matched benign requests rather than refusing indiscriminately. | Safety: benign_helpfulness_non_overrefusal:0.500|ordinary_harm_misuse_resistance:0.500refusalbench/refusalbench.csv | Safety: higher | 0.674% | — |
youden_j
Measures whether the model distinguishes dual-use biological requests from matched benign requests rather than refusing indiscriminately.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | grok-4.20 | 0.7871 | official | |
| 2 | claude-sonnet-4.6 | 0.6766 | official | |
| 3 | claude-opus-4.5 | 0.6709 | official | |
| 4 | claude-opus-4.6 | 0.6667 | official | |
| 5 | gpt-5.4-mini | 0.3787 | official | |
| 6 | gpt-5.5 | 0.2979 | official | |
| 7 | claude-opus-4.7 | 0.234 | official | |
| 8 | gemini-3.1-pro-preview | 0.1319 | official | |
| 9 | glm-5 | 0.08511 | official | |
| 10 | minimax-m2.7 | 0.08085 | official | |
| 11 | qwen3-next-80b-a3b | 0.06383 | official | |
| 12 | nemotron-3-super-120b-a12b | 0.06383 | official | |
| 13 | kimi-k2.6 | 0.0383 | official | |
| 14 | gemini-3.1-flash-lite | 0.02979 | official | |
| 15 | nova-pro | 0.01282 | official | |
| 16 | deepseek-v3.2 | 0.01277 | official | |
| 17 | llama-3.3-70b-instruct | 0.004255 | official | |
| 18 | mistral-large-3 | 0.004237 | official | |
| 19 | deepseek-r1 | -0.004255 | official |