← Evals

Evaluation profile

RefusalBench

1sub-evals
0.674%Safety weight
0%Freedom weight
2components

Weights below are portfolio-specific global index weights.

Model score (higher is better)Predicted score

About this eval

Refusal calibration on biological research prompts.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
youden_jrefusalbench/refusalbench.csv:youden_jMeasures whether the model distinguishes dual-use biological requests from matched benign requests rather than refusing indiscriminately.Safety: benign_helpfulness_non_overrefusal:0.500|ordinary_harm_misuse_resistance:0.500refusalbench/refusalbench.csvSafety: higher0.674%

youden_j

Measures whether the model distinguishes dual-use biological requests from matched benign requests rather than refusing indiscriminately.

RankModelValueRelative performanceProvenance
1grok-4.200.7871official
2claude-sonnet-4.60.6766official
3claude-opus-4.50.6709official
4claude-opus-4.60.6667official
5gpt-5.4-mini0.3787official
6gpt-5.50.2979official
7claude-opus-4.70.234official
8gemini-3.1-pro-preview0.1319official
9glm-50.08511official
10minimax-m2.70.08085official
11qwen3-next-80b-a3b0.06383official
12nemotron-3-super-120b-a12b0.06383official
13kimi-k2.60.0383official
14gemini-3.1-flash-lite0.02979official
15nova-pro0.01282official
16deepseek-v3.20.01277official
17llama-3.3-70b-instruct0.004255official
18mistral-large-30.004237official
19deepseek-r1-0.004255official