← Evals

Evaluation profile

BioSecBench-Refusal

1sub-evals
0.685%total index weight
2components

Within-component eval weight: Benign helpfulness 6.1% · Misuse resistance 0.758%.

Model score (higher is better)Predicted score

About this eval

Calibrated refusal across legitimate biological work and concealed biosecurity hazards.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
balanced_refusal_scorebiosecbench-refusal/biosecbench-refusal.csv:balanced_refusal_scoreMeasures whether the model helps with routine biological research while refusing concealed requests that could enable biological harm.benign_helpfulness_non_overrefusal:0.500|ordinary_harm_misuse_resistance:0.500biosecbench-refusalHigher is better0.685%Benign helpfulness 6.1% · Misuse resistance 0.758%

balanced_refusal_score

Measures whether the model helps with routine biological research while refusing concealed requests that could enable biological harm.

RankModelValueRelative performanceProvenance
1gemini-3.5-flash0.5041official
2claude-sonnet-4.60.4304official
3claude-opus-4.70.4259official
4gpt-5.40.4254official
5claude-opus-4.60.3987official
6gemini-3.1-pro-preview0.3931official
7claude-opus-4.80.3633official
8gpt-5.50.2869official
9grok-4.30.1092official
10grok-4.200.02854official