← Evals

Evaluation profile

BioSecBench-Refusal

1sub-evals
0.489%Safety weight
0%Freedom weight
2components

Weights below are portfolio-specific global index weights.

Model score (higher is better)Predicted score

About this eval

Calibrated refusal across legitimate biological work and concealed biosecurity hazards.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
balanced_refusal_scorebiosecbench-refusal/biosecbench-refusal.csv:balanced_refusal_scoreMeasures whether the model helps with routine biological research while refusing concealed requests that could enable biological harm.Safety: benign_helpfulness_non_overrefusal:0.500|ordinary_harm_misuse_resistance:0.500biosecbench-refusalSafety: higher0.489%

balanced_refusal_score

Measures whether the model helps with routine biological research while refusing concealed requests that could enable biological harm.

RankModelValueRelative performanceProvenance
1gemini-3.5-flash0.5041official
2claude-sonnet-4.60.4304official
3claude-opus-4.70.4259official
4gpt-5.40.4254official
5claude-opus-4.60.3987official
6gemini-3.1-pro-preview0.3931official
7claude-opus-4.80.3633official
8gpt-5.50.2869official
9grok-4.30.1092official
10grok-4.200.02854official