← Evals

Evaluation profile

SOSBench

6sub-evals
0.23%total index weight
1components

Within-component eval weight: Misuse resistance 2.3%.

Model score (lower is better)Predicted score

About this eval

Scientific misuse or unsafe open-science behavior.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
biology_pvrsosbench/sosbench.csv:biology_pvrMeasures whether the model avoids policy-violating assistance on high-risk prompts in biology.ordinary_harm_misuse_resistance:1.000sosbench/sosbench.csvLower is better0.0383%Misuse resistance 0.383%
chemistry_pvrsosbench/sosbench.csv:chemistry_pvrMeasures whether the model avoids policy-violating assistance on high-risk prompts in chemistry.ordinary_harm_misuse_resistance:1.000sosbench/sosbench.csvLower is better0.0383%Misuse resistance 0.383%
medicine_pvrsosbench/sosbench.csv:medicine_pvrMeasures whether the model avoids policy-violating assistance on high-risk prompts in medicine.ordinary_harm_misuse_resistance:1.000sosbench/sosbench.csvLower is better0.0383%Misuse resistance 0.383%
pharmacology_pvrsosbench/sosbench.csv:pharmacology_pvrMeasures whether the model avoids policy-violating assistance on high-risk prompts in pharmacology.ordinary_harm_misuse_resistance:1.000sosbench/sosbench.csvLower is better0.0383%Misuse resistance 0.383%
physics_pvrsosbench/sosbench.csv:physics_pvrMeasures whether the model avoids policy-violating assistance on high-risk prompts in physics.ordinary_harm_misuse_resistance:1.000sosbench/sosbench.csvLower is better0.0383%Misuse resistance 0.383%
psychology_pvrsosbench/sosbench.csv:psychology_pvrMeasures whether the model avoids policy-violating assistance on high-risk prompts in psychology.ordinary_harm_misuse_resistance:1.000sosbench/sosbench.csvLower is better0.0383%Misuse resistance 0.383%

biology_pvr

Measures whether the model avoids policy-violating assistance on high-risk prompts in biology.

RankModelValueRelative performanceProvenance
1claude-sonnet-40.104official
2gpt-50.108official
3claude-opus-4.10.134official
4o30.156official
5claude-3.7-sonnet0.229official
6o4-mini0.262official
7llama-4-maverick0.288official
8gpt-4o0.306official
9gemini-2.5-flash0.336official
10gemini-2.5-pro0.354official
11gpt-4.10.374official
12llama-3.3-70b-instruct0.408official
13llama-4-scout0.488official
14llama-3.1-405b-instruct0.59official
15qwen2.5-72b-instruct0.68official
16grok-3-mini0.758official
17gemma-3-27b-it0.792official
18qwen3-32b0.802official
19deepseek-r10.814official
20deepseek-r1-distill-llama-70b0.838official
21qwen3-235b-a22b0.852official
22deepseek-v30.856official
23grok-30.894official

chemistry_pvr

Measures whether the model avoids policy-violating assistance on high-risk prompts in chemistry.

RankModelValueRelative performanceProvenance
1gpt-50.122official
2claude-opus-4.10.147official
3o30.152official
4o4-mini0.206official
5claude-3.7-sonnet0.208official
6claude-sonnet-40.21official
7llama-4-maverick0.238official
8gpt-4o0.254official
9gpt-4.10.314official
10gemini-2.5-flash0.338official
11gemini-2.5-pro0.342official
12llama-4-scout0.436official
13llama-3.1-405b-instruct0.468official
14llama-3.3-70b-instruct0.54official
15qwen2.5-72b-instruct0.56official
16grok-3-mini0.586official
17deepseek-v30.6official
18grok-30.638official
19gemma-3-27b-it0.646official
20qwen3-235b-a22b0.76official
21qwen3-32b0.784official
22deepseek-r10.834official
23deepseek-r1-distill-llama-70b0.904official

medicine_pvr

Measures whether the model avoids policy-violating assistance on high-risk prompts in medicine.

RankModelValueRelative performanceProvenance
1claude-sonnet-40.213official
2claude-opus-4.10.232official
3gpt-50.332official
4claude-3.7-sonnet0.35official
5o30.372official
6llama-4-maverick0.426official
7gemini-2.5-flash0.462official
7o4-mini0.462official
9gpt-4o0.476official
10gemini-2.5-pro0.492official
11llama-3.3-70b-instruct0.546official
12gpt-4.10.57official
13llama-4-scout0.688official
14llama-3.1-405b-instruct0.69official
15qwen2.5-72b-instruct0.734official
16grok-3-mini0.746official
17qwen3-32b0.774official
18deepseek-r10.806official
19gemma-3-27b-it0.814official
20deepseek-r1-distill-llama-70b0.854official
21grok-30.86official
22qwen3-235b-a22b0.868official
23deepseek-v30.872official

pharmacology_pvr

Measures whether the model avoids policy-violating assistance on high-risk prompts in pharmacology.

RankModelValueRelative performanceProvenance
1claude-sonnet-40.234official
2claude-opus-4.10.249official
3o4-mini0.408official
4gpt-50.418official
5o30.424official
6claude-3.7-sonnet0.579official
7gemini-2.5-pro0.634official
8llama-4-maverick0.652official
9gpt-4o0.676official
10gemini-2.5-flash0.684official
11llama-3.1-405b-instruct0.764official
12llama-3.3-70b-instruct0.812official
13gpt-4.10.85official
14llama-4-scout0.874official
15deepseek-v30.916official
16qwen2.5-72b-instruct0.926official
17grok-3-mini0.93official
18gemma-3-27b-it0.934official
18qwen3-235b-a22b0.934official
20qwen3-32b0.946official
21grok-30.954official
22deepseek-r10.964official
23deepseek-r1-distill-llama-70b0.972official

physics_pvr

Measures whether the model avoids policy-violating assistance on high-risk prompts in physics.

RankModelValueRelative performanceProvenance
1claude-opus-4.10.098official
2gpt-50.104official
3o30.114official
4claude-sonnet-40.145official
5claude-3.7-sonnet0.171official
6gpt-4o0.194official
7o4-mini0.22official
8llama-4-maverick0.24official
9gpt-4.10.41official
10gemini-2.5-flash0.424official
11llama-3.1-405b-instruct0.444official
12gemini-2.5-pro0.466official
13llama-4-scout0.492official
14llama-3.3-70b-instruct0.516official
15qwen2.5-72b-instruct0.678official
16grok-3-mini0.708official
17deepseek-v30.722official
18qwen3-32b0.74official
19qwen3-235b-a22b0.764official
20grok-30.804official
21gemma-3-27b-it0.842official
22deepseek-r10.872official
23deepseek-r1-distill-llama-70b0.886official

psychology_pvr

Measures whether the model avoids policy-violating assistance on high-risk prompts in psychology.

RankModelValueRelative performanceProvenance
1claude-opus-4.10.107official
2claude-sonnet-40.123official
3gpt-50.142official
4claude-3.7-sonnet0.168official
5o30.196official
6llama-4-maverick0.242official
7gemini-2.5-pro0.294official
8o4-mini0.314official
9gemini-2.5-flash0.326official
10gpt-4o0.396official
11llama-3.3-70b-instruct0.446official
12gpt-4.10.498official
13llama-4-scout0.51official
14llama-3.1-405b-instruct0.568official
15grok-3-mini0.7official
16qwen2.5-72b-instruct0.734official
17qwen3-32b0.746official
18gemma-3-27b-it0.792official
19deepseek-r10.806official
20deepseek-r1-distill-llama-70b0.816official
21deepseek-v30.82official
22qwen3-235b-a22b0.852official
23grok-30.89official