← Evals

Evaluation profile

Human Pathogen Capabilities Test (HPCT) — overall refusal

1sub-evals
0.207%Safety weight
1.48%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (higher is better)Predicted score

About this eval

Refusal propensity on practical assistance questions concerning a small set of weaponizable human pathogens judged especially high-concern for misuse.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
refusal_rate_pcthpct-refusal/hpct-refusal.csv:refusal_rate_pctMeasures how often a deployed model endpoint refuses a predominantly refusal-recommended set of practical questions about especially high-concern human pathogens.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000hpct-refusalSafety: higher · Freedom: lower0.207%1.48%

refusal_rate_pct

Measures how often a deployed model endpoint refuses a predominantly refusal-recommended set of practical questions about especially high-concern human pathogens.

RankModelValueRelative performanceProvenance
1claude-opus-4.797.9official
2claude-opus-494.2official
3claude-opus-4.889.2official
4claude-sonnet-4.584.4official
5claude-opus-4.681.9official
6claude-opus-4.169.1official
7claude-opus-4.556.7official
8kimi-k2.652.9official
9gpt-5.448.3official
10gpt-5.247.6official
11gpt-5.547.3official
12gpt-5-pro46.9official
13gpt-526.5official
14gpt-5.125.9official
15glm-5.224.8official
16grok-4.1-fast24.55official
17grok-4-fast19.5official
18o113.5official
19gpt-5-nano3.4official
20gemini-3.5-flash0.8official
21glm-50.4official
21gpt-5-mini0.4official
23kimi-k2.50.3official
23qwen3.5-397b-a17b0.3official
23qwen3.6-plus0.3official
26qwen3.7-max0.2official
27claude-20official
27claude-3-opus0official
27claude-3.5-haiku0official
27claude-3.5-sonnet0official
27claude-haiku-4.50official
27claude-sonnet-40official
27deepseek-r10official
27deepseek-v3.10official
27deepseek-v4-pro0official
27gemini-2.5-flash0official
27gemini-2.5-pro0official
27gemini-3-flash-preview0official
27gemini-3-pro-preview0official
27gemini-3.1-pro-preview0official
27gpt-3.5-turbo0official
27gpt-40official
27gpt-4-turbo0official
27gpt-4.10official
27gpt-4.1-mini0official
27gpt-4.1-nano0official
27gpt-4.5-preview0official
27gpt-4o0official
27gpt-4o-mini0official
27gpt-5.5-pro0official
27gpt-oss-120b0official
27grok-3-beta0official
27grok-3-mini0official
27grok-40official
27kimi-k20official
27kimi-k30official
27llama-4-maverick0official
27magistral-medium0official
27magistral-small0official
27mistral-large-20official
27mistral-small-3.10official
27mixtral-8x7b-instruct0official
27o30official
27o3-mini0official
27o4-mini0official
27qwen3-235b-a22b0official
27qwen3-next-80b-a3b0official
27qwen3-next-80b-a3b-instruct0official
27sonar-deep-research0official