Evaluation profile
Human Pathogen Capabilities Test (HPCT) — overall refusal
1sub-evals
0.207%Safety weight
1.48%Freedom weight
1components
Weights below are portfolio-specific global index weights.
Model score (higher is better)Predicted score
About this eval
Refusal propensity on practical assistance questions concerning a small set of weaponizable human pathogens judged especially high-concern for misuse.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| refusal_rate_pcthpct-refusal/hpct-refusal.csv:refusal_rate_pctMeasures how often a deployed model endpoint refuses a predominantly refusal-recommended set of practical questions about especially high-concern human pathogens. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000hpct-refusal | Safety: higher · Freedom: lower | 0.207% | 1.48% |
refusal_rate_pct
Measures how often a deployed model endpoint refuses a predominantly refusal-recommended set of practical questions about especially high-concern human pathogens.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.7 | 97.9 | official | |
| 2 | claude-opus-4 | 94.2 | official | |
| 3 | claude-opus-4.8 | 89.2 | official | |
| 4 | claude-sonnet-4.5 | 84.4 | official | |
| 5 | claude-opus-4.6 | 81.9 | official | |
| 6 | claude-opus-4.1 | 69.1 | official | |
| 7 | claude-opus-4.5 | 56.7 | official | |
| 8 | kimi-k2.6 | 52.9 | official | |
| 9 | gpt-5.4 | 48.3 | official | |
| 10 | gpt-5.2 | 47.6 | official | |
| 11 | gpt-5.5 | 47.3 | official | |
| 12 | gpt-5-pro | 46.9 | official | |
| 13 | gpt-5 | 26.5 | official | |
| 14 | gpt-5.1 | 25.9 | official | |
| 15 | glm-5.2 | 24.8 | official | |
| 16 | grok-4.1-fast | 24.55 | official | |
| 17 | grok-4-fast | 19.5 | official | |
| 18 | o1 | 13.5 | official | |
| 19 | gpt-5-nano | 3.4 | official | |
| 20 | gemini-3.5-flash | 0.8 | official | |
| 21 | glm-5 | 0.4 | official | |
| 21 | gpt-5-mini | 0.4 | official | |
| 23 | kimi-k2.5 | 0.3 | official | |
| 23 | qwen3.5-397b-a17b | 0.3 | official | |
| 23 | qwen3.6-plus | 0.3 | official | |
| 26 | qwen3.7-max | 0.2 | official | |
| 27 | claude-2 | 0 | official | |
| 27 | claude-3-opus | 0 | official | |
| 27 | claude-3.5-haiku | 0 | official | |
| 27 | claude-3.5-sonnet | 0 | official | |
| 27 | claude-haiku-4.5 | 0 | official | |
| 27 | claude-sonnet-4 | 0 | official | |
| 27 | deepseek-r1 | 0 | official | |
| 27 | deepseek-v3.1 | 0 | official | |
| 27 | deepseek-v4-pro | 0 | official | |
| 27 | gemini-2.5-flash | 0 | official | |
| 27 | gemini-2.5-pro | 0 | official | |
| 27 | gemini-3-flash-preview | 0 | official | |
| 27 | gemini-3-pro-preview | 0 | official | |
| 27 | gemini-3.1-pro-preview | 0 | official | |
| 27 | gpt-3.5-turbo | 0 | official | |
| 27 | gpt-4 | 0 | official | |
| 27 | gpt-4-turbo | 0 | official | |
| 27 | gpt-4.1 | 0 | official | |
| 27 | gpt-4.1-mini | 0 | official | |
| 27 | gpt-4.1-nano | 0 | official | |
| 27 | gpt-4.5-preview | 0 | official | |
| 27 | gpt-4o | 0 | official | |
| 27 | gpt-4o-mini | 0 | official | |
| 27 | gpt-5.5-pro | 0 | official | |
| 27 | gpt-oss-120b | 0 | official | |
| 27 | grok-3-beta | 0 | official | |
| 27 | grok-3-mini | 0 | official | |
| 27 | grok-4 | 0 | official | |
| 27 | kimi-k2 | 0 | official | |
| 27 | kimi-k3 | 0 | official | |
| 27 | llama-4-maverick | 0 | official | |
| 27 | magistral-medium | 0 | official | |
| 27 | magistral-small | 0 | official | |
| 27 | mistral-large-2 | 0 | official | |
| 27 | mistral-small-3.1 | 0 | official | |
| 27 | mixtral-8x7b-instruct | 0 | official | |
| 27 | o3 | 0 | official | |
| 27 | o3-mini | 0 | official | |
| 27 | o4-mini | 0 | official | |
| 27 | qwen3-235b-a22b | 0 | official | |
| 27 | qwen3-next-80b-a3b | 0 | official | |
| 27 | qwen3-next-80b-a3b-instruct | 0 | official | |
| 27 | sonar-deep-research | 0 | official |