← Models

Model profile

GPT 5 Pro

OpenAIdeveloper
2025-08-07release date
#39 / 333Safety rank
#634 / 645Freedom rank

Evidence summary

Safety. GPT 5 Pro has an estimated Safety rank of #39; its 90% source-sensitivity interval is #2–#277. Its behavior-only rank is #44; company governance moves the combined estimate to #39. Published Safety evidence spans 5 eval lineages and 4 of 7 components. Its strongest relative result is SpeciEval (sea_animal_4ns, #9 of 123); its weakest is FORTRESS (over_refusal_score, #44 of 49).

Freedom. GPT 5 Pro has an estimated Freedom rank of #634; its 90% source-sensitivity interval is #344–#639. Published Freedom evidence spans 2 eval lineages and 1 of 1 components. Its strongest relative result is Human Pathogen Capabilities Test (HPCT) — overall refusal (refusal_rate_pct, #58 of 69); its weakest is FORTRESS (over_refusal_score, #44 of 49).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
FORTRESSaverage_risk_score#8 / 5015.2Source ↗official
FORTRESSover_refusal_score#44 / 4912.43Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#12 / 6946.9Source ↗official
Humanity's Last Exam RMS calibration error (Scale Labs)calibrationError#9 / 4249Source ↗official
MASKlying_probability_pct#11 / 5314.01Source ↗official
SpeciEvalbelief_animal_sentience#53 / 1236.83Source ↗official
SpeciEvalland_animal_4ns#21 / 1234.25Source ↗official
SpeciEvalsea_animal_4ns#9 / 1234.25Source ↗official
SpeciEvalspeciesism#15 / 1231.45Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
FORTRESSaverage_risk_score#43 / 5015.2Source ↗official
FORTRESSover_refusal_score#44 / 4912.43Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#58 / 6946.9Source ↗official