← Models

Model profile

GPT 5.4 Pro

OpenAIdeveloper
2026-03-05release date
#60 / 333Safety rank
#600 / 645Freedom rank

Evidence summary

Safety. GPT 5.4 Pro has an estimated Safety rank of #60; its 90% source-sensitivity interval is #1–#275. Its behavior-only rank is #70; company governance moves the combined estimate to #60. Published Safety evidence spans 4 eval lineages and 3 of 7 components. Its strongest relative result is MASK (lying_probability_pct, #3 of 53); its weakest is FORTRESS (over_refusal_score, #38 of 49).

Freedom. GPT 5.4 Pro has an estimated Freedom rank of #600; its 90% source-sensitivity interval is #339–#636. Published Freedom evidence spans 1 eval lineages and 1 of 1 components. Its strongest relative result is FORTRESS (over_refusal_score, #38 of 49); its weakest is FORTRESS (average_risk_score, #45 of 50).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
FORTRESSaverage_risk_score#6 / 5014.84Source ↗official
FORTRESSover_refusal_score#38 / 499.41Source ↗official
Humanity's Last Exam RMS calibration error (Scale Labs)calibrationError#3 / 4238Source ↗official
MASKlying_probability_pct#3 / 538.27Source ↗official
Vectara HHEM Factual Consistencyfactual_consistency_rate#38 / 9491.7Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
FORTRESSaverage_risk_score#45 / 5014.84Source ↗official
FORTRESSover_refusal_score#38 / 499.41Source ↗official